Nuclear-targeted DNA delivery and compositions for use in its implementation
NTDNA with DTS enhances nuclear delivery and expression of cargo nucleic acids in non-dividing cells by utilizing DNA targeting sequences to translocate DNA to the nucleus, addressing limitations of existing methods.
Patent Information
- Application Number
- JP2025524734
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-24
- Filing Date
- 2023-10-25
- Publication Date
- 2025-11-07
AI Technical Summary
Existing nucleic acid delivery methods, such as AAV vectors, are limited by size constraints and induce antibody responses, making them unsuitable for delivering genes larger than 4.7 kB and ineffective in non-dividing or slowly dividing cells, while lipid nanoparticles only deliver cargo to the cytoplasm, failing to reach the nucleus in these cells.
The use of nuclear-targeted deoxyribonucleic acid (NTDNA) comprising a DNA nuclear-targeting sequence (DTS) to facilitate the transport of cargo nucleic acids from the cytosol to the nucleus, utilizing DNA targeting sequences (DTS) that are recognized by cytosolic DNA-binding proteins to translocate DNA cargo to the nucleus.
Enhances nuclear uptake of DNA by 2-fold to 200-fold, improving gene expression in non-dividing cells through methods like DTS-mediated delivery, achieving comparable or better results than traditional methods.
Smart Images

Figure 2025536570000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 63 / 419,890, filed October 27, 2022, U.S. Provisional Application No. 63 / 430,950, filed December 7, 2022, and U.S. Provisional Application No. 63 / 454,505, filed March 24, 2023, each of which is incorporated by reference in its entirety.
[0002] Introduction Nucleic acid delivery to the nucleus is desirable in many cases, including research, diagnostic, and therapeutic applications. An example of such a therapeutic application is gene therapy. In the field of gene therapy, viral vectors, such as vectors based on the virus AAV, are commonly used to deliver genes to the nucleus. However, due to the size limitation of the AAV genome, any gene larger than 4.7 kB is not suitable for use with AAV vectors, limiting the usefulness of such vectors for many indications. In addition, viral vectors such as AAV induce antibody responses, resulting in only one delivery, which is not suitable for some indications, such as liver indications, where cells divide slowly and lose the transduced genome, thereby requiring re-administration. Furthermore, viral vectors such as AAV are toxic at the doses required for therapeutic benefit in some indications. Therefore, what is needed is a new delivery vehicle for delivering nucleic acids, such as DNA, to cells in vitro during research, particularly in patients requiring gene therapy.
[0003] The next generation of delivery vehicles are nanoparticles. Lipid nanoparticles are particularly interesting, given how LNPs have been de-risked by their use in the Onpattro and Covid vaccines. The problem is that nanoparticles only deliver their cargo to the cytoplasm, in contrast to AAV's ability to carry its DNA cargo all the way to the nucleus. For dividing cells, delivery to the cytosol may be sufficient because the nuclear envelope dissolves during replication and new DNA is captured as it reforms. However, this is problematic for non-dividing or slowly dividing cells. What is needed to make LNPs a successful delivery vehicle in such cells is a mechanism that drives DNA transport from the cytoplasm to the nucleus. The inventors have identified methods and compositions to meet the above needs. Summary of the Invention [Means for solving the problem]
[0004] A method for nuclear-targeted DNA delivery is provided. This method embodiment involves contacting a cell with nuclear-targeted deoxyribonucleic acid (NTDNA), which comprises a DNA nuclear-targeting sequence (DTS) and a cargo nucleic acid heterologous to the DTS. Also provided are compositions for use in practicing the methods of the invention. [Brief explanation of the drawings]
[0005] [Figure 1] The premise behind using DTS to increase nuclear uptake is presented: DNA targeting sequences (DTS) encoded on DNA cargo reside in the cytoplasm and are recognized by cytosolic DNA-binding proteins, such as transcription factors, which can bind to and translocate the DNA cargo to the nucleus to facilitate gene expression. [Figure 2]Two library designs are shown for use in screening DTSs. DTS libraries are designed using nuclear targeting factor binding sites (NTFBSs, or more simply, TFBSs) in tandem arrays with spacer sequences of various lengths. In some cases (panel a), a single TFBS is arrayed to understand the contribution of that particular transcription factor to DNA translocation and expression. In other cases (panel b), TFBSs from different transcription factors are arrayed in various combinations to test the ability of cognate transcription factors to cooperate with each other. Putative DTSs are placed upstream of the promoter of an expression cassette (shown here), downstream of the polyA tail of an expression cassette, or on the cargo within an intron of the expression sequence. If desired, a unique molecular identifier (UMI), i.e., an expression barcode, can be included on the cargo within the expression cassette, for example, to screen for DTSs within the pool. DNA is then introduced into cells either in vitro, for example, in primary human hepatocytes or HepG2 cells, or in vivo, for example, in mice or NHPs. Cells are harvested after transfection, and the efficiency of DNA delivery is assessed by intracellular DNA copy number, RNA expression level, or protein quantification analysis. TFBSs selected for screening include those that are binding sites for transcription factors highly expressed in hepatocytes and transcription factors exploited by viruses. [Figure 3] We demonstrate the utility of implementing multiple in vitro experimental designs to identify DTS for in vivo evaluation. (a) Screening of the DTS library in non-dividing primary human hepatocytes allows for the identification of hits (open circles) by both quantitative live-cell imaging and FACS. (b) Screening of the DTS library in serum-starved, growth-arrested cells (HepG2 shown here) recapitulates data from HepG2 cells grown in the presence of serum, but with a larger dynamic range, allowing for the identification of hits (in vitro) with greater sensitivity (open circles). [Figure 4]We show that DTSs containing binding sites for NF-kB or derived from SV40 enhancer elements mediate robust nuclear import of NT DNA and expression of genes encoded by that NT DNA in non-dividing primary human hepatocytes. [Figure 5] We demonstrate the identification of several novel and robust DTSs through a one-by-one arrayed screen of a combinatorial DTS library. Plasmid NT DNA containing the combinatorial DTSs and a GFP reporter expression cassette was delivered to growth-arrested HepG2 cells, and GFP expression was assessed over time. Statistical and hierarchical clustering methods for analysis of kinetic experiments (Panel A), along with endpoint analysis of expression at a final time point of 24 hours post-transfection (Panel B), both reveal large increases in gene expression by the novel DTSs that are comparable to or better than those of the NFKB DTSs. DTS.203: CREB1, PPARA, ONECUT1, HNF4A, PPARA; DTS.233: HNF1A, NR1I3, PPARA, HNF1A, PPARA; DTS.276: FOXA1, HNF4A, HNF1A, HNF4A, HNF1A. [Figure 6] Figure 1 shows a strong correlation between DTSs identified by a pooled DTS screening approach (x-axis), which relies on detecting changes in mRNA abundance, and a single arrayed screening approach (y-axis), which relies on detecting changes in protein abundance. The study was performed by transfecting growth-arrested HepG2 cells with pDNA containing the DTSs. Evaluation of the pooled DTS screening was performed at the mRNA level using RNA-seq analysis to quantify each UMI, and evaluation of the arrayed DTS screening was performed by measuring GFP fluorescence intensity. [Figure 7A]This paper provides data on the activity of DTSs in mouse liver. Plasmids containing DTSs and barcodes embedded in expression cassettes were formulated together as a single formulation and administered to mice via bolus intravenous injection. mRNA levels were quantified using RNA-seq analysis of each barcode within the pool to identify the DTSs driving increased gene expression. (A) RNA abundance of expression cassettes associated with different DTSs in mouse liver 1 day after LNP administration. (B) RNA abundance of DTSs in mouse liver 4 days after LNP administration. [Figure 7B] Same as above [Figure 8] To identify the transcription factors most important in DTS activity, we present the results of a machine learning analysis of all DTS datasets. A gradient-boosted decision tree machine learning algorithm was applied to each individual dataset independently, and the most important features were then identified by the % improvement they provided to the algorithm's predictions. The same top three transcription factors were identified in each dataset, demonstrating their importance within the DTS element. [Figure 9] We demonstrate how an inducible DTS (iDTS) can further increase gene expression when its cognate binding protein is induced to translocate to the nucleus after stimulation with an inducer. In this study, dexamethasone, a known activator of glucocorticoid receptor (GR) relocalization to the nucleus, is used as the inducer. Growth-arrested HepG2 cells are transfected with the NTDNA plasmid containing the GR-responsive DTS and a GFP reporter expression cassette, and the cells are exposed to various amounts of dexamethasone and GFP expression is assessed. [Figure 10]We present data on several iDTSs identified by screening combinatorial libraries that promote significant cargo translocation and expression. Plasmid NTDNA containing a combined DTS and a GFP reporter expression cassette was delivered to growth-arrested HepG2 cells, and GFP expression was assessed 20 hours later. Note that activators do not induce the activity of NTDNAs containing binding sequences for constitutively active transcription factors, but they do induce the translocation of NTDNAs containing DTSs containing binding sequences for inducible transcription factors. iDTS-1-1: NR3C1, NR3C1, NR3C1, NR3C1, NR3C1; iDTS-1-2: HNF4A, NR113, NR3C1, NR3C1, HNF4A; iDTS-1-3: NR3C1, CREB3L3, NR3C1, PPARA, POU2F1; iDTS-1-4: NR3C1, CEBPA, CEBPA, CEBPA, NR3C1. [Figure 11A] Schematics of various TetR proteins and NLS / NES domains tested in mRNA / DNA co-formulations with NTDNA containing the TetO sequence as the DTS are provided. (A) Schematic of TetR-operated proteins. (B) TetR Set 1.X. (C) TetR Set 2.X. (D) TetR Set 3.X. (E) TetR Set 4.X. (F) TetR Set 5.X. [Figure 11B] Same as above [Figure 11C] Same as above [Figure 11D] Same as above [Figure 11E] Same as above [Figure 11F] Same as above [Figure 12]We present data on the use of mRNA expressing a nuclear targeting factor (NTF) to bind to DNA and facilitate its translocation into the nucleus. Plasmid NT DNA containing a DTS with a TetR-binding site, e.g., a tetracycline-responsive element (TRE) containing seven copies of the TetO sequence, was delivered into growth-arrested HepG2 cells along with mRNA expressing a TetR-NLS fusion protein. Forty-eight hours after transfection, GFP expression was assessed by measuring GFP fluorescence intensity, and toxicity was assessed by counting rounded (dead) cells. (A) Illustrations of mRNAs designed to produce different TetR fusion proteins, each with a DNA payload containing a TetR-binding site and a nuclear localization signal (NLS) and nuclear export signal (NES) located at the N-terminus, C-terminus, or both. (B) GFP intensity and rounded cell counts were determined 48 hours after transfection. [Figure 13A] We provide additional data on the use of mRNA expressing a nuclear targeting factor (NTF) to bind to DNA and facilitate its translocation to the nucleus. Plasmid NT DNA containing a DTS with a TetR-binding site (e.g., TetO sequence) was delivered into growth-arrested HepG2 cells along with mRNA expressing a TetR-NLS fusion protein. GFP expression was assessed by measuring GFP fluorescence intensity over a 21-hour period posttransfection. (A) GFP intensity was measured 21 hours after cotransfection of doggybone DNA (dbDNA) and TetR mRNA. (B) GFP intensity was measured 21 hours after cotransfection of plasmid DNA (pDNA) and TetR mRNA. (C) The fold change in GFP intensity increase was calculated relative to DNA alone (e.g., DNA containing the TetO sequence but without the addition of TetR mRNA) and is presented for both pDNA and dbDNA. (D) Time course of GFP intensity up to 21 hours after cotransfection of plasmid DNA (pDNA) and TetR mRNA. (E) Time course of GFP intensity up to 21 h after cotransfection of doggybone DNA (dbDNA) and TetR mRNA. [Figure 13B] Same as above [Figure 13C] Same as above [Figure 13D] Same as above [Figure 13E] Same as above [Figure 14A]We provide additional data on the use of mRNA expressing a nuclear targeting factor (NTF) to bind to DNA and facilitate its translocation to the nucleus. Plasmid NTDNA containing a DTS with a TetR-binding site (e.g., TetO sequence) was delivered into growth-arrested HepG2 cells along with mRNA expressing a TetR-NLS fusion protein. GFP expression was assessed by measuring GFP fluorescence intensity over a 24-hour period posttransfection. (A) GFP intensity was measured 18 hours after cotransfection of dbDNA containing the TetO sequence and TetR mRNA. (B) GFP intensity was measured 18 hours after cotransfection of dbDNA lacking the TetO sequence and TetR mRNA. (C) Time course of GFP intensity up to 24 hours after cotransfection of 17.5 ng of dbDNA containing the TetO sequence and 35 ng of TetR mRNA, compared with several negative controls. (D) GFP intensity 24 hours after co-transfection of 17.5 ng of dbDNA containing the TetO sequence and 35 ng of TetR mRNA, compared to several negative controls. (E) The fold change in GFP intensity increase was calculated relative to dbDNA alone (e.g., dbDNA containing the TetO sequence but without added TetR mRNA) at 9 hours post-transfection. (F) The fold change in GFP intensity increase was calculated relative to dbDNA alone (e.g., dbDNA containing the TetO sequence but without added TetR mRNA) at 18 hours post-transfection. (G) The fold change in GFP intensity increase was calculated relative to dbDNA alone (e.g., dbDNA containing the TetO sequence but without added TetR mRNA) at 24 hours post-transfection. (H) The fold change in GFP intensity increase was calculated relative to dbDNA alone (e.g., dbDNA containing the TetO sequence but without added TetR mRNA) at 24 hours post-transfection. (I) The fold change in the increase in GFP intensity was calculated relative to pDNA alone (e.g., pDNA containing the TetO sequence but without the addition of TetR mRNA) at 24 h post-transfection.(J) Toxicity was assessed by counting the number of rounded (dead) cells 18 hours after co-transfection of dbDNA + TetR mRNA. (K) Toxicity was assessed by counting the number of rounded (dead) cells 18 hours after co-transfection of pDNA + TetR mRNA. [Figure 14B] Same as above [Figure 14C] Same as above [Figure 14D] Same as above [Figure 14E] Same as above [Figure 14F] Same as above [Figure 14G] Same as above [Figure 14H] Same as above [Figure 14I] Same as above [Figure 14J] Same as above [Figure 14K] Same as above [Figure 15A] This figure shows the binding and translocation of NT DNA to the nucleus by several different species of nuclear targeting factors (NTFs). Plasmid NT DNA containing a DTS with an NTF-binding site was co-delivered with mRNA expressing the NTF protein. Each NTF protein contains a DNA-binding domain (DBD) and at least one NLS. The DBD of each NTF mediates binding to the DTS encoded on the NT DNA, and the NLS facilitates nuclear translocation of the NTF-NT DNA complex. (A) GFP expression from NT DNA was assessed by measuring GFP fluorescence intensity 18 hours after co-transfection of growth-arrested HepG2 cells with the NT DNA and various amounts of NTF mRNA. (B) GFP expression from NT DNA was assessed by measuring GFP fluorescence intensity 18 hours after co-transfection of primary human hepatocytes (PHH) with the NT DNA and various amounts of NTF mRNA. [Figure 15B] Same as above [Figure 16A]This figure shows the changes in gene expression observed for two different NTFs, Gal4 and TetR, as a result of varying the number of NTF binding sites encoded in the DTS. Several NTDNA plasmids were constructed with DTSs containing 1 to 7 copies of the TetO sequence (a binding site for TetR) or 1, 2, 3, 4, 5, 6, 7, or 10 copies of the upstream activating sequence (UAS) (a binding site for Gal4). (A) GFP expression from NTDNA containing TetO binding sites was assessed by measuring GFP fluorescence intensity 24 hours after transfection of growth-arrested HepG2 cells with either the NTDNA plasmid alone or with various amounts of TetR mRNA. (B) GFP expression from NTDNA containing UAS binding sites was assessed by measuring GFP fluorescence intensity 12 hours after transfection of growth-arrested HepG2 cells with either the NTDNA plasmid alone or with various amounts of Gal4 mRNA. [Figure 16B] Same as above [Figure 17] We demonstrate that several strategies can be used to co-administer an NTF-expressing mRNA and an NTDNA payload to achieve greater gene expression from NTDNA than from NTDNA administered alone. LNP or PBS was administered to mice by intravenous bolus, and human EPO levels were measured 3 days after administration. These LNP formulations utilize mRNA expressing TetR and DNA expressing EPO and containing a DTS with 7x TetO repeats. Circles: NTDNA-LNP alone; squares: administration of LNPs containing NTDNA and NTF-encoding mRNA; up-pointing triangles: co-administration of NTDNA-containing LNPs and NTF-encoding mRNA; down-pointing triangles: administration first with LNPs containing NTF-encoding mRNA, followed by administration of NTDNA-containing LNPs. [Figure 18A]We demonstrate that nuclear targeting factors (NTFs) can facilitate the nuclear translocation of NT DNA in vivo. In all cases, mice were dosed by intravenous bolus injection, and the gene product (FIX) from the expression cassette of NT DNA was measured for 28 days: (A) LNPs formulated with mRNA encoding a DNA + / - TetR-NLS fusion protein (v1) containing a 7xTetO DTS sequence and an expression cassette encoding human FIX; (B) LNPs co-formulated with mRNA encoding a DNA + / - TetR-NLS fusion protein (v2) containing a 7xTetO DTS sequence and an expression cassette encoding human FIX; (C) LNPs formulated with mRNA encoding a DNA + / - Gal4-NLS fusion protein containing a 5xUAS DTS sequence and an expression cassette encoding human FIX; (D) LNPs formulated with mRNA encoding a DNA + / - Arc-NLS fusion protein containing a 7xArc DTS sequence and an expression cassette encoding human FIX; (E) LNPs formulated with mRNA encoding a DNA + / - Arc-NLS fusion protein containing a 7xArc DTS sequence and an expression cassette encoding human FIX; (F) LNPs formulated with DNA + / - mRNA encoding the Mnt-NLS fusion protein containing an expression cassette encoding DTS and human (FIX), (G) LNPs formulated with DNA + / - mRNA encoding the Bac434-NLS fusion protein containing an expression cassette encoding 7xBac434 DTS and human FIX. Filled circles: LNPs co-formulated with DNA and RNA; open squares: LNPs formulated with DNA alone; filled triangles: PBS. [Figure 18B] Same as above [Figure 18C] Same as above [Figure 18D] Same as above [Figure 18E] Same as above [Figure 18F] Same as above DETAILED DESCRIPTION OF THE INVENTION
[0006] Incorporation by Reference All publications and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.
[0007] definition The terms "polypeptide," "polypeptide sequence," "peptide," "peptide sequence," "protein," "protein sequence," and "amino acid sequence" are used interchangeably herein to designate a linear series of amino acid residues connected to one another by peptide bonds, which series may include proteins, polypeptides, oligopeptides, peptides, and fragments thereof. Proteins may be composed of naturally occurring amino acids and / or synthetic (e.g., modified or non-naturally occurring) amino acids. Thus, as used herein, "amino acid" or "peptide residue" refers to both naturally occurring and synthetic amino acids. The terms "polypeptide," "peptide," and "protein" include fusion proteins, including, but not limited to, fusion proteins with heterologous amino acid sequences, fusion proteins with heterologous and homologous leader sequences with or without an N-terminal methionine residue, immunologically tagged proteins, and fusion proteins with detectable fusion partners, such as fluorescent proteins, beta-galactosidase, luciferase, etc., as fusion partners. Furthermore, it should be noted that a dash at the beginning or end of an amino acid sequence indicates either a peptide bond to a further sequence of one or more amino acid residues or a covalent attachment to the carboxyl or hydroxyl terminal group. However, the absence of a dash should not be construed to mean that no such peptide bond or covalent bond to the carboxyl or hydroxyl terminal group is present, as it is customary to omit such in representing amino acid sequences.
[0008] The terms "polynucleotide," "polynucleotide sequence," "oligonucleotide," "oligonucleotide sequence," "oligomer," "oligo," "nucleic acid sequence," or "nucleotide sequence," used interchangeably herein, refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, the term includes, but is not limited to, single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers containing purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.
[0009] The terms "derivative" and "variant" refer to any compound, such as, but not limited to, a nucleic acid or protein, having a structure or sequence derived from a compound disclosed herein, and whose structure or sequence is sufficiently similar to that disclosed herein such that it has the same or similar activity and utility, or that, based on such similarity, would be expected by one of skill in the art to exhibit the same or similar activity and utility as the referenced compound, and thereby be referred to interchangeably as "functionally equivalent" or "functional equivalent." Modifications to obtain a "derivative" or "variant" can include, for example, addition, deletion, and / or substitution of one or more of the nucleic acid or amino acid residues.
[0010] In the context of proteins, a functional equivalent or a fragment of a functional equivalent may have one or more conservative amino acid substitutions. The term "conservative amino acid substitution" refers to the substitution of an amino acid for another amino acid that has similar properties to the original amino acid. Groups of conservative amino acids are as follows: [Table 14]
[0011] Conservative substitutions can be introduced at any position of a given peptide or fragment thereof, as preferred. However, it may also be desirable to introduce non-conservative substitutions, particularly, but not limited to, non-conservative substitutions, at any one or more positions. Non-conservative substitutions that lead to the formation of functionally equivalent fragments of the peptide will differ substantially, for example, in polarity, charge, and / or steric bulk, while maintaining the functionality of the derivative or variant fragment.
[0012] "Percentage of sequence identity" is determined by comparing two optimally aligned sequences over a comparison window, where the portion of the polynucleotide or polypeptide sequence within the comparison window may have additions or deletions (i.e., gaps) compared to the reference sequence (which has no additions or deletions) due to optimal alignment of the two sequences. In some cases, the percentage can be calculated by determining the number of positions where the same nucleic acid base or amino acid residue occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity.
[0013] In the context of two or more nucleic acid or polypeptide sequences, the terms "identical" or percent "identity" refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (e.g., 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identity over a specified region, e.g., an entire polypeptide sequence or an individual domain of a polypeptide), when compared and aligned for maximum correspondence over a comparison window or designated region, as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection. Such sequences are then said to be "substantially identical." This definition also refers to the complement of a test sequence.
[0014] The terms "complementary" or "substantially complementary," as used interchangeably herein, mean that a nucleic acid (e.g., DNA or RNA) has a sequence of nucleotides that allows it to noncovalently bind, i.e., form Watson-Crick and / or G / U base pairs, to another nucleic acid in a sequence-specific, antiparallel manner (i.e., the nucleic acid specifically binds to a complementary nucleic acid). As known in the art, standard Watson-Crick base pairings include adenine (A) pairing with thymidine (T), adenine (A) pairing with uracil (U), and guanine (G) pairing with cytosine (C).
[0015] A DNA sequence that "encodes" a particular RNA is a DNA nucleic acid sequence that is transcribed into RNA when placed under the control of appropriate regulatory sequences. A DNA polynucleotide can encode an RNA that is translated into protein (mRNA), or a DNA polynucleotide can encode an RNA that is not translated into protein (e.g., tRNA, rRNA, or guide RNA; also called "non-coding" RNA or "ncRNA"). A protein-coding sequence, or a sequence that encodes a particular protein or polypeptide, is a nucleic acid sequence that is transcribed into mRNA (in the case of DNA) and translated into a polypeptide in vitro or in vivo (in the case of mRNA) when placed under the control of appropriate regulatory sequences.
[0016] As used herein, a "codon" refers to a sequence of three nucleotides that together form a unit of the genetic code within a DNA or RNA molecule. As used herein, the term "codon degeneracy" refers to the property of the genetic code that allows for variation in the nucleotide sequence without affecting the amino acid sequence of an encoded polypeptide. Nucleotides are typically referred to as follows: A (adenine), G (guanine), C (cytosine), T (thymine), U (uracil), N (= A or C or G or T / U), B (= C or G or T / U), D (= A or G or T / U), H (= A or C or T / U), K (= G or T / U), M (= A or C), R (= A or G), S (= C or G), V (= A or C or G), W (= A or T / U), (= C or T / U). Where RNA is described herein, it will be understood by those skilled in the art that during the transcription process, the DNA sequence is transcribed into an mRNA sequence having the same nucleotide sequence, except for thymine, which is coded as uracil in the RNA. Thus, one skilled in the art will be able to readily convert a DNA sequence disclosed herein or known in the art into an mRNA sequence.
[0017] The term "codon-optimized" or "codon optimization" refers to the modification of codons in a gene or coding region of a nucleic acid molecule for transformation into various hosts to reflect the typical codon usage of the host organism without altering the polypeptide encoded by the DNA. Such optimization involves replacing at least one, or more than two, or a significant number of codons with one or more codons more frequently used in the genes of that organism. Codon usage tables are readily available, for example, at the "Codon Usage Database," available at www.kazusa.or.jp / codon / (accessed March 20, 2008). By utilizing knowledge of codon usage or codon preferences in each organism, one skilled in the art can apply frequencies to any given polypeptide sequence to produce a nucleic acid fragment with a codon-optimized coding region that encodes the polypeptide but uses the codons optimal for a given species. Codon-optimized coding regions can be designed by a variety of methods known to those skilled in the art.
[0018] The terms "recombinant" or "engineered," when used with reference to, for example, a cell, nucleic acid, protein, or vector, indicate that the cell, nucleic acid, protein, or vector has been modified by, or is the result of, laboratory methods. Thus, for example, a recombinant or engineered protein includes a protein produced by laboratory methods. A recombinant or engineered protein may contain amino acid residues not found in the native (non-recombinant or wild-type) form of the protein, or may contain amino acid residues that have been modified, e.g., labeled. The terms can include any modification to a peptide, protein, or nucleic acid sequence. Such modifications can include: any chemical modification of a peptide, protein, or nucleic acid sequence, including one or more amino acids, deoxyribonucleotides, or ribonucleotides; addition, deletion, and / or substitution of one or more of the amino acids in a peptide or protein; and addition, deletion, and / or substitution of one or more of the nucleic acids in a nucleic acid sequence.
[0019] The term "genomic DNA" or "genomic sequence" refers to the DNA of the genome of an organism, including, but not limited to, the DNA of the genome of a bacterium, fungus, archea, plant, or animal.
[0020] As used herein, in the context of nucleic acids, "transgene," "exogenous gene," or "exogenous sequence" refers to a nucleic acid sequence or gene that is not present in the genome of a cell, but that has been artificially introduced into the genome, for example, via genome editing.
[0021] As used herein, in the context of nucleic acids, "endogenous gene" or "endogenous sequence" refers to a nucleic acid sequence or gene that is naturally present in the genome of a cell, without being introduced through any artificial means.
[0022] The term "expression cassette" refers to a DNA coding sequence operably linked to a promoter. "Operably linked" refers to a juxtaposition wherein the components so described are in a relationship permitting them to function in their intended manner. For example, a promoter is operably linked to a coding sequence if it affects its transcription or expression. The terms "recombinant expression vector" or "DNA construct" are used interchangeably herein to refer to a vector and a DNA molecule having at least one insert. Recombinant expression vectors are typically generated for the purposes of expressing and / or propagating an insert or for the construction of other recombinant nucleotide sequences. A nucleic acid may or may not be operably linked to a promoter sequence and may or may not be operably linked to a DNA regulatory sequence.
[0023] The term "operably linked" means that the nucleotide sequence of interest is linked to a control sequence in a manner that allows expression of the nucleotide sequence. The term "control sequence" is intended to include, for example, promoters, enhancers, and other expression control elements (e.g., polyadenylation signals). Such control sequences are well known in the art and are described, for example, in Goeddel; Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, Calif. (1990). Control sequences include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific control sequences). Those skilled in the art will understand that the design of an expression vector can depend on factors such as the choice of target cells and the level of expression desired.
[0024] A cell has been "genetically modified" or "transformed" or "transfected" with exogenous DNA, e.g., a recombinant expression vector, when such DNA has been introduced inside the cell. The presence of the exogenous DNA results in a permanent or transient genetic change. The transforming DNA may or may not be integrated (covalently linked) into the genome of the cell. Genetically modified (or transformed or transfected) cells that have a therapeutic activity, e.g., to treat hemophilia A, may be used and referred to as therapeutic cells.
[0025] The term "concentration" as used in the context of a molecule such as a peptide fragment refers to the amount of the molecule present in a given volume of solution, e.g., the number of moles of the molecule.
[0026] The terms "individual," "subject," and "host" are used interchangeably herein and refer to any subject for whom diagnosis, treatment, or therapy is desired. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a patient. In some embodiments, the subject is a human patient. In some embodiments, the subject may have or is suspected of having a disorder or condition associated with a gene of interest (GOI). In some embodiments, the subject is a human who has been diagnosed, at the time of diagnosis or thereafter, as being at risk for a disorder or condition associated with the GOI. In some cases, a diagnosis of being at risk for a disorder or condition associated with the GOI can be determined based on the presence of one or more mutations in the endogenous GOI or in genomic sequences near the GOI in the genome that may affect expression of the GOI.
[0027] The term "treatment," with reference to a disease or condition, means that at least an amelioration of symptoms associated with the condition from which an individual suffers is achieved, where remission is used broadly to refer to at least a reduction in a parameter associated with the condition being treated (e.g., hemophilia A), e.g., the magnitude of the symptoms. Accordingly, treatment also includes situations in which a pathological condition, or at least the symptoms associated therewith, are completely inhibited, e.g., prevented from occurring, or eliminated entirely, such that the host no longer suffers from the condition, or at least the symptoms that characterize the condition. Thus, treatment includes (i) prevention, i.e., preventing clinical symptoms from developing, e.g., reducing the risk of developing clinical symptoms, including preventing disease progression; and (ii) inhibition, i.e., preventing the onset or further development of clinical symptoms, e.g., ameliorating or completely inhibiting active disease.
[0028] As used herein, the terms "effective amount," "pharmaceutically effective amount," or "therapeutically effective amount" refer to a sufficient amount of a composition to provide a desired benefit when administered to a subject with a particular condition. Thus, the term "therapeutically effective amount" refers to an amount of therapeutic cells or a composition comprising therapeutic cells that is sufficient to promote a particular effect when administered to a subject in need of treatment. An effective amount also includes an amount sufficient to prevent or delay the onset of disease symptoms, alter the course of disease symptoms (for example, but not limited to, slow the progression of disease symptoms), or reverse disease symptoms. It will be understood that for any given case, an appropriate "effective amount" can be determined by one of ordinary skill in the art using routine experimentation.
[0029] As used herein, the term "pharmaceutically acceptable excipient" refers to any suitable substance that provides a pharmaceutically acceptable carrier, additive, or diluent for administration of a compound of interest to a subject. A "pharmaceutically acceptable excipient" can encompass substances referred to as pharmaceutically acceptable diluents, pharmaceutically acceptable additives, and pharmaceutically acceptable carriers.
[0030] Methods for nuclear-targeted DNA delivery are provided. Aspects of the method include contacting a target cell with nuclear-targeted deoxyribonucleic acid (NTDNA), which comprises a DNA nuclear-targeting sequence (DTS) and a cargo nucleic acid heterologous to the DTS. Also provided are compositions for use in practicing the methods of the invention.
[0031] Before the present invention is described in more detail, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.
[0032] Where a range of values is provided, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specific excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0033] Certain ranges are presented herein with the term "about" preceding the numerical value. The term "about" is used herein to provide literal support for the exact number preceded by the term, as well as a number that is close to or approximately the number preceded by the term. In determining whether a number is close to or approximately a specifically recited number, the unrecited number that is close to or approximately the number may be a number that, in the context provided, provides a substantial equivalent to the specifically recited number.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative exemplary methods and materials are now described.
[0035] All publications and patents cited herein are incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited, as if each individual publication or patent was specifically and individually indicated to be incorporated by reference herein. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the publication dates provided may be different from the actual publication dates, which may need to be independently confirmed.
[0036] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should be further noted that the claims may be drafted to exclude any optional element. Accordingly, this statement is intended to serve as a predicate to the use of such exclusive terminology as "solely," "only," and the like, or the use of a "negative" limitation in connection with the recitation of claim elements.
[0037] As will be apparent to those skilled in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has distinct components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible.
[0038] Although apparatus and methods have been or will be described in functional descriptions for the sake of grammatical fluidity, it is expressly understood that the claims should not be construed as necessarily limited by the syntax of "means" or "step" limitations unless expressly drafted under 35 U.S.C. 112, but should be given the full scope of meaning and equivalents of the definitions provided by the claims under the doctrine of judicial equivalents, and that if the claims are expressly drafted under 35 U.S.C. 112, they should be given the full statutory equivalents under 35 U.S.C. 112.
[0039] method As summarized above, a method for delivering cargo nucleic acid sequence to the nucleus of cell is provided.The method of the present invention provides the transport of cargo nucleic acid from the cytosol of cell to the nucleus of cell.Therefore, the method of the present invention can be characterized as the method for transporting cargo nucleic acid to the nucleus of cell.Therefore, this method can be considered as a method for transporting nuclear-targeted DNA.
[0040] As discussed above, aspects of embodiments of the present invention include contacting a cell with a nuclear-targeted deoxyribonucleic acid (NTDNA) comprising a DNA nuclear targeting sequence (DTS) and a cargo nucleic acid heterologous to the DTS, to deliver the cargo nucleic acid sequence to the nucleus of the cell. Accordingly, aspects of the present invention include methods of DTS-mediated nuclear delivery of DNA.
[0041] The methods of the present invention typically provide improved or enhanced delivery and / or expression of DNA to the nucleus over DNA delivery methods that rely on delivery of DNA that does not contain a DTS. In some cases, the improvement is a two-fold or greater increase in the amount of DNA in the nucleus compared to the amount of DNA that would be found in the nucleus if the DNA did not contain a DTS, e.g., a two-fold, three-fold, four-fold, five-fold, six-fold, seven-fold, eight-fold, ten-fold, fifteen-fold, twenty-fold, thirty-fold, forty-fold, or fifty-fold increase in the amount of DNA in the nucleus. This may be detected as an increase in expression of a cargo nucleic acid heterologous to the DTS, e.g., a two-fold, three-fold, four-fold, five-fold, six-fold, seven-fold, eight-fold, ten-fold, fifteen-fold, twenty-fold, thirty-fold, forty-fold, fifty-fold, one hundred-fold, or two hundred-fold increase in the RNA or protein encoded by the cargo nucleic acid, over that expressed in cells not contacted with the DNA or in cells contacted with DNA that does not contain a DTS.
[0042] Various aspects of the method will now be considered in more detail.
[0043] Nuclear-targeted deoxyribonucleic acid (NTDNA) An embodiment of the method includes contacting a cell with NTDNA. NTDNA is a double-stranded deoxyribonucleic acid that can vary in length, in some cases ranging from 15 nt to 15,000 nt, e.g., 100 to 10,000 nt, including 100 to 5,000 nt. NTDNA used in the methods of the invention includes both a DNA nuclear targeting sequence (DTS) and a cargo nucleic acid sequence heterologous to the DTS. Thus, NTDNA includes a DTS domain and a cargo nucleic acid domain. Each of these domains will now be described in further detail.
[0044] A DTS refers to a nucleotide sequence that mediates the translocation of a polynucleotide containing it into the nucleus of a cell. Without wishing to be bound by theory, it is believed that DTSs utilize the translocation of nuclear-acting DNA-binding proteins when translocating from the cytoplasm to the nucleus. These nuclear-acting DNA-binding proteins act like nuclear targeting factors for DNA, binding to sequences on the DNA and attracting the DNA into the nucleus as the DNA-binding protein translocates there. Thus, the nuclear-acting DNA-binding proteins utilized are referred to herein as nuclear targeting factors (NTFs), and the DNA sequences to which they bind are referred to herein as nuclear targeting factor binding sites (NTFBSs, or more simply, TFBSs).
[0045] In some embodiments, the DTS comprises a TFBS bound by a nuclear targeting factor that is active in a target cell. For example, the nuclear targeting factor may be constitutively expressed in the target cell. As another example, the nuclear targeting factor may typically be latent in the cytoplasm of the target cell, but becomes active when the cell is contacted by NTDNA, e.g., as part of a cellular response to NTDNA or a formulation containing NTDNA, e.g., as part of an inflammatory response, e.g., activating TLR9, cGAS / STING, AIM2, IFI16, or DDX41 transcription factors, e.g., NF-kB, IRF3, IRF7, and others known in the art. As another example, the nuclear targeting factor may be provided to the cell, e.g., as a protein or as an mRNA encoding the protein.
[0046] In instances where the nuclear-acting DNA-binding protein is endogenous to the target cell, delivery of the DNA to the nucleus of the cell occurs without the need for additional exogenous factors. Thus, in some embodiments, a method for delivering NT DNA to the nucleus of a cell utilizing a DTS that utilizes such a DNA-binding protein consists essentially of contacting the cell with NT DNA comprising a DTS. Table 1 provides examples of TFBSs that can be utilized in DTSs and nuclear-acting DNA-binding proteins that they utilize in hepatocytes to achieve nuclear transport of the associated cargo DNA without the need to provide additional agents. [Table 1-1] [Table 1-2] [Table 1-3]
[0047] In other embodiments, the DTS comprises a TFBS bound by a nuclear targeting factor (NTF) that becomes active in a target cell when the target cell is contacted with an inducer. Alternatively, the DTS comprises a sequence that mediates nuclear translocation of a polynucleotide comprising it when the cell is additionally contacted with an inducer. Such a DTS may be referred to as an "inducible DTS" or "iDTS." An inducer refers to an agent that activates a DNA-binding protein present outside the nucleus (e.g., at the plasma membrane, in the cytosol, etc.) to translocate to the nucleus. In some cases, the inducer directly binds to the DNA-binding protein that is induced to translocate. In other cases, the inducer binds to an upstream protein, which then activates the DNA-binding protein that is induced to translocate. As one non-limiting example, a DTS may comprise a nucleic acid sequence that is recognized and bound by a nuclear targeting factor, e.g., a cytosolic protein that translocates from the cytosol to the nucleus in response to inducer activity. Because the DTS of the NTDNA binds to the nuclear targeting factor, translocation of the nuclear targeting factor into the nucleus brings the NTDNA into the nucleus, thereby mediating delivery of the NTDNA (and its cargo nucleic acid) to the nucleus. The DTS sequence found for use in the NTDNA can be varied as desired. In some embodiments, the iDTS sequence used in the NTDNA comprises the DNA-binding domain of the nuclear targeting factor.
[0048] One non-limiting example of an iDTS is a nucleotide sequence containing the binding domain of a nuclear receptor. Nuclear receptors (NRs) are ligand-activated transcription factors that bind small molecules to regulate gene expression and other cellular processes. This family includes receptors for steroid hormones and derivatives (estrogen, progesterone, glucocorticoids, vitamin D, oxysterols, and bile acids, among others), as well as receptors for retinoic acid, thyroid hormones, and fatty acids and their derivatives. These ligands can diffuse directly through the cell membrane as a result of their lipophilic nature (reviewed in Beato et al., Steroids 1996 Apr;61(4):240-51; Holzer et al., Curr Top Dev Biol. 2017;125:1-38, 2017). The 48 human nuclear receptors share a conserved modular structure consisting of a sequence-specific DNA-binding domain and a ligand-binding domain, in addition to various other protein-protein interaction domains. Upon interaction with a ligand, NRs bind to the regulatory regions of target genes as homodimers or heterodimers, or more rarely as monomers. At the promoter, NRs interact with other activators and repressors to control gene expression (Beato et al., supra; Simons et al., Mol Endocrinol. 2014 Feb;28(2):173-82; Hah and Kraus, Mol Cell Endocrinol. 2014 Jan 25;382(1):652-664). Of interest in the present invention is a class of nuclear receptors that reside in the cytoplasm in the absence of ligand. Ligand binding to these receptors promotes nuclear translocation and translocation of the cytoplasmic DNA to which they bind. (https: / / reactome.org / content / detail / R-HSA-9006931). iDTSs that may be present in the NTDNA of the present invention include, but are not limited to, those bound by nuclear receptors. Another non-limiting example of an iDTS is a nucleotide sequence that contains the binding domain of a transcription factor that resides quiescently in the cytoplasm and is indirectly activated when the cell is contacted with an inducing agent.By "indirectly activated" is meant that they are activated by a protein that binds to an inducing agent. Examples of such transcription factors include NF-kB, CREB, etc.
[0049] Table 2 provides examples of TFBSs that can be utilized in iDTSs, nuclear targeting factors that can be leveraged to achieve nuclear transport of the associated cargo DNA, and inducers that can be used to drive the translocation event. [Table 2-1] [Table 2-2]
[0050] In other embodiments, the DTS comprises a TFBS bound by a nuclear targeting factor provided to the cell as an mRNA encoding the protein. As will be understood by those skilled in the art, any protein that binds to DNA and is transported to the nucleus when delivered to the cytoplasm can be provided to the cell to mediate nuclear import. Thus, for example, any of the naturally occurring proteins listed in Tables 1 and 2 can be exogenously provided. As another example, a protein that is not native to the cell, i.e., heterologous to the cell, can be provided. Such proteins can be naturally occurring or engineered. Examples include any of the proteins listed in Table 3. Exemplary proteins of each class and the DNA sequences to which they bind are well known in the art and include those described in more detail below. [Table 3]
[0051] In some embodiments, the TFBS is a binding sequence for a protein selected from the group consisting of HNF1A, PPARA, HNF4A, CEBPA, NR3C1, ONECUT1, TBP, NFkB, NR1I3, FOXA1, and ELF5. In some embodiments, the HNF1A-binding sequence is v4 or v6. In some embodiments, the PPARA-binding sequence is v5, v6, or v7. In some embodiments, the HNF4A-binding sequence is v2, v4, or v7. In some embodiments, the CEBPA-binding sequence is v6. In some embodiments, the NR3C1-binding sequence is v2, v3, v4, or v5. In some embodiments, the ONECUT-binding sequence is v5. In some embodiments, the TBP-binding sequence is v1. In some embodiments, the NFkB-binding sequence is v2. In some embodiments, the NR1I3-binding sequence is v2, v4, or v6. In some embodiments, the FOXA1-binding sequence is v1. In some embodiments, the ELF5 binding sequence is v1.
[0052] In some embodiments, the DTS is a combination of binding sequences for HNF1A, NR1I3, PPARA, HNF1A, and PPARA. In certain embodiments, the DTS is a combination of binding sequences HNF1A.v4, NR1I3.v6, PPARA.v5, HNF1A.v6, and PPARA.v6. In certain embodiments, the DTS is GGTTAATAATTAACAGATTACTACTGATACCTCTTCTCTGTGGGTGACCAGCGTCCTAAAGATTACTACTGATAAACTAGGTCAAAGGTCAAGATTACTACTGATAAGTATGGTTAATGATCCTACAGAGATTACTACTGATACAAAACTAGGTCAAAGGTCA (SEQ ID NO: 138).
[0053] In some embodiments, the DTS is a combination of binding sequences for FOXA1, HNF4A.v4, HNF1A, HNF4A, and HNF1A. In certain embodiments, the DTS is a combination of binding sequences FOXA1.v1, HNF4A.v4, HNF1A.v4, HNF4A.v2, and HNF1A.v6. In certain embodiments, the DTS is TGTTTACTTTAGATTACTACTGATATCGAGCGCAGGTCAAAGGTCACCTGCAGATTACTACTGATAGGTTAATAATTAACAGATTACTACTGATATCGAGCGCTGGGCAAAGGTCACCTGCAGATTACTACTGATAAGTATGGTTAATGATCTACAG (SEQ ID NO: 139).
[0054] In some embodiments, the DTS comprises a binding sequence for NFkB. In certain embodiments, the binding sequence is NFkB.v2. In certain embodiments, the DTS comprises GGGACTTTCCAGATTACTACTGATAGGGACTTTCCAGATTACTACTGATAGGGACTTTCCAGATTACTACTGATAGGGACTTTCCAGATTACTACTGATAGGGACTTTCC (SEQ ID NO: 140).
[0055] In some embodiments, the DTS is a combination of binding sequences for CREB1, PPARA, ONECUT1, HNF4A, and PPARA. In certain embodiments, the DTS is a combination of binding sequences CREB1.v2, PPARA.v6, ONECUT1.v5, HNF4A.v9, and PPARA.v2. In certain embodiments, the DTS is CTGACGTCAGAGATTACTACTGATACAAAACTAGGTCAAAGGTCAAGATTACTACTGATAGTCTGCTAAGTCAATAATCAGAATAGATTACTACTGATACGCCCCAGCACACATGATCAGAAGATTACTACTGATAAGGTCAAAGGTCA (SEQ ID NO: 141).
[0056] In some embodiments, the TFBS is a binding sequence for a Tet repressor (TetR) protein, e.g., YCTATCANTGATAGA (SEQ ID NO: 142), a TetO sequence such as, e.g., TCCCTATCAGTGATAGAGA (SEQ ID NO: 143) or TCGAGTTTACTCCCTATCAGTGATAGAGAACG (SEQ ID NO: 144). In some embodiments, the DTS comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 copies of the TetR TFBS. In certain embodiments, the DTS comprises 7 copies of the TetR TFBS. In certain embodiments, the DTS comprises 10 or more copies of the TetR TFBS. In some embodiments, the DTS comprises a sequence having 80% or more identity to a sequence listed in Table 4, e.g., 85% or 90% or more identity to a sequence in Table 4, and in some cases, 95% or more identity, e.g., 96%, 97%, 98%, or 99% identity. In some cases, the DTS binding sequence is identical to a sequence in Table 4. In some cases, the DTS consists essentially of a sequence in Table 4. In certain embodiments, the DTS comprises a tetracycline response element (TRE) known in the art. In certain embodiments, the DTS consists essentially of a TRE. [Table 4]
[0057] In some embodiments, the TFBS is a binding sequence for a DNA binding domain of a gene editing system, such as a guide RNA of a Cas nuclease, a zinc finger domain of a zinc finger nuclease, or a TALE DNA binding domain of a TALEN, such as those described herein or known in the art.
[0058] In some embodiments, the TFBS is a binding sequence for a zinc finger-containing protein ("ZF protein"). As will be understood by one of skill in the art, any ZF protein and its cognate ZF binding sequence can be used as a nuclear targeting factor (NTF) and cognate TFBS in the compositions and methods of the present disclosure. In some embodiments, the DTS comprises a ZF-responsive TFBS with 90% or greater sequence identity to AAACTGCAAAAG.
[0059] In some embodiments, a TFBS is a binding sequence for a TAL effector DNA-binding domain-containing protein ("TALE protein"). TALEs are proteins of 32 fixed amino acids and two variable residues, where the two residues are engineered to recognize specific nucleotides (e.g., NN for G, NI for A, HD for C, etc.). As will be understood by one of skill in the art, any TALE protein and its cognate TALE binding sequence can be used in the compositions and methods of the present disclosure, examples of which can be found, for example, in Li et al. 2011 (Modularly assembled designer TAL effector nucleases for targeted gene knockout and gene replacement in eukaryotes. Nucleic Acids Research, Volume 39, Issue 14, pp 6315-6325) and Kim et al. 2013 (A library of TAL effector nucleases spanning the human genome. Nature Biotechnology volume 31, pages 251-258 (2013)). In some embodiments, the TALE TFBS has 90% or greater identity to the sequence and is selected from the group consisting of TTCATTACACCTGCAGCT, ATAAACCCCCTCCAA, and TCGAGTTTACTCCCTATCAGTGATAGAGAACG (SEQ ID NO: 154).
[0060] In some embodiments, the TFBS is a binding sequence for a GAL4 protein. In some such embodiments, the GAL4 TFBS has the sequence CGG-N 11 -CCG, e.g., CGGAGGACTGTCCTCCG (SEQ ID NO: 155). In some embodiments, the DTS comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 copies of a GAL4 TFBS. In certain embodiments, the DTS comprises 5 copies of a GAL4 TFBS. In certain embodiments, the DTS comprises 7 copies of a GAL4 TFBS. In certain embodiments, the DTS comprises 10 or more copies of a GAL4 TFBS. In certain embodiments, the DTS comprises an upstream activation sequence (UAS) for a native GAL4 protein known in the art. In certain embodiments, the DTS consists essentially of a UAS. In some embodiments, the DTS comprises a sequence having 80% or more identity to a sequence listed in Table 5, e.g., 85% or 90% or more identity to a sequence in Table 5, and in some cases, 95% or more identity, e.g., 96%, 97%, 98%, or 99% identity. In some cases, the DTS is identical to a sequence in Table 5. [Table 5]
[0061] In some embodiments, the TFBS is a binding sequence for the Arc protein, where Arc is a bacteriophage regulatory protein. In some such embodiments, the Arc TFBS comprises the sequence RYRVTAGANNNNNTCTABYRY (SEQ ID NO: 172), e.g., ATGATAGAAGCACTCTACTAT. In some embodiments, the Arc TFBS comprises a sequence having 80%, 85%, 90% or more identity to ATGATAGAAGCACTCTACTAT (SEQ ID NO: 173). In some embodiments, the DTS comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 copies of the Arc TFBS. In certain embodiments, the DTS comprises 5 copies of the Arc TFBS. In certain embodiments, the DTS comprises 7 copies of the Arc TFBS. In certain embodiments, the DTS comprises 10 or more copies of the Arc TFBS. In some embodiments, the DTS comprises a sequence that is 80%, 85%, or 90% or more identical to the sequence ATGATAGAAGCACTCTACTATTGAGTCCTAGATGATAGAAGCACTCTACTATTCTTCACAGGATGATAGAAGCACTCTACTATTAGGGTTCCTATGATAGAAGCACTCTACTATACACTAGAGTATGATAGAAGCACTCTACTATGATAGTATCAATGATAGAAGCACTCTACTATAGCAAACGAAATGATAGAAGCACTCTACTAT (SEQ ID NO: 174), in some cases a sequence that has 95% or more identity to this sequence, and in certain cases a sequence that shares 100% identity with this sequence.
[0062] In some embodiments, the TFBS is a binding sequence for the Mnt protein, where Mnt is a bacteriophage regulatory protein. In some such embodiments, the Mnt TFBS comprises the sequence GGNCCACNGTGGNCC, e.g., ATAGGTCCACGGTGGACCATA (SEQ ID NO: 175). In some embodiments, the Mnt TFBS comprises a sequence having 80%, 85%, 90% or more identity to ATAGGTCCACGGTGGACCATA. In some embodiments, the DTS comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 copies of the Mnt TFBS. In certain embodiments, the DTS comprises 5 copies of the Mnt TFBS. In certain embodiments, the DTS comprises 7 copies of the Mnt TFBS. In certain embodiments, the DTS comprises 10 or more copies of the Mnt TFBS. In some embodiments, the DTS comprises a sequence having 80%, 85%, or 90% or more identity to the sequence ATAGGTCCACGGTGGACCATATGAGTCCTAGATAGGTCCACGGTGGACCATATCTTCACAGGATAGGTCCACGGTGGACCATATAGGGTTCCTATAGGTCCACGGTGGACCATAACACTAGAGTATAGGTCCACGGTGGACCATAGATAGTATCAATAGGTCCACGGTGGACCATAAGCAAACGAAATAGGTCCACGGTGGACCATA (SEQ ID NO: 176), in some cases a sequence having 95% or more identity to this sequence, and in certain cases a sequence sharing 100% identity with this sequence.
[0063] In some embodiments, the TFBS is a binding sequence for the purine synthesis repressor (PurR) protein. In some such embodiments, the PurR TFBS comprises a sequence having 80%, 85%, 90% or more identity to ACGCAAACGTTTTCGT, in some cases 95% or more identity to this sequence, and in certain cases a sequence sharing 100% identity with this sequence. In some embodiments, the DTS comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 copies of the PurR TFBS. In certain embodiments, the DTS comprises 5 copies of the PurR TFBS. In certain embodiments, the DTS comprises 7 copies of the PurR TFBS. In certain embodiments, the DTS comprises 10 or more copies of the PurR TFBS. In some embodiments, the DTS comprises a sequence having 80%, 85%, or 90% or more identity to the sequence ACGCAAACGTTTTCGTTGAGTCCTAGACGCAAACGTTTTCGTTCTTCACAGGACGCAAACGTTTTCGTTAGGGTTCCTACGCAAACGTTTTCGTACACTAGAGTACGCAAACGTTTTCGTGATAGTATCAACGCAAACGTTTTCGTAGCAAACGAAACGCAAACGTTTTCGT (SEQ ID NO: 177), in some cases a sequence having 95% or more identity to this sequence, and in certain cases a sequence sharing 100% identity with this sequence.
[0064] In some embodiments, the TFBS is a binding sequence for a Bac434 protein, where Bac434 is a bacteriophage regulatory protein. In some such embodiments, the Bac434 TFBS comprises a sequence having 80%, 85%, 90% or more identity to ACAAGAAAGTTTGT (SEQ ID NO: 178), ACAAGATACATTGT (SEQ ID NO: 179), or ACAAGAAAAACTGT (SEQ ID NO: 180), in some cases 95% or more identity to ACAAGAAAGTTTGT (SEQ ID NO: 181), ACAAGATACATTGT (SEQ ID NO: 182), or ACAAGAAAAACTGT (SEQ ID NO: 183), and in certain cases shares 100% identity with ACAAGAAAGTTTGT (SEQ ID NO: 184), ACAAGATACATTGT (SEQ ID NO: 185), or ACAAGAAAAACTGT (SEQ ID NO: 186). In some embodiments, the DTS comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 copies of a Bac434 TFBS. In certain embodiments, the DTS comprises 5 copies of a Bac434 TFBS. In certain embodiments, the DTS comprises 7 copies of a Bac434 TFBS. In certain embodiments, the DTS comprises 10 or more copies of a Bac434 TFBS. In certain embodiments, the DTS comprises a sequence that is 80%, 85%, or 90% or more identical to a sequence listed in Table 6, e.g., 85% or 90% or more identical to a sequence in Table 6, and in some cases, 95% or more identical, e.g., 96%, 97%, 98%, or 99% identical. In some cases, the DTS is identical to a sequence in Table 6. [Table 6]
[0065] In some embodiments, the TFBS is a binding sequence for the GCN4 protein. In some such embodiments, the GCN4 TFBS comprises a sequence having 80%, 85%, 90% or more identity to TGACTC, in some cases 95% or more identity to TGACTC, and in certain cases 100% identity to TGACTC, e.g., AGTGACTCATT. In some embodiments, the DTS comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 copies of the GCN4 TFBS. In certain embodiments, the DTS comprises 5 copies of the GCN4 TFBS. In certain embodiments, the DTS comprises 7 copies of the GCN4 TFBS. In certain embodiments, the DTS comprises 10 or more copies of the GCN4 TFBS. In some embodiments, the DTS comprises a sequence that is 80%, 85%, or 90% or more identical to the sequence AGTGACTCATTTGAGTCCTAGAGTGACTCATTTCTTCACAGGAGTGACTCATTTAGGGTTCCTAGTGACTCATTACACTAGAGTAGTGACTCATTGATAGTATCAAGTGACTCATTAGCAAACGAAAGTGACTCATT (SEQ ID NO: 194), in some cases a sequence that has 95% or more identity to this sequence, and in certain cases a sequence that shares 100% identity with this sequence.
[0066] In some embodiments, the TFBS is a binding sequence for a lactose inhibitor (Lacl) protein, also referred to herein as a lactose repressor (LacR) protein, e.g., a LacO sequence, e.g., TTGTTATCCGCTCACAA (SEQ ID NO: 195). In some such embodiments, the LacR TFBS comprises a sequence having 80%, 85%, 90% or more identity to TTGTTATCCGCTCACAA (SEQ ID NO: 196). In certain embodiments, the DTS comprises the lactose operon (LacO), as known in the art. In certain embodiments, the DTS consists essentially of LacO. In some embodiments, the DTS comprises 2, 3, 4, 5, 6, 7, 8, 9, or 10 copies of the LacR TFBS. In certain embodiments, the DTS comprises 5 copies of the LacR TFBS. In certain embodiments, the DTS comprises 7 copies of the LacR TFBS. In certain embodiments, the DTS comprises 10 or more copies of a LacR TFBS. In some embodiments, the DTS comprises a sequence having 80%, 85%, or 90% or more identity to the sequence TTGTTATCCGCTCACAATTCCACATGTGGCCACAAATTGTTATCCGCTCACAATTCCACATGTGGCCACAAATTGTTATCCGCTCACAATTCCACATGTGGCCACAAATTGTTATCCGCTCACAATTCCACATGTGGCCACAAATTGTTATCCGCTCACAATTCCACATGTGGCCACAAATTGTTATCCGCTCACAATTCCACATGTGGCCACAAATTGTTATCCGCTCACAACA (SEQ ID NO: 197), in some cases 95% or more identity to this sequence, and in certain cases shares 100% identity with this sequence.
[0067] In some embodiments, the TFBS is a binding sequence for an endonuclease, such as the I-SceI D44A protein. In some embodiments, the DTS comprises a TFBS with 90% or greater identity to the sequence TAGGGATAACAGGGTAAT (SEQ ID NO: 198).
[0068] Those skilled in the art will understand that DNA-binding proteins can tolerate some degree of nucleotide substitution in their binding sequences, and the TFBSs provided herein are merely exemplary of sequences that may be used. In some embodiments, the TFBSs share 80% or more identity with the sequences disclosed herein, e.g., 85%, 90%, 95% or more identity to the sequences, e.g., 96%, 97%, 98%, or 99% or more identity, and in certain cases, 100% identity. Publicly available databases, such as Uniprot, Cis-BP, Transfac, and GrassiusX, can be consulted to identify which nucleotides can be varied and which should be conserved in designing DTSs for use in the disclosed inventions.
[0069] A given DTS may include two or more different TFBSs (i.e., TFBSs that differ from one another by nucleotide sequence and are therefore distinct), e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 different TFBSs. A given DTS may include one or more copies of the same TFBS, e.g., 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more copies of the same TFBS, where in some cases the number of copies of the identical TFBS does not exceed 10.
[0070] The length of a given DTS can vary depending on the number of TFBSs it comprises, the length of the TFBS sequences to which each nuclear targeting factor binds, and the number of nucleotides (spacer sequences) between the TFBSs. When two or more TFBSs (either the same or different) are present in a given DTS, the distance between any two TFBSs can vary as desired, in some cases ranging from 5 to 100 bp, e.g., 10 to 75 bp, including 15 to 50 bp. For example, TFBSs can be separated from each other by 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nt, e.g., 0 to 5 nt, 6 to 10 nt, 11 to 15 nt, 16 to 20 nt, 21 to 25 nt, 26 to 30 nt, 31 to 35 nt, 36 to 40 nt, and in some cases, 40 to 50 nt. In some cases, the length of a given DTS ranges from 10 to 500 nt, such as from 100 to 300 nt, including from 150 to 200 nt.
[0071] A given NTDNA may contain a single DTS or multiple DTSs, as desired. When a given NTDNA contains multiple DTSs, the different DTSs may be the same or different. Thus, a given NTDNA may contain two or more different DTSs (i.e., DTSs that differ from each other by nucleotide sequence and are therefore distinct).
[0072] The DTS can be integrated into the NT DNA at any of a variety of locations. For example, the DTS can be located 5' of the promoter, 3' of the promoter and 5' of the expression cassette, within an intron of the expression cassette, or 3' of the expression cassette. In some cases, the DTS is located 5' of the promoter. In some embodiments, the DTS is located 3' of the expression cassette.
[0073] As discussed above, in addition to the DTS components (e.g., composed of one or more TFBS sequences), the NTDNA used in the methods of the present invention also includes a cargo nucleic acid that is heterologous to the DTS. "Heterologous" to the DTS means that the cargo nucleic acid is not naturally associated with the DTS, e.g., is not part of the same gene as the DTS in nature. The cargo nucleic acid can vary as desired. In some cases, the cargo nucleic acid is 500 nt or longer, e.g., 1 kb, 2 kb, 3 kb, 4 kb, or 5 kb or longer, e.g., 6 kb, 7 kb, 8 kb, 9 kb, 10 kb or longer, and in some cases, 15 kb or longer. The cargo nucleic acid can have any desired sequence. In some cases, the cargo nucleic acid can include one or more of a coding sequence, a promoter, a sequence homologous to the genomic DNA of the target nucleus (e.g., to provide genomic integration of the cargo nucleic acid), an untranslated sequence (5'UTR, 3'UTR), a polyadenylation sequence, etc.
[0074] The cargo nucleic acid delivered to the nucleus can be configured to be episomally maintained or integrated into the genome, as desired. Thus, in some cases, the cargo nucleic acid is configured to be episomally maintained in the nucleus of the target cell so that it is not integrated into the genome. In other cases, the cargo nucleic acid can be configured to be integrated into the genome of the target cell. In such cases, integration can be achieved using any convenient protocol, for example, by using a gene editing system, as described in more detail below.
[0075] In some cases, the cargo nucleic acid comprises a coding sequence. By coding sequence is meant a nucleic acid sequence encoding any gene product, for example, a microRNA (miRNA), a small hairpin RNA (shRNA), a circular RNA (circRNA), a long non-coding RNA (lncRNA), an mRNA, a peptide, a polypeptide, or a protein. If desired, a given coding sequence can encode multiple gene products, separated by, for example, an IRES sequence, a 2A sequence, or the like. For example, the coding sequence delivered to the nucleus by NTDNA can be configured so that it can be operably linked to its native promoter and integrated into the genome, for example, to replace a mutant coding sequence or to be operably linked to an active promoter in a safe harbor (in this case, the cargo nucleic acid does not need to include a promoter, and in some cases does not include a promoter).
[0076] In some cases, the cargo nucleic acid may contain a promoter. As used herein, the term "promoter" refers to any nucleic acid sequence that controls the expression of another nucleic acid sequence by driving the transcription of that nucleic acid sequence, which may be a heterologous target gene encoding a protein or RNA. A promoter may be constitutive, inducible, repressible, tissue-specific, or any combination thereof. A promoter is a control region of a nucleic acid sequence that controls the initiation and rate of transcription of the remainder of the nucleic acid sequence. A promoter may also contain genetic elements to which regulatory proteins and molecules, such as RNA polymerase and other transcription factors, can bind. Within the promoter sequence, a transcription initiation site and protein binding domains involved in the binding of RNA polymerase will be found. Eukaryotic promoters often, but not necessarily, contain "TATA" and "CAT" boxes. A variety of promoters, including inducible promoters, can be used to drive transgene expression. The promoter sequence is bounded at its 3' end by a transcription initiation site and extends upstream (5' direction) to include the minimum number of bases or elements necessary to initiate transcription at a level detectable above background. Promoters useful in the NTDNA of the present invention include, for example, constitutively active promoters, such as the CMV promoter, the CAG promoter (a combination of the chicken β-actin (CBA) promoter and the CMV enhancer), the β-actin promoter, the SV-40 promoter, the 4xGRM6-SV40, hTTR, hAAT, 3x-Serpina, ubiquitin B / C, EF1-alpha, EFS, and the HBV promoter. Promoters useful in the NTDNA of the present invention also include promoters with more cell-type-specific expression patterns, such as, but not limited to, TTR, hAAT (and derivatives), 3xSerpina-TTR, HBV, UbiC, and the P3-hybrid promoter for hepatocytes. The cargo nucleic acid may contain a promoter sequence that is delivered to the nucleus so that it can integrate into the genome, e.g., to replace a mutant promoter.
[0077] In some cases, the cargo nucleic acid comprises an expression cassette. An expression cassette refers to a nucleic acid sequence comprising, for example, a promoter as described above, operably linked to, for example, a coding sequence as described above (also referred to herein as a transgene). In some cases, the expression cassette may also comprise one or more nucleic acid sequences comprising a 5' untranslated region (5'UTR), a 3' untranslated region (3'UTR), a polyA tail, and an miRNA regulatory element. In embodiments, the expression cassette may comprise a transgene and one or more regulatory sequences that enable and / or control expression of the transgene, where, for example, the expression cassette may comprise, in this order, one or more of an enhancer / promoter, an ORF reporter (transgene), a post-transcriptional regulatory element (e.g., WPRE), and a polyadenylation and termination signal (e.g., BGH polyA). The expression cassette may also comprise an internal ribosome entry site (IRES) and / or a 2A element. Cis-regulatory elements include, but are not limited to, promoters, riboswitches, insulators, mir-controllable elements, post-transcriptional regulatory elements, tissue- and cell-type-specific promoters, and enhancers. If desired, the expression cassette can contain 4,000 or more nucleotides, 5,000 or more nucleotides, 10,000 or more nucleotides, 20,000 or more nucleotides, 30,000 or more nucleotides, 40,000 or more nucleotides, or 50,000 or more nucleotides, and in some cases, can range from 4,000 to 10,000 nucleotides, or 10,000 to 50,000 nucleotides, or more than 50,000 nucleotides. An expression cassette delivered to the nucleus can be configured so that it can be maintained episomally. Alternatively, an expression cassette delivered to the nucleus can be configured so that it can be integrated into the genome.
[0078] The coding sequence of the expression cassette, e.g., transgene, can vary. In some embodiments, the expression cassette can comprise a transgene ranging from 500 to 50,000 nucleotides in length. In some embodiments, the expression cassette can comprise a transgene ranging from 500 to 75,000 nucleotides in length. In some embodiments, the expression cassette can comprise a transgene ranging from 500 to 10,000 nucleotides in length. In some embodiments, the expression cassette can comprise a transgene ranging from 1,000 to 10,000 nucleotides in length. In some embodiments, the expression cassette can comprise a transgene ranging from 500 to 5,000 nucleotides in length. The NTDNA constructs of the present embodiments do not have the size limitations of encapsidated AAV vectors, thus enabling nuclear delivery of large expression cassettes for efficient transgene delivery.
[0079] A given expression cassette can include, for example, an expressible exogenous sequence (e.g., an open reading frame) or transgene encoding a protein that is either absent, inactive, or insufficiently active in a recipient subject, or a gene encoding a protein with a desired biological or therapeutic effect. A transgene can encode a gene product that can function to correct expression of a defective gene or transcript. In principle, an expression cassette can include any gene that encodes a protein, polypeptide, or RNA that is either reduced or absent due to a mutation, or that provides a therapeutic benefit if overexpression is considered within the scope of this disclosure. An expression cassette can include any transgene useful for treating a disease or disorder in a subject. For example, the NTDNA described herein can be used to deliver and express any gene of interest in a subject, including, but not limited to, exogenous genes and nucleotide sequences, including polypeptide-encoding or non-coding nucleic acids (e.g., RNAi, miR, etc.), as well as viral sequences within the subject's genome, e.g., HIV viral sequences. In some cases, NTDNA (e.g., as disclosed herein) is used for therapeutic purposes (e.g., for medical, diagnostic, or veterinary use) or immunogenic polypeptides. In certain embodiments, NTDNA is useful for expressing any gene of interest in a subject, including one or more polypeptides, peptides, ribozymes, peptide nucleic acids, siRNAs, RNAi, antisense oligonucleotides, antisense polynucleotides, or RNA (coding or non-coding, e.g., siRNAs, shRNAs, microRNAs, and their antisense counterparts (e.g., antagomirs)), antibodies, antigen-binding fragments, or any combination thereof. Thus, the expression cassette can encode a polypeptide, sense or antisense oligonucleotide, or RNA (coding or non-coding, e.g., siRNAs, shRNAs, microRNAs, and their antisense counterparts (e.g., antagomirs)).Expression cassettes can include exogenous sequences encoding reporter proteins used for experimental or diagnostic purposes, such as β-lactamase, β-galactosidase (LacZ), alkaline phosphatase, thymidine kinase, green fluorescent protein (GFP), chloramphenicol acetyltransferase (CAT), luciferase, EPO, and others known in the art.
[0080] The sequences provided in a given expression cassette, e.g., an expression construct of the NTDNA described herein, can be codon-optimized for a target host cell. As used herein, the term "codon-optimized" or "codon optimization" refers to the process of modifying a nucleic acid sequence by replacing at least one, two or more, or a substantial number of codons of a native sequence (e.g., a prokaryotic sequence) with codons more frequently or most frequently used in the genes of a vertebrate of interest, for enhanced expression in the cells of that vertebrate. Different species exhibit particular biases toward certain codons for particular amino acids. Typically, codon optimization does not alter the amino acid sequence of the original translated protein. In some embodiments, the transgene expressed by the NTDNA is a therapeutic gene. In some embodiments, the therapeutic gene is an antibody or antibody fragment, or an antigen-binding fragment thereof, such as a neutralizing antibody or antibody fragment. In some cases, a therapeutic gene is one or more therapeutic agents, including, but not limited to, proteins, polypeptides, peptides, enzymes, antibodies, antigen-binding fragments, and variants and / or active fragments thereof, for use in the treatment, prevention, and / or amelioration of one or more symptoms of, for example, a disease, dysfunction, injury, and / or disorder. In certain embodiments, the transgene is heterologous compared to the DTS of the NT DNA. For example, a coding sequence delivered to the nucleus, e.g., to replace a mutant coding sequence or to be operably linked to an active promoter in a safe harbor (in this case, the polynucleotide does not contain a promoter), so that it can be operably linked to its native promoter and integrated into the genome; or a coding sequence of an expression cassette delivered to the nucleus, either maintained episomally or integrated into the genome. A coding sequence refers to a nucleic acid sequence encoding any gene product, e.g., a microRNA (miRNA), a small hairpin RNA (shRNA), a circular RNA (circRNA), a long non-coding RNA (lncRNA), an mRNA, a peptide, a polypeptide, or a protein.For example, it may encode multiple gene products separated by IRES sequences, 2A sequences, etc.
[0081] If desired, a given cargo nucleic acid may contain flanking sequences homologous to a region of the cell's genome, for example, to facilitate genomic integration of the cargo nucleic acid. If present, such sequences may, in some cases, vary in length from 30 to 5,000 nt, e.g., 50 to 1,000 nt. While the sequence of such regions may vary depending on the intended location of genomic integration, examples of such sequences include, but are not limited to, actin, ADA, albumin, α-globin, β-globin, CD2, CD3, CD5, CD7, CCR5, E1α, IL2RG, Ins1, Ins2, NCF1, p50, p65, PF4, PGC-γ, PTEN, TERT, TRAC, UBC, and VWF.
[0082] In embodiments in which NT DNA is integrated into the genome of a target cell during a process mediated by a gene editing system, e.g., as described below, the NT DNA can function as donor DNA or a donor template in such systems. Site-specific polypeptides, such as DNA endonucleases, can introduce double-strand or single-strand breaks into nucleic acids, e.g., genomic DNA. The double-strand break can stimulate the cell's endogenous DNA repair pathways (e.g., homology-dependent repair (HDR) or non-homologous end joining or alternative non-homologous end joining (A-NHEJ) or microhomology-mediated end joining (MMEJ). NHEJ can repair the cleaved target nucleic acid without the need for a homologous template. This can sometimes result in small deletions or insertions (indels) in the target nucleic acid at the site of the break, leading to disruption or alteration of gene expression. HDR, also known as homologous recombination (HR), can occur when a homologous repair template, or donor, is available.
[0083] A homologous donor template has a sequence homologous to the sequence adjacent to the target nucleic acid cleavage site. Sister chromatids are generally used by cells as repair templates. However, for genome editing purposes, repair templates are often provided as exogenous nucleic acids, such as plasmids, double-stranded oligonucleotides, single-stranded oligonucleotides, double-stranded oligonucleotides, or viral nucleic acids. The exogenous donor template typically introduces additional nucleic acid sequences (e.g., transgenes) or modifications (e.g., single or multiple base changes or deletions) between the adjacent regions of homology, so that the additional or modified nucleic acid sequence is also integrated into the target locus. MMEJ results in genetic outcomes similar to NHEJ in that small deletions and insertions can occur at the cleavage site. MMEJ utilizes several base pairs of homologous sequences adjacent to the cleavage site to promote favorable end-joining DNA repair outcomes. In some cases, it may be possible to predict likely repair outcomes based on analysis of potential microhomologies in the nuclease target region.
[0084] Thus, in some cases, homologous recombination is used to insert an exogenous polynucleotide sequence into the target nucleic acid cleavage site. The exogenous polynucleotide sequence is referred to herein as a donor polynucleotide (or donor or donor sequence or polynucleotide donor template), and in embodiments of the invention, may be NTDNA or a component thereof, such as a cargo nucleic acid component of NTDNA. In some embodiments, a donor polynucleotide, a portion of a donor polynucleotide, a copy of a donor polynucleotide, or a portion of a copy of a donor polynucleotide is inserted into the target nucleic acid cleavage site. In some embodiments, the donor polynucleotide is an exogenous polynucleotide sequence, i.e., a sequence that does not naturally occur at the target nucleic acid cleavage site.
[0085] When exogenous DNA molecules are provided in sufficient concentration inside the nucleus of a cell where a double-strand break occurs, the exogenous DNA can be inserted into the double-strand break during the NHEJ repair process, thus becoming a permanent addition to the genome. These exogenous DNA molecules are referred to as donor templates in some embodiments. When the donor template contains the coding sequence for a gene of interest, optionally together with related regulatory sequences such as a promoter, an enhancer, a polyA sequence, and / or a splice acceptor sequence (also referred to herein as a "donor cassette"), the gene of interest can be expressed from the integrated copy in the genome, resulting in permanent expression throughout the life of the cell. Furthermore, the integrated copy of the donor DNA template can be transmitted to daughter cells when the cell divides.
[0086] In the presence of sufficient concentration of donor DNA template that contains adjacent DNA sequences (referred to as homology arms) that have homology to the DNA sequences on either side of double-strand break, donor DNA template can be integrated through HDR pathway.Homology arms act as the substrate for homologous recombination between donor template and the sequences on either side of double-strand break.This can result in the error-free insertion of donor template, and the sequences on either side of double-strand break are not modified from the sequences in unmodified genome.
[0087] Donors supplied for HDR editing vary significantly, but generally contain the intended sequence with small or large flanking homology arms to allow annealing to genomic DNA. The homology regions flanking the introduced genetic change can be 30 bp or less, or as large as multi-kilobase cassettes that can contain promoters, cDNA, and the like. Both single-stranded and double-stranded oligonucleotide donors can be used. These oligonucleotides range in size from less than 100 nt to over several kb, although longer single-stranded DNA can also be generated and used. Double-stranded donors, including PCR amplicons, plasmids, and minicircles, are often used.
[0088] In some embodiments, the exogenous sequence, e.g., cargo nucleic acid, intended to be inserted into the genome is a gene of interest (GOI) or a functional derivative thereof. The exogenous gene can include a nucleotide sequence encoding a GOI product, e.g., a GOI protein, or a functional derivative thereof. A functional derivative of a GOI can include a nucleic acid sequence encoding a functional derivative of a GOI protein having substantial activity of a wild-type GOI protein, such as a wild-type human GOI protein, e.g., at least about 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the activity exhibited by the wild-type GOI protein. In some embodiments, a functional derivative of a GOI protein can have at least about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, or about 99% amino acid sequence identity to the GOI protein, e.g., the wild-type GOI protein. In some embodiments, one skilled in the art can test the functionality or activity of a compound, e.g., a peptide or protein, using several methods known in the art. A functional derivative of a GOI protein can also include any fragment of the wild-type GOI protein, or a fragment of a modified GOI protein having conservative modifications to one or more of the amino acid residues in the full-length wild-type GOI protein. Thus, in some embodiments, a functional derivative of a nucleic acid sequence of a GOI can have at least about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, or about 99% nucleic acid sequence identity to the GOI, e.g., a wild-type GOI.
[0089] In some embodiments involving the insertion of a GOI or a functional derivative thereof, a cDNA of the GOI or a functional derivative thereof can be inserted into the genome of a patient having a missing GOI or its regulatory sequences. In such cases, the donor DNA or donor template can be an expression cassette or vector construct having a sequence, e.g., a cDNA sequence, encoding the GOI or a functional derivative thereof. In some embodiments, an expression vector can contain and be used a sequence encoding a modified GOI protein as described elsewhere in this disclosure.
[0090] In some embodiments, according to any of the donor templates described herein that include a donor cassette, the donor cassette is flanked on one or both sides by gRNA target sites. For example, such a donor template can include a donor cassette having a 5' gRNA target site and / or a 3' gRNA target site of the donor cassette. In some embodiments, the donor template includes a donor cassette having a 5' gRNA target site of the donor cassette. In some embodiments, the donor template includes a donor cassette having a 3' gRNA target site of the donor cassette. In some embodiments, the donor template includes a donor cassette having a 5' gRNA target site of the donor cassette and a 3' gRNA target site of the donor cassette. In some embodiments, the donor template includes a donor cassette having a 5' gRNA target site of the donor cassette and a 3' gRNA target site of the donor cassette, wherein the two gRNA target sites comprise the same sequence. In some embodiments, the donor template comprises at least one gRNA target site, wherein the at least one gRNA target site in the donor template comprises the same sequence as a gRNA target site in the target locus into which the donor cassette of the donor template will be integrated. In some embodiments, the donor template comprises at least one gRNA target site, wherein the at least one gRNA target site in the donor template comprises the reverse complement of a gRNA target site in the target locus into which the donor cassette of the donor template will be integrated. In some embodiments, the donor template comprises at least one gRNA target site, wherein the at least one gRNA target site in the donor template does not comprise the same sequence as a gRNA target site in the target locus into which the donor cassette of the donor template will be integrated. In some embodiments, the donor template comprises a donor cassette having a gRNA target site 5' of the donor cassette and a gRNA target site 3' of the donor cassette, wherein the two gRNA target sites in the donor template comprise the same sequences as a gRNA target site in the target locus into which the donor cassette of the donor template will be integrated.In some embodiments, the donor template comprises a donor cassette having a gRNA target site 5' of the donor cassette and a gRNA target site 3' of the donor cassette, wherein the two gRNA target sites in the donor template comprise reverse complements of gRNA target sites at a target locus into which the donor cassette of the donor template will be integrated. In some embodiments, the donor template comprises a donor cassette having a gRNA target site 5' of the donor cassette and a gRNA target site 3' of the donor cassette, wherein the two gRNA target sites in the donor template do not comprise the same sequence as the gRNA target sites at a target locus into which the donor cassette of the donor template will be integrated.
[0091] A given NTDNA may or may not be configured to be maintained in bacteria, as desired. For example, in some cases, the NTDNA is configured to be maintained in bacteria. In such cases, the NTDNA may include a bacterial origin of replication and selection elements, such as a plasmid or nanoplasmid. In other cases, the NTDNA may be configured not to be maintained in bacteria. In such cases, the structure of the NTDNA may vary, and examples of such structures include minicircle DNA (mcDNA), linear DNA, covalently closed DNA (doggybone DNA, or "dbDNA," ministring DNA, etc.), double-stranded linear DNA containing terminal repeats at both ends, 3DNA structures, etc.
[0092] inducer As summarized above, in some embodiments, in addition to the NTDNA, cells can also be contacted with an inducer that activates a nuclear targeting factor to which the DTS binds, mediating nuclear entry of the NTDNA. In other words, the inducer modulates the nuclear targeting factor, causing it to translocate from the cytosol to the nucleus, and in so doing, bringing the associated NTDNA to the nucleus (via the binding interaction between the DTS and the nuclear targeting factor). Thus, the activity of the inducer on the nuclear targeting factor causes the nuclear targeting factor to mediate nuclear entry of the NTDNA (and thus its cargo nucleic acid). Inducers can vary depending on the nuclear targeting factor in a given system. Examples of inducers include those that act directly on the nuclear targeting factor, causing its translocation from the cytosol to the nucleus. Examples of such inducers include, but are not limited to, steroid hormones and derivatives (such as estrogen, progesterone, glucocorticoids, vitamin D, oxysterols, and bile acids, among others), as well as receptors for retinoic acid, thyroid hormones, and fatty acids and their derivatives. These ligands, due to their lipophilic nature, can diffuse directly through the cell membrane (reviewed in Beato et al., Steroids 1996 Apr;61(4):240-51; Holzer et al., Curr Top Dev Biol. 2017;125:1-38. 2017). Also of interest as inducers are agents that indirectly activate nuclear-targeted factors. For example, an inducer may indirectly activate a transcription factor that resides quiescently in the cytoplasm and is indirectly activated when the cell contacts the inducer. "Indirectly activated" means that the nuclear-targeted factor, e.g., transcription factor, is activated by a protein that binds to the inducer. Examples of such transcription factors include NF-kB, CREB, etc. Examples of inducible nuclear-targeted factors, the TFBSs to which they bind, and the inducing ligands that activate them can be found in Table 2 above.
[0093] Within embodiments of the present invention, the specific combination of nuclear targeting factor, DTS, and inducer may vary as desired. Examples of combinations of interest that may be used in a given embodiment include, but are not limited to, those embodiments provided in Table 2 above.
[0094] Exogenous nuclear targeting factor As summarized above, in some embodiments, in addition to NTF DNA, cells can also be contacted with mRNA encoding a nuclear targeting factor to which the DTS binds, mediating nuclear entry of the NTF DNA. Without wishing to be bound by theory, it is believed that the exogenously provided nuclear targeting factor is translated from mRNA, translocates from the cytosol to the nucleus, and, in doing so, brings the associated NTF DNA to the nucleus (via a binding interaction between the DTS and the nuclear targeting factor). Typically, the nuclear targeting factor comprises a nuclear localization sequence (NLS) or a fragment thereof. In some cases, the NLS is native or endogenous to the nuclear targeting factor. In other cases, a DNA-binding protein is engineered to contain an NLS to create a nuclear targeting factor. When engineered into an NTF, the NLS can be engineered to be located anywhere within the NTF. In some cases, the NLS is engineered to be at the N-terminus. In some cases, the NLS is engineered to be at the C-terminus. In some cases, the NLS is engineered to be at both the N-terminus and C-terminus. When engineered to be at the terminus, the NLS is typically engineered within 0-20 amino acids of the terminus, in some cases within 0-10 amino acids of the terminus, in certain cases within 0-5 amino acids of the terminus, and in some such cases, at the terminus. Typically, the NLS is engineered to be at a site different from the protein domain that mediates DNA binding. In some cases, the NLS may be adjacent to one or more amino acids, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids. In some cases, the nuclear targeting factor contains one NLS, and in other cases, multiple NLSs. In some cases where multiple NLSs are used, the same NLS is used multiple times. In other cases where multiple NLSs are used, different NLSs are used. Often, when multiple NLSs are used, they are separated by a linker, e.g., a 6xK linker. Exemplary NLSs and their origins are provided in Table 7. [Table 7]
[0095] In some embodiments, the nuclear targeting factor comprises a nuclear export signal. In some embodiments, the NES is an NMD3 ribosomal export adaptor. In some such embodiments, the NES comprises a sequence having 85% or greater identity to NELALKLAGLDINKT (SEQ ID NO: 217). In some such embodiments, the NES comprises a sequence having 85% or greater identity to EHVNKMNSDRVPDVVLIKKSYDRTKRQRRRNWKLKELA (SEQ ID NO: 218).
[0096] In some embodiments, the nuclear targeting element comprises a tag, linker, or other sequence that can be used, for example, to preserve the structure of the added NLS or NES, to purify the protein during recombinant protein production, to visualize the protein within cells, etc. Examples of tags or linkers include, but are not limited to, CS3, HIS, Flag, GGS, RIGID, SpyTag, Strep tag, etc. Exemplary linkers are provided in Table 8. [Table 8]
[0097] Any method for making mRNA can be used to generate the subject mRNA encoding an NTF. As one example, mRNA can be chemically synthesized using solid-phase methods. As another example, mRNA can be synthesized from a DNA template by an in vitro transcription (IVT) reaction, in which a DNA template containing a promoter sequence for RNA polymerase operably linked to a sequence encoding a target mRNA (in this case, a nuclear targeting factor) is contacted with RNA polymerase, and the RNA polymerase binds to the template at the promoter region and initiates RNA synthesis. As yet another example, mRNA can be synthesized in vivo and purified from a biological source. mRNA can contain naturally occurring ribonucleotides and / or chemically modified ribonucleotides. Typically, some chemically modified nucleotides, such as N,N- ... 1-methylpseudouridine, 2-thiouridine (s 2 U), 5-methylcytidine (m 5 C), N 6 -methyladenosine (m 6 A), 2'-O-methyluridine (Um), 2'-O-methylcytidine (Cm), 2'-O-methyladenosine (Am), and 2'-O-methylguanosine (Gm), N 4 -acetyl-cytidine 5'-triphosphate (AC4C) may be included to reduce the immunogenicity of the mRNA. The mRNA may also contain a 5' cap or analog thereof, e.g., m 7 The mRNA may include a GpppG cap, an anti-reverse cap analog (ARCA), a biceps cap, an S cap, a 2S cap, etc. The mRNA typically includes a 5' untranslated region (UTR). The mRNA typically includes a 3' UTR, for example, as described in Table 9. The mRNA may include a tail modification, for example, a ribose-modified adenosine, 8-azaadenosine, hirudycepin, etc. [Table 9-1] [Table 9-2]
[0098] For example, if the DTS contains a binding sequence for the Tet repressor (TetR) protein, e.g., TCCCTATCAGTGATAGAGA (SEQ ID NO: 267) or a variant thereof, the cell can also be contacted with mRNA encoding the TetR protein. In some embodiments, the TetR protein is engineered to contain an NLS. In some embodiments, the NLS is an NLS listed in Table 7. In certain embodiments, the NLS is an NLS from SV40 large T antigen or an optimized variant thereof, an NLS from influenza A nucleoprotein INF-A, an NLS from c-myc, or a fusion of a c-myc NLS to an influenza A nucleoprotein INF-A NLS. In some embodiments, the NLS is located proximal to the N-terminus. In other embodiments, the NLS is located proximal to the C-terminus. In some embodiments, the NLS is fused to the TetR protein with a linker.
[0099] In some embodiments, the TetR protein has the wild-type Escherichia coli transcription regulator TetR protein sequence: The present invention includes sequences having 90% or more, and in some cases 92%, 93%, 94%, 95% or more sequence identity, and in certain cases 96%, 97%, 98%, 99% or more identity, e.g., 100% identity, to MSRLDKSKVINSALELLNEVGIEGLTTRKLAQKLGVEQPTLYWHVKNKRALLDALAIEMLDRHHTHFCPLEGESWQDFLRNNAKSFRCALLSHRDGAKVHLGTRPTEKQYETLENQLAFLCQQGFSLENALYALSAVGHFTLGCVLEDQEHQVAKEERETPTTDSMPPLLRQAIELFDHQGAEPAFLFGLELIICGLEKQLKCESGS (SEQ ID NO: 268).
[0100] In some embodiments, the TetR protein is encoded by a polynucleotide comprising a sequence having 80% or greater sequence identity to a sequence listed in Table 10, e.g., 85%, 90%, or 95% or greater sequence identity, and in some cases 96%, 97%, 98%, or 99% sequence identity, and in some cases 100% sequence identity over the length of the sequence in Table 10. In some embodiments, the polynucleotide sequence is codon-optimized for expression in a particular host cell, e.g., codon-optimized for expression in a human or mouse cell. [Table 10-1] [Table 10-2] [Table 10-3] [Table 10-4]
[0101] As another example, if the DTS contains a binding sequence for a ZF protein, the cell can also be contacted with an mRNA encoding the ZF protein. For example, if the DTS contains the sequence AAACTGCAAAAG or a variant thereof, the cell can also be contacted with an mRNA encoding the ZF protein ZF-CCR5. In some embodiments, the ZF protein is engineered to contain an NLS. In some embodiments, the NLS is proximal to the N-terminus. In other embodiments, the NLS is proximal to the C-terminus. In some embodiments, the ZF protein contains an NLS at both the N-terminus and the C-terminus. In some embodiments, the NLS is fused to the ZF protein with a linker.
[0102] In some embodiments, the ZF protein comprises a sequence having 90% or more, in some cases 92%, 93%, 94%, 95% or more sequence identity, and in certain cases 96%, 97%, 98%, 99% or more identity to the ZF-CCR5 fusion protein: MRPFQCRICMRNFSRSDNLSVHIRTHTGEKPFACDICGRKFAQKINLQVHTKIHTGEKPFQCRICMRNFSRSDVLSEHIRTHTGEKPFACDICGRKFAQRNHRTTHTKIHLR (SEQ ID NO: 284), e.g., the variant MAPKKKRKVGIHGVPAAMAERPFQCRICMRNFSRSDNLSVHIRTHTGEKPFACDICGRKFAQKINLQVHTKIHTGEKPFQCRICMRNFSRSDVLSEHIRTHTGEKPFACDICGRKFAQRNHRTTHTKIHLR (SEQ ID NO: 285), e.g., 100% identity.
[0103] As another example, if the DTS comprises a binding sequence for a TALE protein, e.g., TTCATTACACCTGCAGCT (SEQ ID NO: 286), ATAAACCCCCTCCAA (SEQ ID NO: 287), or TCGAGTTTACTCCCTATCAGTGATAGAGAACG (SEQ ID NO: 288), or variants thereof, the cell can also be contacted with an mRNA encoding the TALE protein. In some embodiments, the TALE protein is engineered to contain an NLS. In some embodiments, the NLS is proximal to the N-terminus. In other embodiments, the NLS is proximal to the C-terminus. In some embodiments, the TALE protein comprises an NLS at both the N-terminus and the C-terminus. In some embodiments, the NLS is fused to the TALE protein with a linker.
[0104] In some embodiments, the TALE protein is TMVLAQNRKKSLHCFEGLFTAVVTSNSDHLVRSQ (SEQ ID NO: 289), For example, MLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIG GKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHG LTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNGGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASHDGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHGLTPDQVVAIASNIGGKQALETVQRLLPVLCQDHG (SEQ ID NO: 290) 90% or more, in some cases 92%, 93%, 94%, 95% or more sequence identity, and in certain cases 96%, 97%, 98%, 99% or more identity, for example, 100% identical. Other exemplary TALEs and methods for manipulating them can be found in Li et al. 2011 (supra) and Kim et al. 2011 (supra).
[0105] In another example, the DTS may contain a binding sequence for the GAL4 protein, e.g., CGG-N 11If the cell contains -CCG or a variant thereof, the cell may also be contacted with mRNA encoding the GAL4 protein. In some embodiments, the GAL4 protein is engineered to contain an NLS. In some embodiments, the NLS is proximal to the N-terminus. In other embodiments, the NLS is proximal to the C-terminus. In some embodiments, the GAL4 protein is engineered to contain NLSs at both the N- and C-termini. In some embodiments, the NLS is fused to the GAL4 protein with a linker. In some embodiments, the NLS is selected from those listed in Table 7. In some specific embodiments, the NLS is the NLS of SV40 large T antigen or an optimized variant thereof, the NLS of influenza A nucleoprotein INF-A, the NLS of c-myc, or a fusion of the c-myc NLS to influenza A nucleoprotein INF-A NLS.
[0106] In some embodiments, the GAL4 protein comprises a sequence that has 90% or more, in some cases 92%, 93%, 94%, 95% or more sequence identity, and in certain cases 96%, 97%, 98%, 99% or more identity, e.g., 100% identity, to the wild-type Saccharomyces cerevisiae GAL4 protein MKLLSSIEQACDICRLKKLKCSKEKPKCAKCLKNNWECRYSPKTKRSPLTRAHLTEVESRLERLE (SEQ ID NO: 291). In some cases, the GAL4 protein comprises a sequence having 90% or more, in some cases 92%, 93%, 94%, 95% or more sequence identity, and in certain cases 96%, 97%, 98%, 99% or more identity, e.g., 100% identity, to MKLLSSIEQACDICRLKKLKCSKEKPKCAKCLKNNWECRYSPKTKRSPLTRAHLTEVESRLERLEQLFLLIFPREDLDMILKMDSLQDIKALLTGLFVQDNVNKDAVTDRLASVETDMPLTLRQHRISATSSSEESSNKGQRQLTVS (SEQ ID NO: 292).
[0107] In some embodiments, the GAL4 protein is encoded by a polynucleotide and comprises a sequence having 80% or greater sequence identity to a sequence listed in Table 11, e.g., 85%, 90%, or 95% or greater sequence identity, and in some cases 96%, 97%, 98%, or 99% sequence identity, and in some cases 100% sequence identity over the length of the sequence in Table 11. In some embodiments, the GAL4 polynucleotide is codon-optimized for expression in a particular host cell, e.g., codon-optimized for expression in human or mouse cells. [Table 11]
[0108] As another example, if the DTS contains a binding sequence for the Arc protein, e.g., RYRVTAGANNNNNTCTABYRY (SEQ ID NO: 298) or a variant thereof, the cell can also be contacted with mRNA encoding the Arc protein. In some embodiments, the Arc protein is engineered to contain an NLS. In some embodiments, the NLS is proximal to the N-terminus. In other embodiments, the NLS is proximal to the C-terminus. In some embodiments, the Arc protein is engineered to contain NLSs at both the N- and C-termini. In some embodiments, the NLS is fused to the Arc protein with a linker. In some embodiments, the NLS is selected from those listed in Table 7. In some specific embodiments, the NLS is the NLS of SV40 large T antigen or an optimized variant thereof, the NLS of influenza A nucleoprotein INF-A, the NLS of c-myc, or a fusion of the c-myc NLS to influenza A nucleoprotein INF-A NLS.
[0109] In some embodiments, the Arc protein comprises the sequence encoding the Salmonella phage Arc-like repressor: The Arc protein may be the st11 variant MKGMSKMPQFNLRWPREVLDLVRKVAEENGRSVNSEIYQRVMESFKKEGRIGA (SEQ ID NO: 299), which has 90% or more, and in some cases 92%, 93%, 94%, 95% or more, and in certain cases 96%, 97%, 98%, 99% or more, e.g., 100% identity. For example, the Arc protein may be the st11 variant MKGMSKMPQFNLRWPREVLDLVRKVAEENGRSVNSEIYQRVMASFAKEGRIAAKNQH E (SEQ ID NO: 300), or a protein having 90% or more, and in some cases 92%, 93%, 94%, 95% or more, and in certain cases 96%, 97%, 98%, 99% or more, identity to the st11 variant.
[0110] In some embodiments, the Arc protein is encoded by a polynucleotide comprising a sequence having 80% or more sequence identity, e.g., 85%, 90%, or 95% or more sequence identity, and in some cases 96%, 97%, 98%, or 99% sequence identity, or in some cases 100% sequence identity, to the Salmonella phage Arc-like suppressor wild-type mRNA sequence: AUGAAGGGCAUGAGCAAGAUGCCCCAGUUCAACCUGCGCUGGCCCCGCGAGGUGCUGGACCUGGUGCGCAAGGUGGCCGAGGAGAACGGCCGCAGCGUGAACAGCGAGAUCUACCAGCGCGUGAUGGAGAGCUUCAAGAAGGAGGGCCGCAUCGGCGCC (SEQ ID NO: 301). In some embodiments, the Arc protein is encoded by a polynucleotide and includes a sequence having 80% or more sequence identity, e.g., 85%, 90%, or 95% or more sequence identity, and in some cases 96%, 97%, 98%, or 99% sequence identity, or in some cases 100% sequence identity, to the sequence encoding the st11 Arc variant: AUGAAGGGCAUGAGCAAGAUGCCCCAGUUCAACCUGCGCUGGCCCCGCGAGGUGCUGGACCUGGUGCGCAAGGUGGCCGAGGAGAACGGCCGCAGCGUGAACAGCGAGAUCUACCAGCGCGUGAUGGCCAGCUUCGCCAAGGAGGGCCGCAUCGCCGCCAAGAACCAGCACGAG (SEQ ID NO: 302). In some cases, the polynucleotide is codon-optimized for expression in a particular host cell, for example, codon-optimized for expression in a human cell.
[0111] As another example, if the DTS contains a binding sequence for the Mnt protein, e.g., GGNCCACNGTGGNCC (SEQ ID NO: 303) or a variant thereof, e.g., ATAGGTCCACGGTGGACCATA (SEQ ID NO: 304), the cell can also be contacted with mRNA encoding the Mnt protein. In some embodiments, the Mnt protein is engineered to contain an NLS. In some embodiments, the NLS is proximal to the N-terminus. In other embodiments, the NLS is proximal to the C-terminus. In some embodiments, the Mnt protein is engineered to contain NLSs at both the N- and C-termini. In some embodiments, the NLS is fused to the Mnt protein with a linker. In some embodiments, the NLS is selected from those listed in Table 7. In some specific embodiments, the NLS is the NLS of SV40 large T antigen or an optimized variant thereof, the NLS of influenza A nucleoprotein INF-A, the NLS of c-myc, or a fusion of the c-myc NLS to influenza A nucleoprotein INF-A NLS.
[0112] In some embodiments, the Mnt protein comprises a sequence that has 90% or more, in some cases 92%, 93%, 94%, 95% or more sequence identity, and in certain cases 96%, 97%, 98%, 99% or more identity, e.g., 100% identity, to the naturally occurring wild-type Mnt: MARDDPHFNFRMPMEVREKLKFRAEANGRSMNSELLQIVQDALSKPSPVTGYRNDAERLADEQSELVKKMVFDTLKDLYKKTT (SEQ ID NO: 305).
[0113] In some embodiments, the Mnt protein is encoded by a polynucleotide comprising a sequence having 80% or more sequence identity, e.g., 85%, 90%, or 95% or more sequence identity, and in some cases 96%, 97%, 98%, or 99% sequence identity, and in some cases 100% sequence identity, over the sequence: AUGGCCCGCGACGACCCCCACUUCAACUUCCGCAUGCCCAUGGAGGUGCGCGAGAAGCUGAAGUUCCGCGCCGAGGCCAACGGCCGCAGCAUGAACAGCGAGCUGCUGCAGAUCGUGCAGGACGCCCUGAGCAAGCCCAGCCCCGUGACCGGCUACCGCAACGACGCCGAGCGCCUGGCCGACGAGCAGAGCGAGCUGGUGAAGAAGAUGGUGUGUUCGACACCCUGAAGGACCUGUACAAGAAGACCACC (SEQ ID NO: 306). In some embodiments, the Mnt polynucleotides are codon-optimized for expression in a particular host cell, for example, codon-optimized for expression in human or mouse cells.
[0114] As another example, if the DTS contains a binding sequence for the PurR protein, e.g., ACGCAAACGTTTTCGT (SEQ ID NO: 307) or a variant thereof, the cell can also be contacted with an mRNA encoding the PurR protein. In some embodiments, the PurR protein is engineered to contain an NLS. In some embodiments, the NLS is proximal to the N-terminus. In other embodiments, the NLS is proximal to the C-terminus. In some embodiments, the PurR protein is engineered to contain an NLS at both the N-terminus and the C-terminus. In some embodiments, the NLS is fused to the PurR protein with a linker. In some embodiments, the NLS is selected from those listed in Table 7.
[0115] In some embodiments, the PurR protein is the wild-type Escherichia coli HTH-type transcriptional repressor PurR:MATIKDVAKRANVSTTTVSHVINKTRFVAEETRNAVWAAIKELHYSPSAVARSLKVNHTKSIGLLATSSEAAYFAEIIEAVEKNCFQKGYTLILGNAWNNLEKQRAYLSMMAQKRVDGLLVMCSEYPEPLLAMLEEYRHIPMVVMDWGEAKADFTDAVIDNAFEGGYMAGRYLIERGHREIGVIPGPLERNTGAGRLAGFMKAMEEAMIKVPESWIV The present invention includes sequences having 90% or more, in some cases 92%, 93%, 94%, 95% or more sequence identity, and in certain cases 96%, 97%, 98%, 99% or more identity, e.g., 100% identity, to QGDFEPESGYRAMQQILSQPHRPTAVFCGGDIMAMGALCAADEMGLRVPQDVSLIGYDNVRNARYFTPALTTIHQPKDSLGETAFNMLLDRIVNKREEPQSIEVHPRLIERRSVADGPFRDYRR (SEQ ID NO: 308).
[0116]
[0117] As another example, if the DTS contains a binding sequence for the Bac434 protein, e.g., ACAAGAAAGTTTGT (SEQ ID NO: 310), ACAAGATACATTGT (SEQ ID NO: 311), or ACAAGAAAAACTGT (SEQ ID NO: 312), or a variant thereof, the cell can also be contacted with mRNA encoding the Bac434 protein. In some embodiments, the Bac434 protein has been engineered to contain an NLS. In some embodiments, the NLS is proximal to the N-terminus. In other embodiments, the NLS is proximal to the C-terminus. In some embodiments, the Bac434 protein is engineered to contain an NLS at both the N-terminus and the C-terminus. In some embodiments, the NLS is fused to the Bac434 protein with a linker. In some embodiments, the NLS is selected from those listed in Table 7.
[0118] In some embodiments, the Bac434 protein comprises a sequence having 90% or more, in some cases 92%, 93%, 94%, 95% or more sequence identity, and in certain cases 96%, 97%, 98%, 99% or more identity, e.g., 100% identity, to the wild-type Escherichia coli phage 434 suppressor protein CI: MSISSRVKSKRIQLGLNQAELAQKVGTTQQSIEQLENGKTKRPRFLPELASALGVSVDWLLNGTSDSNVRFVGHVEPKGKYPLISMVRAGSWCEA (SEQ ID NO: 313).
[0119] In some embodiments, the Bac434 protein is encoded by a polynucleotide comprising a sequence having 80% or more sequence identity, e.g., 85%, 90%, or 95% or more sequence identity, and in some cases 96%, 97%, 98%, or 99% sequence identity, or in some cases 100% sequence identity, to AUGAGCAUCAGCAGCCGCGUGAAGAGCAAGCGCAUCCAGCUGGGCCUGAACCAGGCCGAGCUGGCCCAGAAGGUGGGCACCACCCAGCAGAGCAUCGAGCAGCUGGAGAACGGCAAGACCAAGCGCCCCCGCUUCCUGCCCGAGCUGGCCAGCGCCCUGGGCGUGAGCGUGGACUGGCUGCUGAACGGCACCAGCGACAGCAACGUGCGCUUCGUGGGCCACGUGGAGCCCAAGGGCAAGUACCCCCUGAUCAGCAUGGUGCGCGCCGGCAGCUGGUGCGAGGCC (SEQ ID NO: 314). In some embodiments, the polynucleotide sequence encoding the Bac434 protein is codon-optimized for expression in a particular host cell, for example, codon-optimized for expression in a human or mouse cell.
[0120] As another example, if the DTS contains a binding sequence for GCN4 protein, e.g., AGTGACTCATT (SEQ ID NO: 315) or a variant thereof, the cell can also be contacted with mRNA encoding the GCN4 protein. In some embodiments, the GCN4 protein has been engineered to contain an NLS. In some embodiments, the NLS is proximal to the N-terminus. In other embodiments, the NLS is proximal to the C-terminus. In some embodiments, the GCN4 protein is engineered to contain an NLS at both the N-terminus and the C-terminus. In some embodiments, the NLS is fused to the GCN4 protein with a linker. In some embodiments, the NLS is selected from those listed in Table 7.
[0121] In some embodiments, the GCN4 protein is derived from Saccharomyces cerevisiae The wild-type sequence for GCN4 includes sequences having 90% or more, and in some cases 92%, 93%, 94%, 95% or more sequence identity, and in certain cases 96%, 97%, 98%, 99% or more identity, for example, 100% identity, to the wild-type sequence for GCN4: MSEYQPSLFALNPMGFSPLDGSKSTNENVSASTSTAKPMVGQLIFDKFIKTEEDPIIKQDTPSNLDFDFALPQTATAPDAKTVLPIPELDDAVVESFFSSSTDSTPMFEYENLEDNSKEWTSLFDNDIPVTTDDVSLADKAIESTEEVSLVPSNLEVSTTSFLPTPVLEDAKLTQTRKVKKPNSVVKKSHHVGKDDESRLDHLGVVAYNRKQRSIPLSPIVPESSDPAALKRARNTEAARRSRARKLQRMKQLEDKVEELLSKNYHLENEVARLKKLVGER (SEQ ID NO: 316).
[0122] In some embodiments, the GCN4 protein is encoded by a polynucleotide comprising a sequence having 80% or more sequence identity to (SEQ ID NO: 317), e.g., 85%, 90%, or 95% or more sequence identity, and in some cases 96%, 97%, 98%, or 99% sequence identity, and in some cases 100% sequence identity.In some embodiments, the GCN polynucleotide is codon-optimized for expression in a particular host cell, for example, codon-optimized for expression in a human or mouse cell.
[0123] As another example, if the DTS contains a binding sequence for the Lac repressor protein ("LacR"), e.g., TTGTTATCCGCTCACAA (SEQ ID NO: 318) or a variant thereof, the cell can also be contacted with mRNA encoding the LacR protein. In some embodiments, the LacR protein has been engineered to contain an NLS. In some embodiments, the NLS is proximal to the N-terminus. In other embodiments, the NLS is proximal to the C-terminus. In some embodiments, the LacR protein is engineered to contain an NLS at both the N-terminus and the C-terminus. In some embodiments, the NLS is fused to the LacR protein with a linker. In some embodiments, the NLS is selected from those listed in Table 7.
[0124] In some embodiments, the LacR protein has the wild-type sequence for the Escherichia coli DNA-binding transcriptional repressor LacI:MKPVTLYDVAEYAGVSYQTVSRVVNQASHVSAKTREKVEAAMAELNYIPNRVAQQLAGKQSLLIGVATSSLALHAPSQIVAAIKSRADQLGASVVVSMVERSGVEACKAAVHNLLAQRVSGLIINYPLDDQDAIAVEAACTNVPALFLDVSDQTPINSIIFSHEDGTRLGVEHLVALGHQQIALLAGPLSSVSARLRLAGWHKYLTRNQIQPIAEREGDWS AMSGFQQTMQMLNEGIVPTAMLVANDQMALGAMRAITESGLRVGADISVVGYDDTEDSSCYIPPLTTIKQDFRLLGQTSVDRLLQLSQGQAVKGNQLLPVSLVKRKTTLAPNTQTASPRALADSLMQLARQVSRLESGQ (SEQ ID NO: 319) having 90% or more, in some cases 92%, 93%, 94%, 95% or more sequence identity, and in certain cases 96%, 97%, 98%, 99% or more identity, for example, 100% identity.
[0125]
[0126] As another example, if the DTS contains a binding sequence for the I-SceI D44A protein, e.g., TAGGGATAACAGGGTAAT (SEQ ID NO: 311) or a variant thereof, the cell can also be contacted with mRNA encoding the I-SceI D44A protein. In some embodiments, the I-SceI protein has been engineered to contain an NLS. In some embodiments, the NLS is proximal to the N-terminus. In other embodiments, the NLS is proximal to the C-terminus. In some embodiments, the I-SceI protein is engineered to contain an NLS at both the N-terminus and the C-terminus. In some embodiments, the NLS is fused to the I-SceI protein with a linker.
[0127] In some embodiments, the I-SceI protein has 90% or more, and in some cases 92%, 93%, 94%, 95% or more sequence identity, and in certain cases 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 1109%, 1111, 1120, 1121, 1130, 1131, 1132, 1133, 1140, 1141, 1142, 1143, 1144, 1145, 1146, 1147, 1148, 1149, 1150, 1151, 1152, 1153, 1154, 1155, 1156, 1157, 1158, 1159, 1160, 1161, 1162, 1163, 1164, 1165, 1166, 1167, 1168, 1169, 1170, 1171, 1172, 1173, 1174, 1175, 1176, 1177, 1178, 1179, 1180, 1181, 1182, 1183, 1184, 1185, 1186, 1187, 1188, 1189, 1190, 1191, 1192, 1193, 1194, 1195, 1196, 1197, % or more identity, e.g., a sequence that is 100% identical, e.g., a sequence that is 85% or more identity to MPKKKRKVPKKHAAPPKKKRKVEDPRFMYPYDVPDYAGMKNIKKNQVMNLGPNSKLLKEYKSQLIELNIEQFEAGIGLILGAAYIRSRDEGKTYCMQFEWKNKAYMDHVCLLYDQWVLSPPHKKERVNHLGNLVITWGAQTFKHQAFNKLANLFIVNNKKTIPNNLVENYLTPMSLAYWFMDDGGKWDYNKNSTNKSIVLNTQSFTFEEVEYLVKGLRNKFQLNCYVKINKNKPIIYIDSMSYLIFYNLIKPYLIPQMMYKLPNTISSETFLK (SEQ ID NO: 313).
[0128] As another example, if the TFBS is a binding sequence for a gene editing system, such as a guide RNA for a Cas nuclease, a zinc finger for a zinc finger nuclease, or a TALE for a TALEN, as described below and known in the art, the cell can also be contacted with an mRNA encoding the Cas nuclease, ZF nuclease, or TALEN. In some certain embodiments, the nuclease of the gene editing system is catalytically active. In other words, the nuclease can modify DNA, for example, to nick the DNA and create a double-stranded break and replace one or more nucleotides therein. In other certain embodiments, the gene editing system is catalytically attenuated, e.g., its activity is reduced by 50% or more, in some cases by 60%, 70%, 80%, 90% or more, or in some cases by 95% or more, e.g., catalytic activity is abolished.
[0129] target cell Cells used in embodiments of the present invention, referred to herein as "target cells," can vary. It should be understood that target cells can be of any origin, e.g., from any organism. In some embodiments, target cells are mammalian cells. Some non-limiting examples of mammalian cells include, but are not limited to, mouse cells, rat cells, hamster cells, rodent cells, and non-human primate cells. In some embodiments, target cells are human cells. It should also be understood that target cells can be of any cell type. For example, target cells can be stem cells, which can include embryonic stem cells, induced pluripotent stem cells (iPS cells), fetal stem cells, umbilical cord blood stem cells, or adult stem cells (i.e., tissue-specific stem cells). In other cases, target cells can be any differentiated cell type found in a subject. Cells of interest include both dividing and non-dividing cells. Examples of specific target cells of interest include, but are not limited to, hepatocytes, astrocytes, T lymphocytes, B lymphocytes, NK cells, skeletal muscle cells, cardiomyocytes, neurons, astrocytes, oligodendrocytes, dendritic cells, skin cells, photoreceptors, RPE cells, radial glia, and the like.
[0130] In some embodiments, the target cell is a cell in vitro and the method comprises contacting the cell in vitro. In some embodiments, the target cell is a cell in a subject and the method comprises administering to the subject NTDNA and, optionally, an inducer (where either or both can be present in the same or different suitable delivery vehicles, as desired). In some embodiments, the subject is a mammalian subject, e.g., a rodent, mouse, rat, hamster, or non-human primate. In some embodiments, the subject is a human subject.
[0131] In embodiments of the invention in which an inducing agent is used, the target cells may be contacted with the inducing agent simultaneously or sequentially, as desired. In some cases, the target cells are contacted with the NTDNA and the inducing agent simultaneously. Thus, the NTDNA inducing agent is contacted with the target cells simultaneously. In other cases, the NTDNA and the inducing agent are contacted with the cells sequentially. For example, the target cells may be contacted with the NTDNA before contacting them with the inducing agent. Alternatively, the target cells may be contacted with the NTDNA after contacting them with the inducing agent.
[0132] In embodiments of the present invention in which a nuclear targeting factor is provided to cells, the target cells may be contacted simultaneously or sequentially with the mRNA encoding the NTF. In some cases, the target cells are contacted simultaneously with the NTF DNA and the NTF. In certain such embodiments, the NTF DNA and the NTF are provided as a single formulation in a delivery vehicle (described in more detail below) that is administered to the target cells. In other embodiments, the NTF DNA and the NTF are provided as separate formulations, i.e., as a mixture, that are administered to the cells simultaneously. Preferably, the NTF DNA and the NTF are prepared as a single formulation in a delivery vehicle for contact with the cells.
[0133] Genome integration In some cases, the NT DNA is contacted with a target cell in conjunction with a gene editing system configured to provide, for example, genomic integration of a cargo nucleic acid component of the NT DNA. In such embodiments, when the NT DNA is contacted with a target cell in conjunction with a gene editing system, the NT DNA may be contacted with the target cell simultaneously or sequentially with the gene editing system. In some cases, one or more components of the gene editing system may be present with the NT DNA in a delivery composition, such as a cytosolic delivery composition (e.g., LNP), such as those described in more detail below. The gene editing system that may be used in such embodiments may vary as desired. In some embodiments, the gene editing system used is configured to genomically integrate the cargo nucleic acid into a specific safe harbor location within the genome, for example, a genomic location within or near the endogenous albumin locus. Generally, a safe harbor locus is a location within the genome that can be used to integrate an exogenous nucleic acid, where addition of the exogenous nucleic acid to the safe harbor locus does not significantly affect the growth of the host cell by itself. One of skill in the art will understand that In some embodiments, a cargo nucleic acid may be inserted into a specific safe harbor location within the genome, which may either utilize a promoter found in that safe harbor locus or allow for controlled expression of the cargo nucleic acid's coding sequence by an exogenous promoter fused to the cargo nucleic acid coding sequence prior to insertion.
[0134] Gene editing can be performed using nucleases engineered to target specific sequences. To date, four major types of nucleases exist: meganucleases and their derivatives, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and the CRISPR-Cas9 nuclease system. Nuclease platforms vary in design difficulty, targeting density, and mode of action, particularly because the specificity of ZFNs and TALENs is mediated by protein-DNA interactions, while Cas9 is primarily guided by RNA-DNA interactions. Cas9 cleavage also requires a flanking motif PAM, which differs between different CRISPR systems. Cas9 from Streptococcus pyogenes cleaves using the NRG PAM, while CRISPR from Neisseria meningitidis can cleave at sites with PAMs including NNNNGATT, NNNNGTTTT, and NNNNGCTT. Some other Cas9 orthologues target protospacers adjacent to alternative PAMs. CRISPR endonucleases such as Cas9 can be used in various embodiments of the disclosed methods. However, the teachings described herein, such as therapeutic targeting sites, can be applied to other forms of endonucleases, such as ZFN, TALEN, HE, or MegaTAL, or using combinations of nucleases. These different systems are now described in more detail.
[0135] A given NT DNA delivery cell substrate delivery composition may contain one or more elements of a given gene editing system. The gene editing system used in embodiments of the present invention may contain several different elements, such as a nucleic acid element, such as a genome-targeting nucleic acid or guide RNA, a nucleic acid, such as an mRNA encoding an endonuclease, a polypeptide component, such as an endonuclease, etc. These various components will now be discussed in more detail in conjunction with a description of a representative endonuclease-based genome integration system.
[0136] Thus, in some embodiments, genome editing and composition methods use nucleic acid sequences (or oligonucleotides) encoding site-directed polypeptides or DNA endonucleases. The nucleic acid sequence encoding the site-directed polypeptide can be DNA or RNA. If the nucleic acid sequence encoding the site-directed polypeptide is RNA, the nucleic acid sequence can be covalently linked to the gRNA sequence or can exist as a separate sequence. In some embodiments, the peptide sequence of the site-directed polypeptide or DNA endonuclease can be used in place of the nucleic acid sequence.
[0137] Modifications of target DNA resulting from NHEJ and / or HDR can lead to, for example, mutations, deletions, alterations, integrations, gene corrections, gene replacements, gene tagging, transgene insertions, nucleotide deletions, gene disruptions, translocations, and / or gene mutations. The process of integrating a cargo nucleic acid, such as NTDNA, into genomic DNA is an example of genome editing. Site-specific polypeptides are nucleases used in genome editing to cleave DNA. Site-specific polypeptides can be administered to cells or patients as either one or more polypeptides or one or more mRNAs encoding the polypeptides. In some embodiments, the site-specific polypeptide has multiple nucleic acid cleavage (i.e., nuclease) domains. Two or more nucleic acid cleavage domains can be linked together via a linker. In some embodiments, the linker has a flexible linker. The linker can have a length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40 or more amino acids.
[0138] In the context of a CRISPR / Cas or CRISPR / Cpf1 system, a site-directed polypeptide can bind to a guide RNA, which then specifies the site within the target DNA to which the polypeptide is directed. In some embodiments of the CRISPR / Cas or CRISPR / Cpf1 system herein, the site-directed polypeptide is an endonuclease, such as a DNA endonuclease.
[0139] Naturally occurring wild-type Cas9 enzymes have two nuclease domains, an HNH nuclease domain and a RuvC domain. As used herein, "Cas9" refers to both naturally occurring and recombinant Cas9. Cas9 enzymes contemplated herein have an HNH or HNH-like nuclease domain and / or a RuvC or RuvC-like nuclease domain.
[0140] The HNH or HNH-like domain has an McrA-like fold. The HNH or HNH-like domain has two antiparallel β-strands and an α-helix. The HNH or HNH-like domain has a metal-binding site (e.g., a divalent cation-binding site). The HNH or HNH-like domain can cleave one strand of a target nucleic acid (e.g., the complementary strand of the crRNA target strand).
[0141] The RuvC or RuvC-like domain has an RNaseH or RNaseH-like fold. The RuvC / RNaseH domain is involved in a diverse set of nucleic acid-based functions, including acting on both RNA and DNA. The RNaseH domain has five β-strands surrounded by multiple a-helices. The RuvC / RNaseH or RuvC / RNaseH-like domain has a metal-binding site (e.g., a divalent cation-binding site). The RuvC / RNaseH or RuvC / RNaseH-like domain can cleave one strand of a target nucleic acid (e.g., the non-complementary strand of a double-stranded target DNA).
[0142] In some embodiments, the site-directed polypeptide has an amino acid sequence that has at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% amino acid sequence identity to a wild-type exemplary site-directed polypeptide (e.g., Cas9 from S. pyogenes, US2014 / 0068797 SEQ ID NO:8 or Sapranauskas et al., Nucleic Acids Res, 39(21):9275-9282(2011))] and various other site-directed polypeptides). In some embodiments, the site-directed polypeptide has an amino acid sequence that has at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% amino acid sequence identity to the nuclease domain of a wild-type exemplary site-directed polypeptide (e.g., Cas9 from S. pyogenes, supra). In some embodiments, the site-directed polypeptide has at least 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to the wild-type site-directed polypeptide (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids. In some embodiments, the site-directed polypeptide has up to 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type site-directed polypeptide (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids. In some embodiments, the site-directed polypeptide has at least 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type site-directed polypeptide (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in the HNH nuclease domain of the site-directed polypeptide.In some embodiments, the site-directed polypeptide has up to 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type site-directed polypeptide (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in the HNH nuclease domain of the site-directed polypeptide. In some embodiments, the site-directed polypeptide has at least 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type site-directed polypeptide (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in the RuvC nuclease domain of the site-directed polypeptide. In some embodiments, the site-directed polypeptide has up to 70, 75, 80, 85, 90, 95, 97, 99, or 100% identity to a wild-type site-directed polypeptide (e.g., Cas9 from S. pyogenes, supra) over 10 contiguous amino acids in the RuvC nuclease domain of the site-directed polypeptide.
[0143] In some embodiments, the site-directed polypeptide comprises a modified form of a wild-type exemplary site-directed polypeptide. The modified form of the wild-type exemplary site-directed polypeptide comprises a mutation that reduces the nucleic acid cleavage activity of the site-directed polypeptide. In some embodiments, the modified form of the wild-type exemplary site-directed polypeptide has less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nucleic acid cleavage activity of the wild-type exemplary site-directed polypeptide (e.g., Cas9 from S. pyogenes, supra). The modified form of the site-directed polypeptide may not have substantial nucleic acid cleavage activity. When a site-directed polypeptide is a modified form that does not have substantial nucleic acid cleavage activity, it is referred to herein as "enzymatically inactive."
[0144] In some embodiments, the modified form of the site-directed polypeptide has a mutation such that it is capable of inducing a single-strand break (SSB) on the target nucleic acid (e.g., by cleaving only one of the sugar-phosphate backbones of a double-stranded target nucleic acid). In some embodiments, the mutation results in less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the nucleic acid cleavage activity in one or more of the multiple nucleic acid cleavage domains of a wild-type site-directed polypeptide (e.g., Cas9 from S. pyogenes, supra). In some embodiments, the mutation results in one or more of the multiple nucleic acid cleavage domains retaining the ability to cleave the complementary strand of the target nucleic acid but reducing their ability to cleave the non-complementary strand of the target nucleic acid. In some embodiments, the mutation results in one or more of the multiple nucleic acid cleavage domains retaining the ability to cleave the non-complementary strand of the target nucleic acid but reducing their ability to cleave the complementary strand of the target nucleic acid. For example, residues in a wild-type exemplary S. pyogenes Cas9 polypeptide, such as AsplO, His840, Asn854, and Asn856, are mutated to inactivate one or more of the nucleic acid cleavage domains (e.g., nuclease domains). In some embodiments, the mutated residues correspond to residues AsplO, His840, Asn854, and Asn856 in a wild-type exemplary S. pyogenes Cas9 polypeptide (e.g., as determined by sequence and / or structural alignment). Non-limiting examples of mutations include D10A, H840A, N854A, or N856A. One of skill in the art will recognize that mutations other than alanine substitutions are suitable.
[0145] In some embodiments, the D10A mutation is combined with one or more of an H840A, an N854A, or an N856A mutation to produce a site-directed polypeptide that substantially lacks DNA cleavage activity. In some embodiments, the H840A mutation is combined with one or more of a D10A, an N854A, or an N856A mutation to produce a site-directed polypeptide that substantially lacks DNA cleavage activity. In some embodiments, the N854A mutation is combined with one or more of an H840A, a D10A, or an N856A mutation to produce a site-directed polypeptide that substantially lacks DNA cleavage activity. In some embodiments, the N856A mutation is combined with one or more of an H840A, an N854A, or a D10A mutation to produce a site-directed polypeptide that substantially lacks DNA cleavage activity. Site-directed polypeptides with one substantially inactive nuclease domain are referred to as "nickases."
[0146] In some embodiments, RNA-guided endonuclease variants, such as Cas9, can be used to increase the specificity of CRISPR-mediated genome editing. Wild-type Cas9 is generally guided by a single guide RNA designed to hybridize with a specific sequence of about 20 nucleotides in a target sequence (such as an endogenous genomic locus). However, some mismatches can be tolerated between the guide RNA and the target locus, effectively reducing the required homology length at the target site to, for example, only 13 nt of homology, thereby increasing the possibility of CRISPR / Cas9 complex binding and double-stranded nucleic acid cleavage at other locations in the target genome, also known as off-target cleavage. Because Cas9 nickase variants each cleave only one strand, to create a double-stranded break, a pair of nickases must be adjacent and bind to opposite strands of the target nucleic acid, thereby creating a pair of nicks, which is the equivalent of a double-stranded break. This requires that two separate guide RNAs, one for each nickase, must be adjacent and bind to opposite strands of the target nucleic acid. This requirement essentially doubles the minimum length of homology required for double-strand breaks to occur, thereby reducing the likelihood that double-strand breaks will occur elsewhere in the genome, where the two guide RNA sites, if present, are unlikely to be close enough to each other to allow double-strand breaks to form. As described in the art, nickases can also be used to promote HDR versus NHEJ. HDR can be used to introduce selected changes into target sites in the genome through the use of specific donor sequences that effectively mediate the desired changes. Descriptions of various CRISPR / Cas systems for use in gene editing can be found, for example, in International Patent Application Publication No. 2013 / 176772 and Nature Biotechnology 32, 347-355 (2014), as well as the references cited therein.
[0147] In some embodiments, the site-directed polypeptide (e.g., a variant, mutated, enzymatically inactive, and / or conditionally enzymatically inactive site-directed polypeptide) targets a nucleic acid. In some embodiments, the site-directed polypeptide (e.g., a variant, mutated, enzymatically inactive, and / or conditionally enzymatically inactive endoribonuclease) targets DNA. In some embodiments, the site-directed polypeptide (e.g., a variant, mutated, enzymatically inactive, and / or conditionally enzymatically inactive endoribonuclease) targets RNA.
[0148] In some embodiments, the site-directed polypeptide has one or more non-native sequences (e.g., the site-directed polypeptide is a fusion protein). In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 15% amino acid identity to Cas9 from a bacterium (e.g., S. pyogenes), a nucleic acid binding domain, and two nucleic acid cleavage domains (i.e., an HNH domain and a RuvC domain). In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 15% amino acid identity to Cas9 from a bacterium (e.g., S. pyogenes), and two nucleic acid cleavage domains (i.e., an HNH domain and a RuvC domain). In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 15% amino acid identity to Cas9 from a bacterium (e.g., S. pyogenes), and two nucleic acid cleavage domains, one or both of the nucleic acid cleavage domains having at least 50% amino acid identity to the nuclease domain from Cas9 from a bacterium (e.g., S. pyogenes). In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 15% amino acid identity to Cas9 from a bacterium (e.g., S. pyogenes), two nucleic acid cleavage domains (i.e., an HNH domain and a RuvC domain), and a non-native sequence (e.g., a nuclear localization signal) or a linker connecting the site-directed polypeptide to the non-native sequence. In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 15% amino acid identity to Cas9 from a bacterium (e.g., S. pyogenes), two nucleic acid cleavage domains (i.e., an HNH domain and a RuvC domain), and the site-directed polypeptide has a mutation in one or both of the nucleic acid cleavage domains that reduces the cleavage activity of the nuclease domain by at least 50%.In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 15% amino acid identity to Cas9 from a bacterium (e.g., S. pyogenes) and two nucleic acid cleavage domains (i.e., an HNH domain and a RuvC domain), wherein one of the nuclease domains has a mutation at aspartic acid 10 and / or one of the nuclease domains has a mutation at histidine 840, which mutation reduces the cleavage activity of the nuclease domain by at least 50%.
[0149] In some embodiments, the one or more site-specific polypeptides, e.g., DNA endonucleases, comprise two nickases that together create one double-stranded break at a specific locus within the genome, or four nickases that together create two double-stranded breaks at a specific locus within the genome. Alternatively, one site-specific polypeptide, e.g., DNA endonuclease, affects one double-stranded break at a specific locus within the genome.
[0150] In some embodiments, a polynucleotide encoding a site-directed polypeptide can be used to edit a genome. In some of these embodiments, the polynucleotide encoding the site-directed polypeptide is codon-optimized according to standard methods in the art for expression in cells containing the target DNA of interest. For example, when the intended target nucleic acid is in a human cell, it is contemplated to use a human codon-optimized polynucleotide encoding Cas9 to produce Cas9 polypeptide.
[0151] CRISPR (clustered regularly interspaced short palindromic repeats) genomic loci can be found in the genomes of many prokaryotes (e.g., bacteria and archaea). In prokaryotes, CRISPR loci encode products that function as a type of immune system that helps defend prokaryotes against foreign invaders such as viruses and phages. There are three stages of CRISPR locus function: integration of new sequences into the CRISPR locus, expression of CRISPR RNA (crRNA), and silencing of the foreign invader nucleic acid. Five types of CRISPR systems (e.g., type I, type II, type III, type U, and type V) have been identified.
[0152] CRISPR loci contain several short repetitive sequences called "repeats." When expressed, the repeats can form secondary hairpin structures (e.g., hairpins) and / or have unstructured single-stranded sequences. The repeats usually occur in clusters and frequently diverge between species. The repeats are regularly spaced with unique intervening sequences called "spacers," resulting in a repeat-spacer-repeat locus architecture. The spacers are identical to or highly homologous to known foreign invader sequences. The spacer-repeat units encode crisprRNAs (crRNAs), which are processed into the mature form of the spacer-repeat units. The crRNA has a "seed" or spacer sequence responsible for targeting the target nucleic acid (in the naturally occurring form in prokaryotes, the spacer sequence targets the foreign invader nucleic acid). The spacer sequence is located at the 5' or 3' end of the crRNA.
[0153] CRISPR loci also contain polynucleotide sequences encoding CRISPR-associated (Cas) genes. Cas genes encode endonucleases involved in the biosynthesis and interference steps of crRNA function in prokaryotes. Some Cas genes share homologous secondary and / or tertiary structures.
[0154] In natural type II CRISPR systems, crRNA biogenesis requires a trans-activating CRISPR RNA (tracrRNA). The tracrRNA is modified by endogenous RNase III and then hybridizes to crRNA repeats within the pre-crRNA array. Endogenous RNase III is recruited to cleave the pre-crRNA. The cleaved crRNA is subjected to exoribonuclease trimming to produce the mature crRNA form (e.g., 5' trimming). The tracrRNA remains hybridized to the crRNA, and the tracrRNA and crRNA associate with a site-specific polypeptide (e.g., Cas9). The crRNA in the crRNA-tracrRNA-Cas9 complex guides the complex to a target nucleic acid to which the crRNA can hybridize. Hybridization of the crRNA to the target nucleic acid activates Cas9 for target nucleic acid cleavage. The target nucleic acid in type II CRISPR systems is referred to as a protospacer adjacent motif (PAM). In nature, PAM is essential for facilitating the binding of site-specific polypeptides (e.g., Cas9) to target nucleic acids. Type II systems (also called Nmeni or CASS4) are further subdivided into type II-A (CASS4) and type II-B (CASS4a). Jinek et al., Science, 337(6096):816-821(2012) demonstrated that the CRISPR / Cas9 system is useful for RNA-programmable genome editing, and International Patent Application Publication No. 2013 / 176772 provides numerous examples and applications of CRISPR / Cas endonuclease systems for site-specific gene editing.
[0155] Type V CRISPR systems have several key differences from type II systems. For example, Cpf1 is a single RNA-guided endonuclease that lacks a tracrRNA, in contrast to type II systems. Indeed, Cpf1-associated CRISPR arrays are processed into mature crRNAs without the need for an additional transactivating tracrRNA. Type V CRISPR arrays are processed into short mature crRNAs, 42–44 nucleotides in length, each of which begins with a 19-nucleotide direct repeat followed by a 23–25-nucleotide spacer sequence. In contrast, mature crRNAs in type II systems begin with a 20–24-nucleotide spacer sequence followed by approximately 22-nucleotide direct repeats. Furthermore, Cpf1 utilizes a T-rich protospacer-adjacent motif, resulting in the Cpf1-crRNA complex efficiently cleaving target DNA preceded by a short T-rich PAM, in contrast to the G-rich PAM that follows target DNA in type II systems. Thus, type V systems cleave at points distal to the PAM, whereas type II systems cleave at points adjacent to the PAM. Additionally, in contrast to type II systems, Cpf1 cleaves DNA via staggered DNA double-strand breaks in 4- or 5-nucleotide 5' overhangs. Type II systems cleave via blunt double-strand breaks. Like type II systems, Cpf1 contains a predicted RuvC-like endonuclease domain but lacks the second HNH endonuclease domain, which is in contrast to type II systems.
[0156] Exemplary CRISPR / Cas polypeptides include the Cas9 polypeptide in Figure 1 of Fonfara et al., Nucleic Acids Research, 42:2577-2590 (2014). The CRISPR / Cas gene nomenclature system has been extensively rewritten since the discovery of Cas genes.
[0157] The genome-targeting nucleic acid interacts with a site-specific polypeptide (e.g., a nucleic acid-guided nuclease such as Cas9), thereby forming a complex. The genome-targeting nucleic acid (e.g., a gRNA, such as those described in more detail below) guides the site-specific polypeptide to the target nucleic acid.
[0158] In some embodiments, the site-directed polypeptide and the genome-targeting nucleic acid can each be administered separately to a cell or patient. While in some other embodiments, the site-directed polypeptide can be pre-complexed with one or more crRNAs along with one or more guide RNAs or tracrRNAs. The pre-complexed material can then be administered to a cell or patient. Such pre-complexed material is known as a ribonucleoprotein particle (RNP).
[0159] Genome-targeting nucleic acid or guide RNA. Genome editing components may include a genome-targeting nucleic acid that can direct the activity of an associated polypeptide (e.g., a site-specific polypeptide or a DNA endonuclease) to a specific target sequence within a target nucleic acid. In some embodiments, the genome-targeting nucleic acid is RNA. Genome-targeting RNA is referred to herein as a "guide RNA" or "gRNA." A guide RNA has at least a spacer sequence that hybridizes to a target nucleic acid sequence of interest and a CRISPR repeat sequence. In type II systems, the gRNA also has a second RNA called a tracrRNA sequence. In type II guide RNA (gRNA), the CRISPR repeat sequence and the tracrRNA sequence hybridize to each other to form a duplex. In type V guide RNA (gRNA), the crRNA forms a duplex. In both systems, the duplex binds to the site-specific polypeptide such that the guide RNA and the site-specific polypeptide form a complex. The genome-targeting nucleic acid provides target specificity to the complex through its association with the site-specific polypeptide. Thus, the genome-targeting nucleic acid directs the activity of the site-specific polypeptide.
[0160] In some embodiments, the genome-targeting nucleic acid is a dual-molecule guide RNA. In some embodiments, the genome-targeting nucleic acid is a single-molecule guide RNA. A dual-molecule guide RNA has two RNA strands. The first strand has, from 5' to 3', an optional spacer extension sequence, a spacer sequence, and a minimal CRISPR repeat sequence. The second strand has a minimal tracrRNA sequence (complementary to the minimal CRISPR repeat sequence), a 3' tracrRNA sequence, and an optional tracrRNA extension sequence. In Type II systems, a single-molecule guide RNA (sgRNA) has, from 5' to 3', an optional spacer extension sequence, a spacer sequence, a minimal CRISPR repeat sequence, a single-molecule guide linker, a minimal tracrRNA sequence, a 3' tracrRNA sequence, and an optional tracrRNA extension sequence. The optional tracrRNA extension may have elements that contribute additional functionality (e.g., stability) to the guide RNA. The single molecule guide linker connects the minimal CRISPR repeat and the minimal tracrRNA sequence to form a hairpin structure. The optional tracrRNA extension has one or more hairpins. In the V-type system, the single molecule guide RNA (sgRNA) has, from 5' to 3', the minimal CRISPR repeat sequence and a spacer sequence.
[0161] By way of example, guide RNAs or other smaller RNAs used in the CRISPR / Cas / Cpf1 system can be readily synthesized by chemical means, as exemplified below and described in the art. While chemical synthesis procedures are continually expanding, purification of such RNAs by procedures such as high-performance liquid chromatography (HPLC) (which avoids the use of gels such as PAGE) tends to become more difficult as polynucleotide lengths increase significantly beyond about 100 nucleotides. One approach used to generate longer RNAs is to produce two or more molecules that are ligated together. Much longer RNAs, such as those encoding Cas9 or Cpf1 endonucleases, are more easily produced enzymatically. Various types of RNA modifications, for example, modifications that enhance stability, reduce the likelihood or severity of innate immune responses, and / or enhance other attributes, as described in the art, can be introduced during or after chemical synthesis and / or enzymatic production of the RNA.
[0162] In some embodiments of genome-targeting nucleic acids, a spacer extension sequence can modify activity, provide stability, and / or provide a location for modification of the genome-targeting nucleic acid. The spacer extension sequence can modify on- or off-target activity or specificity. In some embodiments, a spacer extension sequence is provided. The spacer extension sequence can have a length of 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 1000, 2000, 3000, 4000, 5000, 6000, or 7000 or more nucleotides. The spacer extension sequence can have a length of about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 1000, 2000, 3000, 4000, 5000, 6000, or 7000 or more nucleotides. The spacer extension sequence can have a length of less than 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 1000, 2000, 3000, 4000, 5000, 6000, 7000, or more nucleotides. In some embodiments, the spacer extension sequence is less than 10 nucleotides in length. In some embodiments, the spacer extension sequence is 10-30 nucleotides in length. In some embodiments, the spacer extension sequence is 30-70 nucleotides in length.
[0163] In some embodiments, the spacer extension sequence comprises another moiety (e.g., a stability control sequence, an endoribonuclease binding sequence, a ribozyme). In some embodiments, the moiety decreases or increases the stability of the nucleic acid targeting nucleic acid. In some embodiments, the moiety is a transcription terminator segment (i.e., a transcription termination sequence). In some embodiments, the moiety functions in eukaryotic cells. In some embodiments, the moiety functions in prokaryotic cells. In some embodiments, the moiety functions in both eukaryotic and prokaryotic cells. Non-limiting examples of suitable moieties include a 5' cap (e.g., a 7-methylguanylate cap (m7G)), a riboswitch sequence (e.g., to allow for controlled stability and / or controlled accessibility by proteins and protein complexes), a sequence that forms a dsRNA duplex (i.e., a hairpin), a sequence that targets the RNA to a subcellular location (e.g., the nucleus, mitochondria, chloroplasts, etc.), a modification or sequence that provides tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows for fluorescent detection, etc.), and / or a modification or sequence that provides a binding site for a protein (e.g., a protein that acts on DNA, including a transcriptional activator, a transcriptional repressor, a DNA methyltransferase, a DNA demethylase, a histone acetyltransferase, a histone deacetylase, etc.).
[0164] Spacer sequence hybridizes with the sequence in the target nucleic acid of interest.The spacer of genome targeting nucleic acid interacts with target nucleic acid in a sequence-specific manner through hybridization (i.e., base pairing).Therefore, the nucleotide sequence of spacer varies according to the sequence of the target nucleic acid of interest.
[0165] In the CRISPR / Cas system herein, a spacer sequence is designed to hybridize with the target nucleic acid located 5' of the PAM of the Cas9 enzyme used in the system. The spacer can perfectly match the target sequence or have a mismatch. Each Cas9 enzyme has a specific PAM sequence that it recognizes in the target DNA. For example, S. pyogenes recognizes a PAM with the sequence 5'-NRG-3' in the target nucleic acid, where R is either A or G, and N is any nucleotide, and N is immediately 3' of the target nucleic acid sequence targeted by the spacer sequence.
[0166] In some embodiments, the target nucleic acid sequence has 20 nucleotides. In some embodiments, the target nucleic acid has fewer than 20 nucleotides. In some embodiments, the target nucleic acid has more than 20 nucleotides. In some embodiments, the target nucleic acid has at least 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or more nucleotides. In some embodiments, the target nucleic acid has up to 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or more nucleotides. In some embodiments, the target nucleic acid sequence has 20 bases immediately 5' to the first nucleotide of the PAM. For example, in a sequence having 5'-NNNNNNNNNNNNNNNNNNNNNRG-3', the target nucleic acid has a sequence corresponding to N, where N is any nucleotide, and the underlined NRG sequence (R is G or A) is a Streptococcus pyogenes Cas9 PAM. In some embodiments, the PAM sequence used in the compositions and methods of the disclosure, such as the sequence recognized by SpCas9, is NGG.
[0167] In some embodiments, the spacer sequence that hybridizes to the target nucleic acid has a length of at least about 6 nucleotides (nt), such as at least about 6 nt, about 10 nt, about 15 nt, about 18 nt, about 19 nt, about 20 nt, about 25 nt, about 30 nt, about 35 nt, or about 40 nt, or about 6 nt to about 80 nt, about 6 nt to about 50 nt, about 6 nt to about 45 nt, about 6 nt to about 40 nt, about 6 nt to about 35 nt, about 6 nt to about 30 nt, about 6 nt to about 25 nt, about 6 nt to about 20 nt, about 6 nt to about 19 nt, about 10 nt to about 50 nt, about 10 nt to about 45 nt, about 10 nt to about 40 nt, about 10 nt to about 35 nt, The spacer sequence may be about 10 nt to about 30 nt, about 10 nt to about 25 nt, about 10 nt to about 20 nt, about 10 nt to about 19 nt, about 19 nt to about 25 nt, about 19 nt to about 30 nt, about 19 nt to about 35 nt, about 19 nt to about 40 nt, about 19 nt to about 45 nt, about 19 nt to about 50 nt, about 19 nt to about 60 nt, about 20 nt to about 25 nt, about 20 nt to about 30 nt, about 20 nt to about 35 nt, about 20 nt to about 40 nt, about 20 nt to about 45 nt, about 20 nt to about 50 nt, or about 20 nt to about 60 nt. In some embodiments, the spacer sequence has 20 nucleotides. In some embodiments, the spacer has 19 nucleotides. In some embodiments, the spacer has 18 nucleotides. In some embodiments, the spacer has 17 nucleotides. In some embodiments, the spacer has 16 nucleotides. In some embodiments, the spacer has 15 nucleotides.
[0168] In some embodiments, the percent complementarity between the spacer sequence and the target nucleic acid is at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98%, at least about 99%, or 100%. In some embodiments, the percent complementarity between the spacer sequence and the target nucleic acid is up to about 30%, up to about 40%, up to about 50%, up to about 60%, up to about 65%, up to about 70%, up to about 75%, up to about 80%, up to about 85%, up to about 90%, up to about 95%, up to about 97%, up to about 98%, up to about 99%, or 100%. In some embodiments, the percent complementarity between the spacer sequence and the target nucleic acid is 100% over the 6 contiguous 5'-most nucleotides of the target sequence of the complementary strand of the target nucleic acid. In some embodiments, the percent complementarity between the spacer sequence and the target nucleic acid is at least 60% over about 20 contiguous nucleotides. In some embodiments, the length of the spacer sequence and the target nucleic acid may differ by 1-6 nucleotides, which can be considered a bulge or bulges.
[0169] In some embodiments, spacer sequences are designed or selected using a computer program that can use variables such as predicted melting temperature, secondary structure formation, predicted annealing temperature, sequence identity, genomic context, chromatin accessibility, %GC, genomic frequency of occurrence (e.g., of sequences that are identical or similar but vary in one or more spots as a result of mismatches, insertions, or deletions), methylation status, and the presence of SNPs.
[0170] Zinc finger nucleases. Zinc finger nucleases (ZFNs) are modular proteins with engineered zinc finger DNA-binding domains linked to the catalytic domain of the type II endonuclease FokI. Because FokI functions only as a dimer, a pair of ZFNs must be engineered to bind to cognate target "half-site" sequences on opposite DNA strands with precise spacing between them to allow the formation of catalytically active FokI dimers. Upon dimerization of the FokI domains, which themselves have no sequence specificity, a DNA double-strand break is generated between the ZFN half-sites as the initiating step in genome editing.
[0171] The DNA-binding domain of each ZFN generally has three to six zinc fingers with a rich Cys2-His2 architecture. Each finger primarily recognizes a triplet of nucleotides on one strand of the target DNA sequence, although interstrand interactions with a fourth nucleotide may also be important. Modifying the amino acids of a finger at positions that make critical contacts with DNA alters the sequence specificity of a given finger. Thus, a four-finger zinc finger protein selectively recognizes a 12-bp target sequence, where the target sequence is a composite of triplet preferences contributed by each finger, but triplet preferences can be influenced to varying degrees by adjacent fingers. An important aspect of ZFNs is that they can be easily retargeted to almost any genomic address simply by modifying individual fingers, although doing so successfully requires considerable expertise. Most applications of ZFNs use proteins with four to six fingers, each recognizing 12 to 18 bp. Thus, a pair of ZFNs generally recognizes a combined target sequence of 24–36 bp, not including the 5–7 bp spacer between the half sites. The binding sites can be further separated by larger spacers, including 15–17 bp. Target sequences of this length are likely unique in the human genome, assuming repetitive sequences or gene homologs are excluded during the design process. Nevertheless, ZFN protein-DNA interactions are not absolute in their specificity, and off-target binding and cleavage events can occur either as heterodimers between two ZFNs or as homodimers of one or the other ZFN. The latter possibility has been effectively eliminated by engineering the dimerization interface of the FokI domain to create "plus" and "minus" variants, also known as obligate heterodimer variants, that can dimerize only with each other and not with themselves. Enforcing obligate heterodimerization prevents homodimer formation. This significantly enhances the specificity of ZFNs and any other nucleases that employ these FokI variants.
[0172] Various ZFN-based systems have been described in the art, and modifications are regularly reported, and numerous references describe the rules and parameters used to guide the design of ZFNs. See, for example, Segal et al., Proc Natl Acad Sci USA 96(6):2758-63(1999); Dreier B et al., J Mol Biol.303(4):489-502(2000); Liu Q et al., J Biol Chem.277(6):3850-6(2002); Dreier et al., J Biol Chem 280(42):35588-97(2005); and Dreier et al., J Biol Chem.276(31):29466-78(2001).
[0173] Transcription activator-like effector nucleases (TALENs). TALENs, like ZFNs, represent another form of modular nuclease in which an engineered DNA-binding domain is linked to a FokI nuclease domain, and a pair of TALENs work in tandem to achieve targeted DNA cleavage. Their primary difference from ZFNs is the nature of the DNA-binding domain and the associated target DNA sequence recognition properties. TALEN DNA-binding domains are derived from the TALE protein, originally described in the plant bacterial pathogen Xanthomonas sp. TALEs contain a tandem array of 33-35 amino acid repeats, each of which recognizes a single base pair in the target DNA sequence, typically up to 20 bp long, giving a total target sequence length of up to 40 bp. The nucleotide specificity of each repeat is determined by a repeat variable dyad (RVD), which contains only two amino acids at positions 12 and 13. The bases guanine, adenine, cytosine, and thymine are primarily recognized by four RVDs: Asn-Asn, Asn-Ile, His-Asp, and Asn-Gly, respectively. This constitutes a much simpler recognition code than zinc fingers and therefore represents an advantage over the latter for nuclease design. Nevertheless, like ZFNs, TALEN protein-DNA interactions are not absolute in their specificity, and TALENs also benefit from the use of obligate heterodimeric variants of the FokI domain to reduce off-target activity.
[0174] Additional variants of the FokI domain have been created that are inactivated in its catalytic function. When either half of a TALEN or ZFN pair contains an inactive FokI domain, only single-strand DNA nicking occurs at the target site, rather than a DSB. The outcome is comparable to the use of CRISPR / Cas9 / Cpf1 "nickase" mutants in which one of the Cas9 cleavage domains is inactivated. DNA nicks can be used to drive genome editing via HDR, but with lower efficiency than DSBs. A key benefit is that off-target nicks are repaired quickly and accurately, unlike DSBs, which are prone to NHEJ-mediated misrepair.
[0175] A variety of TALEN-based systems have been described in the art, and their modifications are regularly reported.For example, see Boch, Science 326(5959):1509-12(2009); Mak et al., Science 335(6069):716-9(2012); and Moscou et al., Science 326(5959):1501(2009).The use of TALEN based on "Golden Gate" platform or cloning scheme has been described by several groups. See, for example, Cermak et al., Nucleic Acids Res. 39(12):e82 (2011); Li et al., Nucleic Acids Res. 39(14):6315-25 (2011); Weber et al., PLoS One. 6(2):e16765 (2011); Wang et al., J Genet Genomics 41(6):339-47, Epub 2014 Can 17 (2014); and Cermak T et al., Methods Mol Biol. 1239:133-59 (2015).
[0176] Homing endonucleases. Homing endonucleases (HEs) are sequence-specific endonucleases that have long recognition sequences (14–44 base pairs) and often cleave DNA with high specificity at unique sites within the genome. There are at least six known families of HEs, classified by their structure, including LAGLIDADG (SEQ ID NO: 6), GIY-YIG, His-Cis box, HNH, PD-(D / E)xK, and VSR-like, derived from a wide range of hosts, including eukaryotes, protists, bacteria, archaea, cyanobacteria, and phages. Similar to ZFNs and TALENs, HEs can be used to create DSBs at target loci as an initial step in genome editing. In addition, some natural and engineered HEs cleave only a single strand of DNA, thereby functioning as site-specific nickases. The large target sequences of HEs and the specificity they offer make them attractive candidates for creating site-specific DSBs.
[0177] Various HE-based systems have been described in the art, and modifications are regularly reported. See, for example, the reviews by Steentoft et al., Glycobiology 24(8):663-80 (2014); Belfort and Bonocora, Methods Mol Biol. 1123:1-26 (2014); Hafez and Hausner, Genome 55(8):553-69 (2012), and the references cited therein.
[0178] MegaTAL / Tev-mTALEN / MegaTev. As further examples of hybrid nucleases, the MegaTAL and Tev-mTALEN platforms utilize a fusion of a TALE DNA-binding domain and a catalytically active HE, taking advantage of both the tunable DNA binding and specificity of TALEs and the cleavage sequence specificity of HEs. See, e.g., Boissel et al., NAR 42:2591-2601 (2014); Kleinstiver et al., G3 4:1155-65 (2014); and Boissel and Scharenberg, Methods Mol. Biol. 1239:171-96 (2015).
[0179] In a further variation, the MegaTev architecture is a fusion of a meganuclease (Mega) with a nuclease domain derived from the GIY-YIG homing endonuclease I-TevI (Tev). The two active sites are spaced approximately 30 bp apart on the DNA substrate, generating two DSBs with incompatible sticky ends. See, for example, Wolfs et al., NAR 42, 8816-29 (2014). It is anticipated that other combinations of existing nuclease-based approaches will evolve and be useful in achieving the targeted genome modifications described herein.
[0180] dCas9-FokI or dCpf1-FokI and other nucleases. Combining the structural and functional properties of the above nuclease platforms offers an additional approach to genome editing that could potentially overcome some of their inherent deficiencies. As an example, CRISPR genome editing systems generally use a single Cas9 endonuclease to create DSBs. Targeting specificity is driven by a 20- or 22-nucleotide sequence in the guide RNA that undergoes Watson-Crick base pairing with the target DNA (plus, in the case of Cas9 from S. pyogenes, an additional two bases in the adjacent NAG or NGG PAM sequence). While such sequences are long enough to be unique in the human genome, the specificity of the RNA / DNA interaction is not absolute, and considerable perturbation can be tolerated, especially in the 5' half of the target sequence, effectively reducing the number of bases driving specificity. One solution to this problem is to completely inactivate the catalytic function of Cas9 or Cpf1, retaining only the RNA-guided DNA binding function, and instead fuse the FokI domain to the inactivated Cas9. See, e.g., Tsai et al., Nature Biotech 32:569-76 (2014); and Guilinger et al., Nature Biotech. 32:577-82 (2014). Because FokI must dimerize to become catalytically active, two guide RNAs are required to bring two FokI fusions into close proximity, dimerize, and cleave DNA. This essentially doubles the number of bases within the combined target site, thereby increasing the targeting stringency of CRISPR-based systems.
[0181] As a further example, fusion of a TALE DNA binding domain to a catalytically active HE such as I-TevI is expected to take advantage of both the tunable DNA binding and specificity of the TALE and the cleavage sequence specificity of I-TevI, further reducing off-target cleavage.
[0182] Additional details regarding gene editing systems that find use in embodiments of the present invention can be found in U.S. Published Patent Application No. 2021 / 0348159.
[0183] Delivery Compositions and Methods If desired, the NT DNA can be present in a delivery composition comprising the NT DNA and a delivery vehicle component, e.g., a delivery vehicle component that mediates entry of the NT DNA into the cytosol from an extracellular location. In some aspects, the methods provided herein include delivering the NT DNA to a target cell. Also provided herein are cells produced by such methods and organisms (such as animals, plants, or fungi) containing or produced from such cells. Nucleic acid delivery methods can include lipofection, nucleofection, microinjection, biolistics, liposomes, immunoliposomes, polycations, or lipid:nucleic acid conjugates, naked DNA, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Delivery can be to a cell (eg, in vitro or ex vivo administration) or to a target tissue (eg, in vivo administration), as indicated above.
[0184] Various techniques and methods for delivering nucleic acids to cells are known in the art. For example, NTNDAs can be delivered to cells by conjugating the nucleic acid with a ligand that is internalized by the cell. For example, the ligand can bind to a receptor on the cell surface and be internalized via endocytosis. The ligand can be covalently bound to a nucleotide in the nucleic acid. Exemplary conjugates for delivering nucleic acids to cells are described in, for example, WO2015 / 006740, WO2014 / 025805, WO2012 / 037254, WO2009 / 082606, WO2009 / 073809, WO2009 / 018332, WO2006 / 112872, WO2004 / 090108, WO2004 / 091515, and WO2017 / 177326.
[0185] NT DNA can also be delivered to cells by transfection, for example, as described herein. Useful transfection methods include, but are not limited to, lipid-mediated transfection, cationic polymer-mediated transfection, or calcium phosphate precipitation. Transfection reagents are well known in the art and include TurboFect transfection reagent (Thermo Fisher Scientific), Pro-Ject reagent (Thermo Fisher Scientific), TRANSPASS™ P protein transfection reagent (New England Biolabs), CHARIOT™ protein delivery reagent (Active Motif), PROTEOJUICE™ protein transfection reagent (EMD Millipore), 293fectin, LIPOFECTAMINE™ 2000, LIPOFECTAMINE™ 3000 (Thermo Fisher Scientific), LIPOFECTAMINE™ (Thermo Fisher Scientific), LIPOFECTIN™ (Thermo Fisher Scientific), DMRIE-C, CELLFECTIN™ (Thermo Fisher Scientific), OLIGOFECTAMINE™ (Thermo Fisher Scientific), and LIPOFECTAMINE™ (Thermo Fisher Scientific). Scientific), LIPOFECTACE(TM), FUGENE(TM)(Roche, Basel, Switzerland), FUGENE(TM) HD(Roche), TRANSFECTAM(TM)(Transfectam, Promega, Madison, Wis.), TFX-10(TM)(Pr omega), TFX-20(TM) (Promega), TFX-50(TM) (Promega), TRANSFECTIN(TM) (BioRad, Hercules, Calif.), SILENTFECT(TM) (Bio-Rad), Effectene(TM) (Qiagen, Valencia, Calif.).), DC-chol (Avanti Polar Lipids), GENEPORTER™ (Gene Therapy Systems, San Diego, Calif.), DHARMAFECT 1™ (Dharmacon, Lafayette, Colo.), DHARMAFECT 2™ (Dharmacon), DHARMAFECT 3™ (Dharmacon), DHARMAFECT 4™ (Dharmacon), ESCORT™ III (Sigma, St. Louis, Mo.), and ESCORT™ IV (Sigma Chemical Co.). Nucleic acids, such as NT DNA, can also be delivered to cells via microfluidic methods, such as those known to those of skill in the art.
[0186] Non-viral methods for in vivo or ex vivo delivery of nucleic acids include electroporation, lipofection (see U.S. Pat. Nos. 5,049,386, 4,946,787, and commercially available reagents such as Transfectam™ and Lipofectin™), microinjection, biolistics, LNPs, virosomes, liposomes (see, e.g., Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Ther. 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Res.52:4817-4820(1992); U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787), immunoliposomes, polycation or lipid:nucleic acid conjugates, naked DNA, and drug-enhanced uptake of DNA. For example, sonoporation using the Sonitron 2000 system (Rich-Mar) can also be used for the delivery of nucleic acids.
[0187] For example, NTNDAs can be formulated into lipid nanoparticles (LNPs), lipidoids, liposomes, lipoplexes, or core-shell nanoparticles. Delivery reagents such as liposomes, nanocapsules, microparticles, microspheres, lipid nanoparticles, and vesicles can be used for the introduction of the compositions of the present disclosure into suitable host cells. Specifically, nucleic acids can be formulated for delivery either encapsulated in lipid particles, liposomes, vesicles, nanospheres, nanoparticles, gold particles, and the like. Such formulations may be preferred for the introduction of pharmaceutically acceptable formulations of the nucleic acids disclosed herein.
[0188] The NTDNA described herein can be delivered in vitro or in vivo using various delivery methods known in the art or modifications thereof. For example, in some embodiments, the NTDNA is delivered by creating transient penetrations in the cell membrane using mechanical, electrical, ultrasonic, hydrodynamic, or laser-based energy to facilitate DNA entry into the target cell. For example, the NTDNA can be delivered by squeezing the cell through a size-restricted channel or by transiently disrupting the cell membrane by other means known in the art. In some cases, the NTDNA is directly injected as naked DNA alone into skin, thyroid, cardiac, skeletal muscle, or liver cells.
[0189] In some cases, NTDNA is delivered by gene gun: gold or tungsten spherical particles (1-3 μm diameter) coated with NTDNA are accelerated to high velocities by pressurized gas and can penetrate into target tissue cells.
[0190] In some embodiments, electroporation is used to deliver NT DNA to target cells. Electroporation involves the insertion of an electrode pair into the tissue, causing temporary destabilization of the cell membrane of the target cell tissue, allowing DNA molecules in the medium surrounding the destabilized membrane to penetrate into the cytoplasm and nucleoplasm of the cell. Electroporation has been used in vivo in many types of tissue, such as skin, lung, and muscle.
[0191] In some cases, NTDNA is delivered by hydrodynamic injection, a simple and highly efficient method for the direct intracellular delivery of any water-soluble compound and particle to skeletal muscle in the viscera and entire limbs.
[0192] In some cases, NTDNA is delivered by ultrasound, creating nanoscopic pores in the membrane to facilitate intracellular delivery of DNA particles to organ or tumor cells, so the size and concentration of the plasmid DNA play a major role in the efficiency of this system. In other cases, NTDNA is delivered by magnetofection, using a magnetic field to concentrate the nucleic acid-containing particles in the target cells.
[0193] In some cases, chemical delivery systems can be used, for example, by using nanomer complexes, including the compression of negatively charged nucleic acids with cationic liposomes / micelles or polycationic nanomer particles belonging to cationic polymers. Cationic lipids used for delivery methods include, but are not limited to, monovalent cationic lipids, polyvalent cationic lipids, guanidine-containing compounds, cholesterol-derivative compounds, cationic polymers (e.g., poly(ethyleneimine), poly-L-lysine, protamine, other cationic polymers), and lipid-polymer hybrids.
[0194] NTDNA can also be administered directly to an organism for in vivo cell transduction, for example, as described herein. Administration is by any of the routes typically used to introduce molecules into ultimate contact with blood or tissue cells, including, but not limited to, injection, infusion, topical application, and electroporation. Suitable methods for administering such nucleic acids are available and well known to those skilled in the art, and while more than one route can be used to administer a particular composition, certain routes can often provide a more immediate and effective response than another route.
[0195] For example, compositions comprising the NT DNA described herein and a cytosolic delivery vehicle are specifically contemplated herein. In some embodiments, the NT DNA is formulated in a lipid delivery system, e.g., the LNP described herein. In some embodiments, such compositions are administered by any route desired by a skilled practitioner. The compositions can be administered to a subject by different routes, including orally, parenterally, sublingually, transdermally, rectally, transmucosally, topically, via inhalation, via buccal administration, intrapleurally, intravenously, intraarterially, intraperitoneally, subcutaneously, intramuscularly, intranasally, intrathecally, and intraarticularly, or combinations thereof. For veterinary use, the compositions can be administered in a suitably tolerated formulation in accordance with standard veterinary practice. A veterinarian can readily determine the dosing regimen and route of administration that is most appropriate for a particular animal. The compositions can be administered by conventional syringes, needleless injection devices, "microparticle gene guns," or other physical methods such as electroporation ("EP"), hydrodynamic methods, or ultrasound.
[0196] In some cases, NTDNA is delivered by hydrodynamic injection, a simple and highly efficient method for the direct intracellular delivery of any water-soluble compound and particle to skeletal muscle in the viscera and entire limbs.
[0197] In some cases, NTDNA is delivered by ultrasound by creating nanoscopic pores in the membrane, facilitating intracellular delivery of the DNA particles to cells of internal organs or tumors; the size and concentration of NTDNA play a large role in the efficiency of this system. In some cases, NTDNA described herein is delivered by magnetofection, using a magnetic field to concentrate the nucleic acid-containing particles in the target cells.
[0198] In some cases, chemical delivery systems can be used, for example, by using nanomer complexes, including the compression of negatively charged nucleic acids with cationic liposomes / micelles or polycationic nanomer particles belonging to cationic polymers. Cationic lipids used for delivery methods include, but are not limited to, monovalent cationic lipids, polyvalent cationic lipids, guanidine-containing compounds, cholesterol-derivative compounds, cationic polymers (e.g., poly(ethyleneimine), poly-L-lysine, protamine, other cationic polymers), and lipid-polymer hybrids.
[0199] Microparticles / Nanoparticles. In some embodiments, the NTDNA described herein is delivered by nanoparticles. One example of a nanoparticle that finds use in delivering the subject NTDNA is a lipid nanoparticle (LNP). Generally, the LNPs of the present disclosure may be composed of nucleic acid (NTDNA) molecules, one or more ionized or cationic lipids (or salts thereof), one or more nonionic or neutral lipids (e.g., phospholipids), a molecule that prevents aggregation (e.g., PEG or PEG-lipid conjugate), and optionally a sterol (e.g., cholesterol). For example, the lipid nanoparticles can include an ionizable amino lipid (e.g., heptatriaconta-6,9,28,31-tetraen-19-yl 4-(dimethylamino)butanoate, DLin-MC3-DMA, phosphatidylcholine (1,2-distearoyl-sn-glycero-3-phosphocholine, DSPC), cholesterol, and a coating lipid (polyethylene glycol-dimyristolglycerol, PEG-DMG), as disclosed, for example, by Tam et al. (2013). Advances in Lipid Nanoparticles for siRNA delivery. Pharmaceuticals 5(3):498-507.
[0200] In some embodiments, the lipid nanoparticles have an average diameter of about 10 to about 1000 nm. In some embodiments, the lipid nanoparticles have a diameter that is less than 300 nm. In some embodiments, the lipid nanoparticles have a diameter of about 10 to about 300 nm. In some embodiments, the lipid nanoparticles have a diameter that is less than 200 nm. In some embodiments, the lipid nanoparticles have a diameter of about 25 to about 200 nm. In some embodiments, the lipid nanoparticle preparation (e.g., a composition comprising a plurality of lipid nanoparticles) has a size distribution, with an average size (e.g., diameter) of about 70 nm to about 200 nm, more typically, the average size is about 100 nm or less.
[0201] In some embodiments, the present disclosure provides lipid nanoparticles comprising the NTDNA described herein and ionized lipids.Ionized lipids are typically used to condense nucleic acid cargoes, such as NTDNA, at low pH and to drive membrane association and fusion.Generally, ionized lipids are lipids that are positively charged or contain at least one amino group that is protonated under acidic conditions, for example, at a pH of 6.5 or less.Ionized lipids are also referred to herein as cationic lipids. Exemplary ionizable lipids are those described in International PCT Patent Publication Nos. 2015 / 095340, 2015 / 199952, 2018 / 011633, 2017 / 049245, 2015 / 061467, 2012 / 040184, 2012 / 000104, 2015 / 074085, 2016 / 081029, 2017 / 004143, 2017 / 075531, 2017 / 081046, 2017 / 081052, 2017 / 081062, 2017 / 081072, 2017 / 081082, 2017 / 081092, 2017 / 081094, 2017 / 081096, 2017 / 081098, 2017 / 081099, 2017 / 081099, 2017 / 081091, 2017 / 081092, 2017 / 081094, 2017 / 081096, 2017 / 081097, 2017 / 081098, 2017 / 081099, 2017 / 081099, 2017 / 081099, 2017 / 081091, 2017 / 081092, 2017 / 081093, 2017 / 081094, 2017 / 081095, 2017 / 081096, 2017 / 081 117528, 2011 / 022460, 2013 / 148541, 2013 / 116126, 2011 / 153120, 2012 / 044638, 2012 / 054365 No. 2011 / 090965, No. 2013 / 016058, No. 2012 / 162210, No. 2008 / 042973, No. 2010 / 129709, No. 2010 / 144740, No. 201 2 / 099755, 2013 / 049328, 2013 / 086322, 2013 / 086373, 2011 / 071860, 2009 / 132131, 2010 / 0485 No. 36, No. 2010 / 088537, No. 2010 / 054401, No. 2010 / 054406, No. 2010 / 054405, No. 2010 / 054384, No. 2012 / 016184, No. 2009 / 086558, 2010 / 042877, 2011 / 000106, 2011 / 000107, 2005 / 120152, 2011 / 141705, 2013 / 1 26803, 2006 / 007712, 2011 / 038160, 2005 / 121348, 2011 / 066651, 2009 / 127060, 2011 / 141704,Nos. 2006 / 069782, 2012 / 031043, 2013 / 006825, 2013 / 033563, 2013 / 089151, 2017 / 099823, 2015 / 095346, and 2013 / 086354, and U.S. Patent Publication Nos. 2016 / 0311759, 2015 / 0376115, 2016 / 0151284, 2017 / 0210697, 2015 / 0140070, 2013 / 0178541, 2013 / 03 No. 03587, No. 2015 / 0141678, No. 2015 / 0239926, No. 2016 / 0376224, No. 2 017 / 0119904, 2012 / 0149894, 2015 / 0057373, 2013 / 0090372 No. 2013 / 0274523, No. 2013 / 0274504, No. 2013 / 0274504, No. 2009 / 00 No. 23673, No. 2012 / 0128760, No. 2010 / 0324120, No. 2014 / 0200257, No. 20 15 / 0203446, 2018 / 0005363, 2014 / 0308304, 2013 / 0338210 No. 2012 / 0101148, No. 2012 / 0027796, No. 2012 / 0058144, No. 2013 / 03 No. 23269, No. 2011 / 0117125, No. 2011 / 0256175, No. 2012 / 0202871, No. 20 No. 11 / 0076335, No. 2006 / 0083780, No. 2013 / 0123338, No. 2015 / 0064242 , 2006 / 0051405, 2013 / 0065939, 2006 / 0008910, 2003 / 0022649, 2010 / 0130588, U52013 / 0116307, 2010 / 0062967, 2013 / 0202684, 2014 / 0141070, 2014 / 0255472, 2014 / 0039032, 2018 / 0028664, U52016 / 0317458, and 2013 / 0195920.
[0202] Various LNP formulations known in the art can be used to deliver the NTDNA described herein.For example, various LNP formulations and delivery methods using lipid nanoparticles are described in U.S. Patent Nos. 9,404,127, 9,006,417, 9,518,272, and U.S. Patent Application No. 63 / 415,229.Such particles can be prepared by high-energy mixing of aqueous NTDNA with ethanolic lipids at low pH, which protonates the ionized lipids and provides favorable energetics for NTDNA / lipid association and particle nucleation.Particles can be further stabilized through aqueous dilution and removal of organic solvent.Particles can be concentrated to a desired level.
[0203] Another example of a nanoparticle that finds use in delivering the subject NTDNA is a metal nanoparticle. In some embodiments, the NTDNA described herein is delivered by gold nanoparticles. Generally, nucleic acids can be covalently bound to gold nanoparticles, as described, for example, by Ding et al. (2014). Gold Nanoparticles for Nucleic Acid Delivery. Mol. Ther. 22(6); 1075-1083, or non-covalently bound to gold nanoparticles (e.g., bound via charge-charge interactions). In some embodiments, gold nanoparticle-nucleic acid conjugates are produced using methods described, for example, in U.S. Patent No. 6,812,334.
[0204] In some embodiments, the NTDNA described herein can be readily formulated in highly concentrated chitosan-nucleic acid polyplex compositions and orally administered in DNA enteric-coated pills as described in U.S. Patent Nos. 8,846,102, 9,404,088, and 9,850,323, each of which is incorporated herein in its entirety.
[0205] Exosomes. In some embodiments, the NTDNA described herein is delivered by packaging it into exosomes. Exosomes are small membrane vesicles of endocytic origin that are released into the extracellular environment following fusion of multivesicular bodies with the plasma membrane. Their surface consists of a lipid bilayer from the plasma membrane of the donor cell, and they contain cytosolic material from the cell that produced the exosome and exhibit membrane proteins from the parent cell on their surface. Exosomes are produced by various cell types, including epithelial cells, B and T lymphocytes, mast cells (MCs), and dendritic cells (DCs). In some embodiments, exosomes with diameters of 10 nm to 1 μm, 20 nm to 500 nm, 30 nm to 250 nm, or 50 nm to 100 nm are contemplated for use. Exosomes can be isolated for delivery to target cells either by using their donor cells or by introducing specific nucleic acids into them. Various approaches known in the art can be used to produce exosomes containing capsid-free AAV vectors of the invention.
[0206] Conjugates. In some embodiments, the NTDNAs disclosed herein are conjugated (e.g., covalently attached) to an agent that increases cellular uptake. An "agent that increases cellular uptake" is a molecule that facilitates transport of the nucleic acid across a lipid membrane. For example, the nucleic acid can be conjugated to a lipophilic compound (e.g., cholesterol, tocopherol, etc.), a cell-penetrating peptide (CPP) (e.g., penetratin, TAT, Syn1B, etc.), and a polyamine (e.g., spermine). Further examples of agents that increase cellular uptake are disclosed, for example, in Winkler (2013). Oligonucleotide conjugates for therapeutic applications. Ther. Deliv. 4(7); 791-809.
[0207] In some embodiments, the NTDNA disclosed herein is conjugated to a polymer (e.g., a polymer molecule) or a folate molecule (e.g., a folic acid molecule). Generally, delivery of polymer-conjugated nucleic acids is known in the art, for example, as described in WO 2000 / 34343 and WO 2008 / 022309. In some embodiments, the NTDNA disclosed herein is conjugated to a poly(amide) polymer, for example, as described by U.S. Pat. No. 8,987,377. In some embodiments, the nucleic acids described by the present disclosure are conjugated to a folic acid molecule, for example, as described in U.S. Pat. No. 8,507,455. In some embodiments, the NTDNA disclosed herein is conjugated to a carbohydrate, for example, as described in U.S. Pat. No. 8,450,467.
[0208] Nanocapsules. Alternatively, nanocapsule formulations of NTDNA can be used. Nanocapsules are generally capable of entrapping substances in a stable and reproducible manner. To avoid side effects due to intracellular polymer overload, such fine particles (approximately 0.1 μm in size) should be designed using polymers that can be degraded in vivo. Biodegradable polyalkyl-cyanoacrylate nanoparticles that meet these requirements are contemplated for use.
[0209] Liposomes. The NTDNA described herein can be loaded into liposomes for delivery to cells or target organs in a subject. Liposomes are vesicles with at least one lipid bilayer and an aqueous core. Liposomes are typically used as carriers for drug / therapeutic delivery in the context of formulation development. They act by fusing with cell membranes and rearranging their lipid structure to deliver drugs or active pharmaceutical ingredients (APIs). Liposome compositions for such delivery are composed of phospholipids, particularly compounds with phosphatidylcholine groups, although these compositions may also contain other lipids.
[0210] The formation and use of liposomes are generally known to those skilled in the art. Liposomes with improved serum stability and circulation half-lives have been developed (U.S. Patent No. 5,741,516). Furthermore, various methods for preparing liposomes and liposome-like preparations as potential drug carriers have been described (U.S. Patent Nos. 5,567,434, 5,552,157, 5,565,213, 5,738,868, and 5,795,587).
[0211] Additional Components In some embodiments, the delivery composition may include NTDNA and one or more additional components. In some cases, the one or more additional components may be, for example, one or more additional components of the gene editing system described above, where the one or more additional components mediate the genome integration of the NTDNA cargo nucleic acid. Thus, the NTDNA lipid nanoparticle may further include one or more of a guide RNA, an endonuclease, or a nucleic acid coding sequence, such as RNA (such as mRNA) or DNA.
[0212] One or more additional compounds may be a therapeutic agent. The therapeutic agent may be selected from any class suitable for therapeutic purposes. In other words, the therapeutic agent may be selected from any class suitable for therapeutic purposes. In other words, the therapeutic agent may be selected according to the desired therapeutic purpose and biological effect. For example, if the NTDNA in the LNP is useful for treating cancer, the additional compound can be an anti-cancer agent (e.g., a chemotherapeutic agent, a targeted cancer therapy (including, but not limited to, a small molecule, an antibody, or an antibody-drug conjugate)). In another example, if the LNP containing NTDNA is useful for treating an infectious disease, the additional compound can be an antimicrobial agent (e.g., an antibiotic or an antiviral compound). In yet another example, if the LNP containing NTDNA is useful for treating an immune disease or disorder, the additional compound can be a compound that modulates the immune response (e.g., an immunosuppressant, an immunostimulatory compound, or a compound that modulates one or more specific immune pathways). In some embodiments, different cocktails of different lipid nanoparticles containing different compounds, such as NTDNA encoding different proteins or different compounds, such as therapeutic agents, can be used in the compositions and methods of the invention. In some embodiments, the additional compound is an immunomodulatory agent. For example, the additional compound is an immunosuppressant. In some embodiments, the additional compound is an immunostimulatory agent.
[0213] composition Also provided are compositions that find use in practicing embodiments of the present invention.Compositions of the present invention include, for example, those having NTDNA as described, wherein NDTNA can be present in combination with one or more additional components, such as, for example, components of the gene editing system described above, such as, but not limited to, gRNA, endonuclease, or nucleic acid encoding it.
[0214] In some embodiments, the composition can comprise, for example, an NT DNA delivery vehicle, such as a liposome or lipid nanoparticle, as described above. Thus, in some embodiments, any of the components of the composition (e.g., the DNA endonuclease or the nucleic acid encoding it, the gRNA, and the NT DNA) can be formulated in a liposome or lipid nanoparticle. In some embodiments, one or more such components are associated with the liposome or lipid nanoparticle via a covalent or non-covalent bond. In some embodiments, any of the components can be contained separately or together in the liposome or lipid nanoparticle. Thus, in some embodiments, the DNA endonuclease or the nucleic acid encoding it, the gRNA, and the NT DNA (donor template) are each separately formulated in a liposome or lipid nanoparticle. In some embodiments, the DNA endonuclease is formulated together with the gRNA in a liposome or lipid nanoparticle. In some embodiments, the DNA endonuclease or the nucleic acid encoding it, the gRNA, and the donor template are formulated together in a liposome or lipid nanoparticle.
[0215] In some embodiments, the above compositions further comprise one or more additional reagents, wherein such additional reagents are selected from buffers, buffers for introducing polypeptides or polynucleotides into cells, wash buffers, control reagents, control vectors, control RNA polynucleotides, reagents for in vitro production of polypeptides from DNA, adaptors for sequencing, etc. The buffer may be a stabilization buffer, a reconstitution buffer, a dilution buffer, etc. In some embodiments, the composition also comprises one or more components that can be used to facilitate or enhance on-target binding or cleavage of DNA by the endonuclease, or to improve targeting specificity.
[0216] Also provided herein are pharmaceutical compositions comprising NTDNA encapsulated in a delivery vehicle and a pharmaceutically acceptable carrier or excipient. In some aspects, the present disclosure provides lipid nanoparticle formulations further comprising one or more pharmaceutical excipients. In some embodiments, the lipid nanoparticle formulation further comprises sucrose, Tris, trehalose, and / or glycine.
[0217] In some embodiments, any component of the composition is formulated with a pharmaceutically acceptable excipient, such as, for example, a carrier, solvent, stabilizer, adjuvant, or diluent, depending on the particular mode of administration and dosage form. In some embodiments, the guide RNA composition is generally formulated to achieve a physiologically compatible pH, ranging from about pH 3 to about pH 11, or from about pH 3 to about pH 7, depending on the formulation and route of administration. In some embodiments, the pH is adjusted to a range from about pH 5.0 to about pH 8. In some embodiments, the composition comprises a therapeutically effective amount of at least one compound described herein together with one or more pharmaceutically acceptable excipients. Optionally, the composition can comprise a combination of compounds described herein, or can include a second active ingredient useful in the treatment or prevention of bacterial growth (e.g., without limitation, an antibacterial or antimicrobial agent), or can include a combination of reagents of the present disclosure. In some embodiments, the gRNA is formulated with one or more other oligonucleotides, such as a nucleic acid encoding a DNA endonuclease and / or a donor template. Alternatively, the nucleic acid encoding the DNA endonuclease and the donor template are formulated separately or in combination with other oligonucleotides using the methods described above for gRNA formulation.
[0218] Suitable excipients may include, for example, carrier molecules including large, slowly metabolized macromolecules such as proteins, polysaccharides, polylactic acids, polyglycolic acids, polymeric amino acids, amino acid copolymers, and inactive virus particles. Other exemplary excipients include antioxidants (e.g., but not limited to, ascorbic acid), chelating agents (e.g., but not limited to, EDTA), carbohydrates (e.g., but not limited to, dextrin, hydroxyalkylcellulose, and hydroxyalkylmethylcellulose), stearic acid, liquids (e.g., but not limited to, oils, water, saline, glycerol, and ethanol), wetting or emulsifying agents, pH buffering substances, etc.
[0219] In some embodiments, the term "composition" refers to a therapeutic composition having therapeutic cells modified via a method of the invention, such as those described herein, for use in ex vivo therapeutic methods. In some embodiments, the therapeutic composition contains a physiologically acceptable carrier along with the cell composition, and optionally, at least one additional bioactive agent described herein dissolved or dispersed therein as an active ingredient. In some embodiments, the therapeutic composition is substantially non-immunogenic when administered to a mammalian or human patient for therapeutic purposes, unless so desired. Generally, the genetically modified therapeutic cells described herein are administered as a suspension with a pharmaceutically acceptable carrier. Those skilled in the art will recognize that pharmaceutically acceptable carriers used in cell compositions do not contain buffers, compounds, cryopreservatives, preservatives, or other agents in amounts that substantially interfere with the viability of the cells delivered to a subject. Cell-containing formulations can include, for example, an osmotic buffer that allows for maintaining cell membrane integrity, and optionally, nutrients to maintain cell viability or enhance engraftment upon administration. Such formulations and suspensions are known to those skilled in the art and / or can be adapted for use with progenitor cells as described herein using routine experimentation. In some embodiments, the cell composition can also be emulsified or presented as a liposomal composition, provided that the emulsification procedure does not adversely affect cell viability. The cells and any other active ingredients can be mixed with excipients that are pharmaceutically acceptable, compatible with the active ingredients, and in amounts suitable for use in the therapeutic methods described herein. Additional agents included in the cell composition can include pharmaceutically acceptable salts of the components therein. Pharmaceutically acceptable salts include, for example, acid addition salts (formed with the free amino groups of the polypeptide) formed with inorganic acids such as hydrochloric acid or phosphoric acid, or organic acids such as acetic acid, tartaric acid, mandelic acid, and the like.Salts formed with free carboxyl groups can also be derived from inorganic bases such as sodium, potassium, ammonium, calcium, or ferric hydroxide, and organic bases such as isopropylamine, trimethylamine, 2-ethylaminoethanol, histidine, procaine, and the like. Physiologically acceptable carriers are well known in the art. Exemplary liquid carriers are sterile aqueous solutions containing no materials in addition to the active ingredient and water, or containing a buffer such as sodium phosphate at a physiological pH value, physiological saline, or both, e.g., phosphate-buffered saline. Furthermore, aqueous carriers can contain two or more buffer salts, as well as salts such as sodium chloride and potassium chloride, dextrose, polyethylene glycol, and other solutes. Liquid compositions can also contain liquid phases in addition to and to the exclusion of water. Examples of such additional liquid phases are glycerin, vegetable oils such as cottonseed oil, and water-oil emulsions. The amount of active compound used in the cell composition that will be effective in treating a particular disorder or condition will depend on the nature of the disorder or condition and can be determined by standard clinical techniques.
[0220] kit Embodiments of the present disclosure also include kits. Some embodiments of the present disclosure provide kits, for example, including the NT DNA described above. The kits of the present disclosure may further include one or more additional components, such as an inducer (e.g., as described above), such as a gene editing system or a component thereof described above, etc. In some cases, the various components of a given kit may be combined into a single composition, such as a pharmaceutical composition, together with, for example, a cytosolic delivery vehicle, such as described above.
[0221] In some cases, the kit may include or contain an article of manufacture containing materials useful for treating the above-mentioned diseases. In some embodiments, the article of manufacture includes a container and a label. Suitable containers include, for example, bottles, vials, syringes, and test tubes. The container may be formed from a variety of materials, such as glass or plastic. In some embodiments, the container holds a composition effective for treating a disease described herein and may have a sterile access port. For example, the container may be an intravenous solution bag or vial having a stopper pierceable by a hypodermic injection needle. The active agent in the composition is a compound of the present invention. In some embodiments, a label on or associated with the container indicates that the composition is used to treat the selected disease. The article of manufacture may further include a second container containing a pharmaceutically acceptable buffer, such as phosphate-buffered saline, Ringer's solution, or dextrose solution. It may further include other materials desirable from a commercial and user standpoint, including other buffers, diluents, filters, needles, syringes, and package inserts containing instructions for use.
[0222] The components of the kit may be present in separate containers, or multiple components may be present in a single container. In addition to the components mentioned above, the subject kits may further include instructions for using the components of the kit, e.g., to practice the subject methods. The instructions are generally recorded on a suitable recording medium. For example, the instructions may be printed on a substrate such as, for example, paper or plastic. Thus, the instructions may be present in the kit as a package insert, on labeling of the container of the kit or its components (i.e., associated with the packaging or subpackaging), or the like. In other embodiments, the instructions are present as an electronic storage data file present on a suitable computer-readable storage medium, e.g., a CD-ROM, a diskette, a hard disk drive (HDD), a portable flash drive, or the like. In still other embodiments, the actual instructions are not present in the kit, but means for obtaining the instructions from a remote source, e.g., via the Internet, are provided. An example of this embodiment is a kit that includes a web address where the instructions can be viewed and / or from which the instructions can be downloaded. As with the instructions, this means that the means for obtaining the instructions is recorded on a suitable substrate.
[0223] usefulness The subject methods and compositions can be used in any application in which nuclear delivery of a cargo nucleic acid is desired, for example, as described above. Desired applications include both research and therapeutic applications. Desired applications include, but are not limited to, research, diagnostic, and therapeutic applications. In some cases, cargo nucleic acids that can be introduced into the nucleus via the methods of the present invention include those encoding research proteins, diagnostic proteins, and therapeutic proteins.
[0224] A research protein is a protein whose activity finds use in a research protocol. Thus, a research protein is a protein used in an experimental procedure. A research protein can be any protein with such utility, although in some cases, a research protein is a protein domain that is also provided in a research protocol by expressing it in cells from an encoding vector. Examples of specific types of research proteins include, but are not limited to, transcriptional regulators of inducible expression systems, members of signal production systems such as enzymes and their substrates, hormones, prohormones, proteases, enzyme activity regulators, perturbimers and peptide aptamers, antibodies, regulators of protein-protein interactions, genome-modifying proteins such as CRE recombinases, meganucleases, zinc finger nucleases, CRISPR / Cas-9 nucleases, TAL effector nucleases, and cell reprogramming proteins such as Oct3 / 4, Sox2, Klf4, c-Myc, Nanog, Lin-28, and the like.
[0225] A diagnostic protein is a protein whose activity finds use in diagnostic protocols. Thus, a diagnostic protein is a protein used in a diagnostic procedure. A diagnostic protein can be any protein that has such utility. Examples of specific types of diagnostic proteins include, but are not limited to, members of signal producing systems, such as enzymes and their substrates, labeled binding members, such as labeled antibodies and their binding fragments, peptide aptamers, and the like.
[0226] Proteins of interest further include therapeutic proteins. Therapeutic proteins are proteins that provide a therapeutic benefit to a patient, and include secreted proteins, transmembrane proteins, and intracellularly acting proteins. As will be understood by those skilled in the art, cargo nucleic acids encoding any protein that is associated with liver disease or that, upon secretion from the liver, finds use in treating another organ in the body can be delivered using the subject compositions and methods.
[0227] Target cells to which nucleic acids can be delivered in accordance with the present invention can vary widely. Target cells of interest include, but are not limited to, cell lines such as HeLa, HEK, CHO, and 293, mouse embryonic stem cells, human stem cells, mesenchymal stem cells, primary cells, tissue samples, and the like. Some non-limiting examples of mammalian cells include, but are not limited to, mouse cells, rat cells, hamster cells, rodent cells, and non-human primate cells. In some embodiments, the target cells are human cells. It should also be understood that the target cells can be of any cell type. For example, the target cells can be stem cells, which can include embryonic stem cells, induced pluripotent stem cells (iPS cells), fetal stem cells, umbilical cord blood stem cells, or adult stem cells (i.e., tissue-specific stem cells). In other cases, the target cells can be any differentiated cell type found in a subject. Cells of interest include both dividing and non-dividing cells. Examples of specific target cells of interest include, but are not limited to, hepatocytes, astrocytes, T lymphocytes, B lymphocytes, NK cells, skeletal muscle cells, cardiac muscle cells, neurons, astrocytes, oligodendrocytes, dendritic cells, skin cells, and the like.
[0228] In some cases, the intended use is a therapeutic use, for example, in the treatment of disease. For example, the compositions and methods of the present application can be used to deliver nucleic acid sequences to the nucleus of a cell to complement a genetic defect. As one non-limiting example, the compositions of the present application can be used in the treatment of a genetic defect that affects the function of liver cells, or in the treatment of a genetic defect elsewhere in the body that can be repaired by utilizing liver cells as biofactories to secrete defective proteins.
[0229] The following examples are offered by way of illustration and not by way of limitation. [Example]
[0230] The following examples are presented to provide those of ordinary skill in the art with a complete disclosure and description of how to make and use the present invention, and are not intended to limit the scope of what the inventors regard as their invention, nor are they intended to represent that the following experiments are all or the only experiments performed. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperatures, etc.), but some experimental error and deviation should be accounted for. Unless otherwise indicated, parts are parts by weight, molecular weight is weight average molecular weight, temperature is in degrees Celsius, and pressure is at or near atmospheric.
[0231] Materials and Methods General methods in molecular and cellular biochemistry are described in Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harvard Laboratory Press 2001); Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999); Protein Methods (Bollag et al., John Wiley & Sons 1996); Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999); Viral Vectors (Kaplift & Loewy eds., Academic Press 1995); Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997); and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons Reagents, cloning vectors, cells, and kits for the methods referred to or related to in this disclosure are available from commercial suppliers such as BioRad, Agilent Technologies, Thermo Fisher Scientific, Sigma-Aldrich, New England Biolabs (NEB), Takara Bio USA, Inc., and repositories such as Addgene, Inc., American Type Culture Collection (ATCC), for example.
[0232] DNA Construct Design. Plasmids were designed to test the effects of DTS elements on nuclear import and expression, as shown in Figure 2. Each DTS was constructed with pentameric repeats, with a 15-bp spacer between each TFBS. The DTS was added immediately 5' of the promoter, essentially within the putative enhancer region. Various promoters, including CBh (a variant of the CAG promoter) and TTR, were used in these constructs. Several reporter genes, including enhanced green fluorescent protein (eGFP), were used in these constructs. The polyadenylation (polyA) signal was the BGH-polyA sequence. For pooled screening, unique molecular identifiers (UMIs; also known as barcodes) were used to distinguish between DNA constructs within the pool. The UMI was added to the 3'UTR immediately downstream of the eGFP stop codon, immediately preceding the BGH-polyA sequence.
[0233] Library cloning. Pooled DTS libraries were designed as described above. To ensure the highest quality library production, oligo lengths for cloning were standardized to 300 bp. To accommodate this standardized length, the core sequence of each DTS was centered and embedded in a fragment of neutral random DNA sequence. Each oligo contained a unique Golden Gate restriction enzyme digestion site (Bsa1) for cloning double-stranded DNA into a screening vector. To convert the resulting oligo pool to double-stranded DNA for cloning, the oligos were PCR-amplified with primers specific to the flanking regions using Superfi polymerase for 16 cycles of PCR using the following program: 95°C, 16-fold (95°C 0:15, 58°C 0:15, 68°C 0:15), 68°C 2:00. Golden Gate vector assembly was then performed by mixing the following components (each with a unique Golden Gate restriction enzyme overhang) with Bsa1 restriction enzyme and T4 ligase in T4 ligase buffer: 1) double-stranded pooled DTS library DNA, 2) hTTR promoter, 3) eGFP gene, and 4) BGH-PolyA fragment containing the mechanically mixed N12 UMI. After 10 cycles of 37°C / 16°C corresponding to vector digestion and ligation, the ligated DNA was transformed into NEB-stable, chemically competent E. coli and plated on LB-agarose with chloramphenicol antibiotic (34 μg / mL). After overnight incubation at 37°C and counting a minimum of 40,000 colonies, bacterial colonies were scraped, aggregated, and DNA extracted using endotoxin-free maxiprep. Appropriate UMI / DTS barcode correlations (referred to here as "library accessions") were determined via excision of the hTTR / eGFP fragment, vector religation, PCR amplification of the corresponding proximal DTS and UMI regions, and sequencing (2 × 150 bp reads) on an Illumina Miseq. The DTS library plasmid pools were then transfected in vitro or formulated into LNPs for in vivo administration.
[0234] Cell culture. HepG2 cells were plated in 384-well plates and grown for 48 hours to achieve confluency. After removing the cell culture medium and replacing it with culture medium supplemented with aphidicolin (1 μM) using an EL406 automated plate washer and dispenser, the cells were transfected with 17.5 ng of DTS-containing plasmid using Lipofectamine 3000 without supplementation. GFP expression was monitored over 24 hours using a Biotek Cytation 5 cell imaging reader to assess DTS activity. Total GFP per well was calculated and averaged across four-well replicates.
[0235] LNP formulation. LNPs encapsulating nucleic acid payloads were prepared by mixing an organic solution of lipids with an aqueous solution of nucleic acid (e.g., DNA only, mRNA only, or a DNA / mRNA mixture) as described by Prud'homme et al. (J Pharm Sci 2018). Briefly, a lipid excipient mixture (ionizable lipids, helper lipids, cholesterol, PEG-lipids, and potentially other targeting moieties) is dissolved in an organic solvent. An aqueous solution of nucleic acid is prepared in a low pH buffer ranging from 3.0 to 4.0. The lipid mixture is then mixed with the aqueous nucleic acid solution at a flow rate ratio of 1:3 (V / V) using a commercially available mixer device. The resulting solution is immediately diluted with a buffer pH range of 5.0 to 6.5. The diluted LNPs are subjected to dialysis purification against a secondary buffer having a pH range of 7.0 to 8.0. The LNP solution was concentrated using a 100,000 MWCO Amicon Ultra centrifuge tube (Millipore Sigma) and then filtered through a 0.2 μm PES sterilizing grade filter. Particle size was determined by dynamic light scattering (Horiba nanoPartica SZ-100). Encapsulation efficiency was calculated using the Quant-it RiboGreen assay kit.
[0236] Quantification of mRNA abundance from pooled DTS screening. For in vitro studies, RNA was isolated by standard Trizol extraction methods and prepared using Zymo RNA extraction columns. For in vivo studies, liver tissue was homogenized in Trizol solution and then prepared using Zymo RNA extraction columns to obtain purified RNA. The purified RNA was transformed into cDNA using Maximus H-minus Reverse Transcriptase (Thermofisher) with a custom RT primer targeting the BGH-polyA sequence. The cDNA was PCR amplified using primers specific for GFP and UMI, which were then used to estimate transcript copy number and generate amplicons for library preparation and Illumina sequencing. Total transcript copy number was estimated at 200–30,000 transcripts per sample in each PCR reaction. The PCR-amplified cDNA was then used for library preparation (NEBNext Library Preparation for Illumina) and sequenced on an Illumina Miseq (2 × 150 bp reads). Counts for each UMI were converted to the corresponding DTS using the UMI / barcode accession, and the abundance change for each DTS was statistically analyzed using a custom R script. Log2 fold changes are calculated by comparing RNA levels within a pool to their expected abundance, derived from their abundance in the corresponding input DNA pool.
[0237] In vitro transcription of mRNA. mRNA was generated by in vitro transcription (IVT) using the commercially available mMESSAGE mMACHINE™ T7 transcription kit from ThermoFisher or the Hiscrib T7 mRNA kit from New England Biolabs. Incubation times, reagent concentrations, and reagent titrations were optimized to increase mRNA yield and purity. The IVT reaction uses a DNA template that can be a linearized plasmid, PCR product, or gene fragment. The mRNA sequence was codon-optimized for expression in human and mouse cells. To facilitate mRNA expression and stability, engineered UTRs were added to the template (Table 8). To stabilize the RNA and reduce its immunogenicity, a combination of chemically modified nucleotides (e.g., m6A, m6Am, 20meA, Ac4C, m5C, pseudoU, m1pseudo1, 5moU) and caps (e.g., Arca, CleanCap) was used. PolyA tails were either encoded in the plasmid template or added via a PolyA reaction kit commercially available from ThermoFisher, and mRNA was purified by LiCl precipitation or using the GeneJet RNA cleanup and concentration kit commercially available from ThermoFisher.
[0238] Co-transfection of DNA and mRNA into HepG2 cells. HepG2 cells were purchased from ATCC and grown in complete medium (EMEM + 10% fetal bovine serum, ATCC). For co-transfection experiments, cells were seeded at confluence onto collagen-coated 384-well plates and cultured for 48 hours in complete medium, followed by 48 hours in complete medium supplemented with 0.75 μM aphidicolin to prevent cell division. After the growth period, cells were transfected using Lipofectamine 3000 (Invitrogen) transfection reagent according to the manufacturer's instructions. RNA and DNA were independently complexed with Lipofectamine 3000 in Opti-MEM I (Gibco) for 15 minutes. The RNA and DNA complexes were then mixed and subsequently diluted 1:10 in complete medium supplemented with 0.75 μM aphidicolin. The medium was aspirated from the plated HepG2 cells and replaced with the transfection mixture. DNA was transfected at 17.5 ng to 4.38 ng per well. mRNA was transfected at 70 ng to 0.07 ng per well. Cells were live-imaged using a Cytation5 automated microscope (Agilent) mounted in a BioSpa8 automated incubator (Agilent) using a 4x objective with a GFP filter set and bright field at 3-hour intervals for 72 hours. Gen5IPrime (Agilent) image analysis software was used to quantify the number of GFP cells per well and GFP intensity.
[0239] Co-administration of DNA and mRNA to hepatocytes in vivo. DNA and mRNA are formulated independently into different LNP formulations, as described elsewhere, and the two formulations are mixed to create a complex LNP mixture. The complex LNP preparation is administered to mice. Additionally, DNA and mRNA are co-formulated into a single LNP formulation, and the co-formulation is administered to mice.
[0240] Quantification of EPO levels in serum. Blood is collected via retro-orbital bleeding into serum separator tubes and processed into serum. Serum samples can be stored at -80°C from collection until analysis. Serum levels of human EPO protein driven by expression from the DNA payload are quantified using the U-PLEX Human EPO Assay from MSD according to the manufacturer's instructions.
[0241] Quantification of FIX levels in plasma. Blood was collected via retro-orbital bleeding into K2EDTA tubes and processed to plasma. Plasma samples were stored at -80°C from collection to analysis. Plasma levels of human FIX after administration of LNPs were quantified using the U-Plex assay on the MSD platform. Briefly, a monoclonal mouse anti-human FIX antibody (Prolytix, clone AHIX-5041) was conjugated to biotin and used as a capture reagent on a streptavidin-coated plate. A polyclonal goat anti-human FIX antibody (Cedarlane, clone CL20040AP) was conjugated to sulfo-TAG and used as a detection reagent in standard settings for quantification of electrochemiluminescence (ECL) signals using a QuickPlex SQ 120MM instrument from MSD. Pooled normal human plasma (Affinity Biologicals, FRNCP0125), a pool of normal citrated human plasma collected from a minimum of 20 donors, was used to generate a standard curve and calculate % normal human FIX levels. The assay was confirmed to be specific for human FIX, did not cross-react with mouse FIX, and showed very low levels of background in untreated mouse plasma samples.
[0242] Example 1: Analysis of DTS activity using multiple in vitro experimental designs. Plasmid DNA (pDNA) containing various DTSs was transfected into primary human hepatocytes. GFP expression was quantified 24 hours after transfection by measuring GFP intensity using quantitative live-cell imaging. Cells were then harvested and the percentage of GFP-positive cells was measured by FACS. A strong correlation was observed between total GFP intensity and the percentage of GFP-positive cells (Figure 3A). DTSs identified as hits (open circles) increased both GFP intensity and the percentage of GFP-positive cells compared to the spacer-negative control (Figure 3A). The increase in the percentage of GFP-positive cells is significant because it indicates a nuclear translocation mechanism for DTS activity, rather than simple enhancer activity, which would have only increased GFP fluorescence intensity.
[0243] pDNA containing various DTSs was transfected into HepG2 cells grown with or without serum. GFP expression was quantified 24 hours after transfection by measuring GFP intensity using quantitative live-cell imaging. HepG2 cells grown in the presence of serum were actively dividing, resulting in DTS activity producing less than a 6-fold change compared to the spacer-negative control (Figure 3B). However, HepG2 cells that were serum-starved were not actively dividing, resulting in DTS activity producing up to a 25-fold increase in GFP intensity compared to the spacer-negative control (Figure 3B). DTSs identified as hits (open circles) produced increased GFP intensity in both experimental settings, with good correlation between conditions (Figure 3B). Importantly, non-dividing cells were observed to yield a greater dynamic range for assessing DTS function due to the nuclear membrane remaining intact because the cells do not undergo mitosis.
[0244] Example 2: Establishing benchmark DTS function in primary human hepatocytes. Primary human hepatocytes were transfected with pDNA containing a DTS with an NF-kB binding site, a DTS derived from the SV40 enhancer, or no DTS. GFP expression was quantified 24 hours after transfection by measuring GFP intensity using quantitative live-cell imaging. pDNA containing an NF-kB or SV40-derived DTS produced a robust increase in gene expression compared to pDNA without a DTS (Figure 4). pDNA with a DTS containing an NF-kB binding site drove a nearly 10-fold increase in GFP intensity compared to pDNA without a DTS (Figure 4).
[0245] Example 3: Arrayed screening of novel DTSs that drive improved gene expression in vitro. pDNAs containing various DTSs (Table 12) were individually transfected into growth-arrested HepG2 cells. Cell growth was inhibited to prevent nuclear envelope dissolution. GFP expression was quantified over a 24-hour period posttransfection by measuring GFP intensity using quantitative live-cell imaging. Using statistical analysis and hierarchical clustering methods, strong and weak DTSs were easily distinguished within the heatmap (Figure 5A). The top DTSs at 24 hours performed equally or better than NFKB, driving robust increases in gene expression compared to the spacer-negative control (Figure 5B). Unlike NFKB, which shows increased activity in response to inflammation, the top three novel DTSs (DTS.203, DTS.233, and DTS.276) contain TFBSs for constitutively expressed hepatocyte-specific transcription factors that are not responsive to inflammation. DTS.203: CREB1, PPARA, ONECUT1, HNF4A, PPARA. DTS.233:HNF1A, NR1I3, PPARA, HNF1A, PPARA. DTS.276:FOXA1, HNF4A, HNF1A, HNF4A, HNF1A.
[0246] Example 4: Establishing correlation between arrayed and pooled screening. For the arrayed screening, pDNAs containing various DTSs were individually transfected into growth-arrested HepG2 cells. For the pooled screening, pDNAs containing various DTSs were pooled together and transfected as a pool into growth-arrested HepG2 cells. Cell growth was inhibited to prevent nuclear envelope dissolution, which improves the accuracy of measuring DTS activity associated with nuclear translocation. For the arrayed screening, GFP protein expression was quantified 24 hours after transfection by measuring GFP intensity using quantitative live-cell imaging. For the pooled screening, mRNA abundance for each DTS variant was determined using RNA-seq by deconvoluting the read counts for each UMI. A strong correlation was observed between protein expression in the arrayed screening and mRNA expression in the pooled screening (Figure 6). Several DTSs, including NFKB and the novel DTS.233, showed robust increases in mRNA abundance and GFP intensity compared to the spacer-negative control (Figure 6). Thus, pooled screening can be used to rapidly and reliably identify potent DTSs.
[0247] Example 5: In vivo confirmation of DTS activity in mouse liver. pDNA containing various DTSs was pooled together and formulated into lipid nanoparticles (LNPs) containing ALC-0315 ionizable lipid, as described in the methods above. Wild-type female BALB / c mice (approximately 8-12 weeks old) were administered a single intravenous bolus injection into the tail vein at 5 mL / kg body weight. pDNA-LNPs were administered to groups of five mice at 0.66 mg / kg based on the weight of the DNA payload. At both 1 and 4 days post-dosing, mice were euthanized, and their livers were harvested and homogenized. After RNA extraction from liver homogenates, mRNA abundance for each DTS variant was determined using RNA-seq by deconvoluting the read counts for each UMI. At 1 day post-dosing, NFKB produced the largest and most statistically significant increase in mRNA abundance compared to the spacer-negative control (Figure 7A). At 1 day post-administration, the best novel DTSs (DTS.233 and DTS.276) produced an 8- to 10-fold increase in mRNA abundance, dramatically improving gene expression compared to the spacer-negative control (Figure 7A). At 4 days post-administration, the best DTSs continued to show increased mRNA abundance, although the magnitude of the increase was likely lower due to gene silencing driven by the immunostimulatory plasmid backbone (Figure 7B). It can also be observed that at 4 days post-administration, many DTSs produced reduced mRNA abundance, likely due to repressor-like activity (Figure 7B). Because the combinatorial design of the novel DTSs employed several TFBSs placed together, it was important to determine combinations that resulted in both increased and decreased gene expression. Positive and negative data can be used to train machine learning algorithms, enabling the extraction of key trends. We developed a gradient-boosted decision tree machine learning algorithm and applied it independently to each dataset generated in the DTS pool. We identified the same three transcription factors (HNF1A, PPARA, and HNF4A) as overexpressed in strong DTS and underexpressed in weak DTS (Figure 8).
[0248] Example 6: Use of inducers to improve DTS activity. pDNA containing various DTSs was individually transfected into growth-arrested HepG2 cells. At the same time, various concentrations of an activator were added to the cells. Dexamethasone, a known activator of the glucocorticoid receptor (GR, also known as NR3C1), was used as the activator. GFP expression was quantified 24 hours after transfection by measuring GFP intensity using quantitative live-cell imaging. An inducible DTS (iDTS-1) containing five copies of the GR response element (GRE) in pentameric repeats produced increased gene expression compared to the spacer-negative control as a function of increasing dexamethasone concentrations (Figure 9). Another construct of iDTS-1, containing two copies of the GRE as part of a combined pentameric repeat, also showed increased gene expression in response to dexamethasone as an activator (Figure 10). The iDTS-1-2 construct produced an increase in gene expression equivalent to a robust NFKB DTS and significantly higher than the spacer-negative control (Figure 10). [Table 12-1] [Table 12-2] [Table 12-3] [Table 12-4] [Table 12-5] [Table 12-6] [Table 12-7] [Table 12-8] [Table 12-9] [Table 12-10] [Table 12-11] [Table 12-12]
[0249] Example 7: In vitro co-transfection of NTF-expressing mRNA improves DNA nuclear import and gene expression. mRNAs were designed to express TetR fusion proteins with nuclear localization signals (NLSs) and, in some cases, nuclear export signals (NESs) located at the N-terminus, C-terminus, or both, as outlined in Table 13 (Figures 11A-C). [Table 13-1] [Table 13-2]
[0250] Growth-arrested HepG2 cells were cotransfected with these different TetR fusion proteins and the plasmid NTDNA containing the tetracycline response element (TRE) as a DTS (containing seven copies of the TetO sequence TCGAGTTTACTCCCTATCAGTGATAGAGAACG (SEQ ID NO: 444) with a four-nucleotide spacer tatg between each TRE) and a GFP expression cassette. Forty-eight hours after transfection, GFP expression was assessed by measuring GFP fluorescence intensity, and toxicity was assessed by counting rounded (dead) cells. C-terminal NLS conjugates 2.2, 2.6, and 2.7 produced the highest increase in GFP intensity compared to the DNA alone and DNA + mCherry mRNA negative controls (Figure 12).
[0251] In follow-up experiments, growth-arrested HepG2 cells were transfected with mRNAs expressing TetR fusion proteins in family 2.X (C-terminal NLS) and family 5.X (C-terminal NLS and NES). Plasmid NT DNA containing a DTS with a TetR-binding site, e.g., a tetracycline-responsive element (TRE) containing seven copies of the TetO sequence, was delivered into growth-arrested HepG2 cells along with mRNA expressing the TetR-NLS fusion protein. GFP expression was assessed by measuring GFP fluorescence intensity for 21 hours after transfection. Cotransfecting TetR mRNA with TetO DNA significantly increased GFP expression compared to the negative DNA control for both dbDNA containing the TetO sequence (Figure 13A) and pDNA containing the TetO sequence (Figure 13B-C). Several TetR mRNAs produced significant increases in gene expression (including 2.2, 2.7, 2.11, 5.4, and 5.6), with the magnitude of gene expression increasing with increasing mRNA amounts (Figures 13A and 13B). mRNA produced with the IVT kit from NEB (2.2 and 2.7 labeled "NEB") was superior to mRNA produced with the IVT kit from ThermoFisher (all others). Compared to TetO DNA alone, cotransfection with TetR mRNA 2.2 and 2.7 increased gene expression by approximately 10- to 40-fold at 21 hours (Figure 13C). Over the entire 21-hour time course, TetO DNA cotransfected with TetR mRNA drove faster and higher GFP expression than DNA alone (Figures 13D and 13E).
[0252] The dose response and time course of gene expression from the two best TetR-NLS fusion proteins (2.2 and 2.7) were further evaluated. Plasmid NTDNA containing a DTS with a TetR binding site, e.g., a tetracycline response element (TRE) containing seven copies of the TetO sequence, was delivered into growth-arrested HepG2 cells along with mRNA expressing the TetR-NLS fusion protein. GFP expression was assessed by measuring GFP fluorescence intensity for 24 hours after transfection. dbDNA containing the TetO sequence cotransfected with TetR mRNA significantly increased GFP expression compared to the negative control of DNA alone (Figure 14A). The increase in gene expression was dependent on the TetO sequence, as dbDNA lacking the TetO sequence did not show increased gene expression when cotransfected with TetR mRNA (Figure 14B). Furthermore, the increase in gene expression also depends on the NLS fused to TetR, as Tet_0.0 mRNA, which lacks an NLS, did not improve gene expression (Figure 14A). The magnitude of gene expression increased with increasing amounts of mRNA across a wide range of RNA levels (Figure 14A). Over the entire 24-hour time course, TetO dbDNA cotransfected with TetR mRNA drives faster and higher GFP expression than DNA alone (Figure 14C). Furthermore, the level of gene expression achieved with TetO dbDNA + TetR mRNA was significantly at least 10-fold better than all negative controls, including plasmid DNA containing the NFKB DTS (Figure 14D). The fold change in GFP intensity increase with TetO dbDNA + TetR mRNA versus dbDNA alone varies over time. For high RNA inputs of Tet_2.2, as the signal from DNA alone increases, the fold change decreases from 60-fold at 9 hours (Figure 14E), to 15-fold at 18 hours (Figure 13F), to 10-fold at 24 hours (Figure 14G) (likely due to a low level of cell division allowing DNA access to the nuclei in a small fraction of cells).The DNA backbone architecture was not found to have an effect; both dbDNA and pDNA containing the TetO sequence showed similar increases in GFP expression when combined with TetR mRNA (Figures 14H and 14I). Finally, DNA + mRNA cotransfection was very well tolerated at all DNA:mRNA ratios, resulting in very low rounded cell counts (Figures 14J and 14K).
[0253] Co-delivery of NTF-expressing mRNA and NT DNA results in nuclear translocation and gene expression in vivo. mRNA encoding the TetR protein described above is formulated into LNPs to produce mRNA / LNP formulations. DNA (plasmid, doggybone, or nanoplasmid) containing a tetracycline response element (TRE) as a DTS and a reporter expression cassette encoding GFP or factor IX is formulated into LNPs to produce DNA / LNP formulations. The mRNA / LNP and DNA / LNP formulations are then mixed to produce mRNA / DNA / LNP preparations. Three cohorts of adult mice (cohort 1: mRNA / LNP; cohort 2: DNA / LNP; cohort 3: mRNA / DNA / LNP mixture) are injected at three doses (0.1 mg / kg, 0.3 mg / kg, and 1 mg / kg). Serum levels of factor IX are measured at 4 hours, 3 days, and 7 days. GFP is detected by postmortem histology. No Factor IX or GFP reporter is detected in samples from Cohort 1. Very low levels of Factor IX reporter are detected in very few cells with GFP in serum Cohort 2. A 2-5 fold increase in Factor IX reporter and GFP+ cells is observed in samples from Cohort 3 relative to Cohort 2, demonstrating a dose response.
[0254] Co-administration of NTF-expressing mRNA and NT DNA as a single LNP formulation improves DNA nuclear translocation and gene expression. mRNA encoding the TetR protein is co-formulated with DNA (plasmid, doggybone, or nanoplasmid) containing a tetracycline response element (TRE) as a DTS and a reporter expression cassette encoding GFP or factor IX to create an mRNA / DNA / LNP formulation. Three cohorts of adult mice (cohort 1: mRNA / LNP; cohort 2: DNA / LNP; cohort 3: mRNA / DNA / LNP formulation) are injected at three doses (0.1 mg / kg, 0.3 mg / kg, and 1 mg / kg). Serum levels of factor IX are measured at 4 hours, 3 days, and 7 days. GFP is detected by postmortem histology. No factor IX or GFP reporter is detected in samples from cohort 1. Very low levels of Factor IX reporter, along with GFP, are detected in very few cells in serum Cohort 2. A 20-fold increase in Factor IX reporter and GFP+ cells relative to Cohort 2 is observed in samples from Cohort 3, a 5-10-fold increase in efficacy relative to the combined preparation.
[0255] Example 8: Evaluation of different NTFs for their ability to facilitate DNA nuclear import and gene expression in vitro. Co-transfection of DNA with mRNA expressing several different NTFs improves DNA nuclear import and gene expression. In this study, the ability of several NTFs to promote DNA nuclear import in vitro was evaluated. To this end, DNA containing a GFP expression cassette was engineered to contain a DTS containing seven repeats of the DNA-binding domains for TetR, Gal4, Arc, Mnt, PurR, or Bac434. mRNA containing sequences encoding the TetR, Gal4, Arc, Mnt, PurR, and Bac434 proteins fused to a nuclear localization signal (NLS) (containing a DNA-binding domain capable of recognizing and binding to the aforementioned DNA-binding domains) was synthesized.
[0256] Growth-arrested HepG2 cells and primary human hepatocytes (PHH) were cotransfected with mRNAs encoding NTF proteins along with DNA containing the cognate DTS and a GFP expression cassette. Each NTF mRNA was paired with its cognate DNA in the cotransfection. For example, TetR mRNA was paired with DNA containing a DTS with seven repeats of the TetR TFBS, Gal4 mRNA was paired with DNA containing a DTS with seven repeats of the Gal4 TFBS, Arc mRNA was paired with DNA containing a DTS with seven repeats of the Arc TFBS, Mnt mRNA was paired with DNA containing a DTS with seven repeats of the Mnt TFBS, PurR mRNA was paired with DNA containing a DTS with seven repeats of the PurR TFBS, and Bac434 mRNA was paired with DNA containing a DTS with seven repeats of the Bac434 TFBS. Eighteen hours after transfection, GFP expression was assessed by measuring GFP fluorescence intensity. All of the NTF proteins increased nuclear translocation and gene expression when paired with their cognate DNA in growth-arrested HepG2 cells (Figure 15A) and PHH cells (Figure 15B). The magnitude of the improvement in gene expression increased with increasing amounts of mRNA co-transfection (Figures 15A and 15B). Depending on the NTF and DNA used, approximately 10- to 40-fold increases in gene expression were observed with the maximum amount of mRNA co-transfection (e.g., 17.5 ng of RNA) compared to DNA alone (e.g., 0 ng of RNA) in both cell types (Figures 15A and 15B).
[0257] Example 9: Investigation of the effect of including multiple TFBS repeats in the DTS on the efficacy of DNA transfer. In this study, DNA was engineered to contain DTSs with different numbers of TFBS repeats for TetR or Gal4, and the effects of multiple sequences were assessed in vitro by cotransfection with mRNAs encoding TetR or Gal4 NTFs. Growth-arrested HepG2 cells were cotransfected with mRNAs encoding NTF proteins along with DNAs containing the cognate DTSs and a GFP expression cassette. Each NTF mRNA was paired with its cognate DNA in the cotransfection, where the cognate DNA contained DTSs with various numbers of cognate TFBS repeats. For example, TetR mRNA was paired with DNA containing DTSs with 2, 3, 4, 5, 6, or 7 repeats of the TetR TFBS (e.g., TetO). 24 hours after transfection, GFP expression was assessed by measuring GFP fluorescence intensity.
[0258] Cotransfection of TetR mRNA with DNA containing a DTS with two repeats of TetO did not substantially increase gene expression relative to DNA alone, whereas cotransfection of TetR mRNA with DNA containing a DTS with three, four, five, six, or seven repeats of TetO substantially increased gene expression compared to DNA alone (Figure 16A). The increase in gene expression appeared to increase with the number of TetO repeats in the DTS, with seven repeats producing a more than 10-fold increase in GFP expression at the highest amount of RNA (e.g., 17.5 ng of RNA) compared to DNA alone (Figure 16A).
[0259] Gal4 mRNA was paired with DNA containing a DTS with 1, 2, 3, 4, 5, 7, or 10 repeats of the Gal4 TFBS (e.g., UAS). 12 hours after transfection, GFP expression was assessed by measuring GFP fluorescence intensity. Co-transfection of Gal4 mRNA with DNA containing a DTS with 1 or 2 repeats of the UAS did not substantially increase gene expression relative to DNA alone, whereas co-transfection of Gal4 mRNA with DNA containing a DTS with 3, 4, 5, 7, or 10 repeats of the UAS substantially increased gene expression compared to DNA alone (Figure 16B). The increase in gene expression appeared to increase with the number of UAS repeats in the DTS up to 5 repeats, but the increase in gene expression appeared to be saturable, as 5, 7, and 10 repeats showed similar performance (Figure 16B). Cotransfection of Gal4 mRNA and DNA containing 5, 7, or 10 repeats of the DTS produced a more than 10-fold increase in GFP expression in HepG2 cells with the highest amount of RNA (e.g., 17.5 ng of RNA) compared with DNA alone (e.g., 0 ng of RNA) (Figure 16B).
[0260] This study shows that increasing the number of repeats improves the efficiency of nuclear import, although the effect may be saturable.
[0261] Example 10: Evaluation of different strategies for in vivo administration of DNA and mRNA. Several different strategies for administering DNA encoding an EPO transgene and mRNA encoding a TetR translocator were investigated. First, control LNPs were formulated with EPO-TetO DNA alone, where the DNA contained a DTS with 7 TetO repeats. The same EPO-TetO DNA was also co-formulated into LNPs with mRNA expressing TetR protein. The same EPO-TetO DNA and mRNA expressing TetR protein were also independently formulated into separate LNPs and then mixed together in the same dosing solution for co-dosing. Finally, in a separate group, LNPs carrying TetR mRNA were pre-administered 1 hour before LNPs carrying EPO-TetO DNA. Adult BALB / c mice were administered LNPs via tail vein injection at a dose volume of 5 mL / kg. Serum levels of human EPO protein were measured 3 days after administration. In all groups, the total amount of DNA administered was 0.5 mg / kg (mpk), and the total amount of mRNA administered was 0.5 mg / kg (mpk) except for the DNA-only group, in which no mRNA was administered.
[0262] When TetR mRNA was co-formulated into LNPs with EPO-TetO DNA or separately formulated into LNPs, mixed, and co-administered, TetR protein expression from mRNA increased EPO expression approximately 5-fold compared to LNPs with DNA alone (Figure 17A). When TetR mRNA-containing LNPs were pre-administered 1 hour before EPO-TetO DNA-containing LNPs, TetR protein expression from mRNA increased EPO expression compared to DNA alone, but not to quite the same extent as observed with co-formulation or co-administration (Figure 17A). All three strategies for incorporating TetR mRNA successfully increased TetO DNA expression.
[0263] Example 11: Evaluation of different NTFs for their ability to facilitate DNA nuclear import and gene expression. LNPs were co-formulated with DNA and mRNA expressing several different NTFs to evaluate the ability of these other NTFs to improve DNA nuclear import and gene expression. Adult BALB / c mice were administered LNPs formulated with either DNA alone or co-formulated with DNA and mRNA via tail vein injection at a dose volume of 5 mL / kg. LNPs containing DNA alone were administered at 0.5 mg / kg DNA, and co-formulated LNPs were administered at 0.5 mg / kg DNA and 1.5 mg / kg mRNA (3:1 RNA:DNA weight ratio). In this study, each DNA payload contained an expression cassette for human factor IX (FIX). Plasma levels of human FIX protein were measured up to 28 days after administration.
[0264] Administration of LNPs co-formulated with TetR mRNA and DNA containing a TetO DTS increased FIX levels approximately 5-fold compared to LNPs containing TetO DNA alone. Interestingly, "v2" TetR mRNA produced a faster and more sustained increase in gene expression compared to "v1" TetR mRNA (Figures 18A and 18B). Administration of LNPs co-formulated with Gal4 mRNA and DNA containing a DTS containing a UAS sequence increased FIX levels approximately 5-fold compared to LNPs containing UAS DNA alone (Figure 18C). Administration of LNPs co-formulated with Arc mRNA and DNA containing a DTS containing an Arc binding sequence increased FIX levels approximately 3-fold compared to LNPs containing Arc DNA alone (Figure 18D). Administration of LNPs co-formulated with Mnt mRNA and DNA containing a DTS containing an Mnt binding sequence increased FIX levels approximately 3-fold compared to LNPs containing Mnt DNA alone (Figure 18E). Administration of LNPs co-formulated with Bac434 mRNA and DNA bearing a DTS containing a Bac434-binding sequence increased FIX levels approximately two-fold compared to LNPs containing Bac434 DNA alone (Figure 18F). Similar to in vitro observations, this data indicates that multiple families of NTFs can improve nuclear translocation and gene expression in mice.
[0265] In at least some of the above-described embodiments, one or more elements used in one embodiment may be used interchangeably in another embodiment, provided that such substitution is technically feasible. Those skilled in the art will appreciate that various other omissions, additions, and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and variations are intended to fall within the scope of the subject matter defined by the appended claims.
[0266] It will be understood by those skilled in the art that the terms used herein generally, and in the appended claims in particular (e.g., the body of the appended claims), are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including, but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “including, but not limited to,” etc.). It will be further understood by those skilled in the art that where a specific number of introduced claim recitations is intended, such intention will be explicitly recited in the claim; in the absence of such recitation, such intention does not exist. For example, as an aid to understanding, the appended claims below may include the use of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be interpreted as meaning that the introduction of a claim recitation with the indefinite article “a” or “an” limits a particular claim that includes such introduced claim recitation to embodiments that include only one such recitation. The same claim includes the introductory phrase "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" means "at least one" or "one or more"); the same applies to the use of definite articles used to introduce claim recitations. Additionally, even if a specific number of recitations of an introduced claim are explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., the bare recitation of "two recitations" without other modifiers means at least two recitations, or more than two recitations).Furthermore, when a convention similar to "at least one of A, B, and C, etc." is used, such a configuration is generally intended in the sense that one of ordinary skill in the art would understand the convention (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having A alone, B alone, C alone, both A and B, both A and C, both B and C, and / or both A, B, and C, etc.). When a convention similar to "at least one of A, B, or C, etc." is used, such a configuration is generally intended in the sense that one of ordinary skill in the art would understand the convention (e.g., "a system having at least one of A, B, or C" includes, but is not limited to, systems having A only, B only, C only, both A and B, both A and C, both B and C, and / or both A, B, and C, etc.). It will be further understood by those skilled in the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."
[0267] In addition, when features or aspects of the disclosure are described in terms of a Markush group, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.
[0268] As will be understood by those skilled in the art, for any and all purposes, e.g., in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of those subranges. Any listed range can be readily recognized as fully indicating and allowing the same range to be broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third, upper third, etc. As will also be understood by those skilled in the art, all terms such as "up to," "at least," "greater than," "less than," etc., refer to ranges that are inclusive of the recited numbers and that can subsequently be broken down into subranges as described above. Finally, as will be understood by those skilled in the art, a range includes each individual member. Thus, for example, a group having 1 to 3 items refers to a group having 1, 2, or 3 items. Similarly, a group having 1 to 5 items refers to a group having 1, 2, 3, 4, or 5 items, etc.
[0269] Although the foregoing invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it will be readily apparent to those skilled in the art in light of the teachings of the invention that certain changes and modifications can be made thereto without departing from the spirit or scope of the appended claims.
[0270] Accordingly, the foregoing merely illustrates the principles of the present invention. It will be appreciated that those skilled in the art will be able to devise various arrangements, not explicitly described or shown herein, which embody the principles of the present invention and are within its spirit and scope. Furthermore, all examples and conditional language set forth herein are intended primarily to aid the reader in understanding the principles of the present invention and the concepts the inventors contributed to furthering the art, and should not be construed as being limited to such specifically described examples and conditions. Furthermore, all statements herein describing principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, such equivalents are intended to include both currently known equivalents and future-developed equivalents, regardless of structure, i.e., any elements developed to perform the same function, regardless of structure. Furthermore, nothing disclosed herein is intended to be dedicated to the public, regardless of whether such disclosure is expressly recited in the claims.
[0271] Accordingly, the scope of the present invention is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of the present invention is embodied by the appended claims. In the claims, 35 U.S.C. 112(f) or 35 U.S.C. 112(6) is expressly defined to be invoked for a limitation in a claim only when the precise phrase "means for" or the precise phrase "step for" is recited at the beginning of that limitation; if such precise phrase is not used in the limitation in the claim, 35 U.S.C. 112(f) or 35 U.S.C. 112(6) is not invoked.
Claims
1. 1. A method for delivering a cargo nucleic acid sequence to the nucleus of a cell, said method comprising: The cells a) a DNA nuclear targeting sequence (DTS) comprising a targeting factor binding sequence (TFBS) for a nuclear targeting factor (NTF); b) contacting said DTS with nuclear-targeted deoxyribonucleic acid (NTDNA), comprising a cargo nucleic acid heterologous to said DTS; delivering said cargo nucleic acid sequence to the nucleus of said cell.
2. The method of claim 1 , wherein the cargo nucleic acid sequence comprises an expression cassette.
3. The method of claim 1 , wherein the cargo nucleic acid comprises a promoter that is heterologous to the DTS.
4. The method of claim 1 , wherein the cargo nucleic acid comprises a coding sequence that is heterologous to the DTS.
5. 10. A method according to any one of the preceding claims, wherein the cargo nucleic acid sequence is flanked by sequences that are homologous to genomic sequences of the cell.
6. 10. The method of any one of the preceding claims, wherein the DTS comprises two or more different targeting factor binding sequences (TFBS).
7. The method of any one of claims 1 to 6, wherein the DTS comprises two or more copies of the same TFBS.
8. The method of claim 6 or 7, wherein the DTS comprises 5 to 100 nucleotides between each TFBS.
9. 9. The method of any one of claims 1 to 8, wherein the DTS comprises one or more TFBS selected from Table 1 or Table 2.
10. 9. The method of claim 1, wherein the DTS comprises a TFBS for an NTF selected from the group consisting of HNF1A, PPARA, HNF4A, CEBPA, NR3C1, ONECUT1, TBP, NFkB, TetR protein, ZF-CCR5 protein, TALE protein, GAL4 protein, Arc protein, Mnt protein, PurR protein, Bac434 protein, GCN4 protein, LacR protein, I-SceI D44A protein, and Cas9 protein.
11. (a) when the TFBS is for HNF1A, the sequence is selected from GGTTAATAATTAAC and AGTATGGTTAATGATCTACAG; (b) when the TFBS is for PPARA, the sequence is selected from AACTAGGTCAAAGGTCA, CAAAACTAGGTCAAAGGTCA, and AACTAGGTCAAAGGTCAAAG; (c) when the TFBS is for HNF4A, the sequence is selected from TCGAGCGCTGGGCAAAGGTCACCTGC, TCGAGCGCAGGTCAAAGGTCACCTGC, and AGGTCAAAGTCCA; (d) when the TFBS is for CEBPA, the sequence is TGGTATGATTTTGTAATGGGGTAGGA; (e) when the TFBS is for NR3C1, the sequence is selected from AGAACAAAATGTTCT, AAGAACAAAATGTTCTT, AGAACATTTTGTACG, and AAGAACATTTTGTACGT; (f) when the TFBS is for ONECUT1, the sequence is GTCTGCTAAGTCAATAATCAGAAT; (g) when the TFBS is for TBP, the sequence is TATAAAA; (h) if the TFBS is for NFkB, the sequence is GGGACTTTCC; (i) when the TFBS is for a TetR protein, the sequence comprises TCCCTATCAGTGATAGAGA; (j) when the TFBS is for a ZF-CCR5 protein, the sequence comprises AAACTGCAAAG; (k) if the TFBS is for a TALE protein, the sequence comprises TTCATTACACCTGCAGCT, ATAAACCCCCTCCAA, or Tcgagtttactccctatcagtgatagagaacg; (l) When the TFBS is for a GAL4 protein, the sequence is CGG-N 11 - includes CCG, (m) when the TFBS is for an Arc protein, the sequence comprises RYRVTAGANNNNNNTCTABYRY; (n) when the TFBS is for an Mnt protein, the sequence comprises GGNCCACNGTGGNCC; (o) when the TFBS is for a PurR protein, the sequence comprises ACGCAAACGTTTTCGT; (p) when the TFBS is for a Bac434 protein, the sequence comprises ACAAGAAAGTTTGT, ACAAGATACATTGT, or ACAAGAAAACTGT; (q) if the TFBS is for a GCN4 protein, the sequence comprises TGACTC; (r) the sequence comprises TTGTTATCCGCTCACAA when the TFBS is for a LacR protein, and TAGGGATAACAGGGTAAT when the TFBS is for an I-SceI protein.
12. The DTS is (a) HNF1A, NR1I3, PPARA, HNF1A, and PPARA; (b) HNF4A, NR1I3, NR3C1, NR3C1, and HNF4A; (c) FOXA1, HNF4A, HNF1A, HNF4A, and HNF1A; (d) NR3C1, CEBPA, HNF1A, CEBPA, and PPARA; (e) ETS1, NR3C1, ELF5, CEBPA, and NR1I3; (f) NR1I3, HNF4A, CREB1, NR1I3, and ETS1; (g) HNF4A, PPARA, CEBPA, HNF1A, and NR3C1; (h) NFKB1, NFKB1, NFKB1, NFKB1, and NFKB1; (i) SREBF1, MLXIPL, HNF1A, CREB3L3, and ONECUT1; and (j) The method of claim 10, comprising a TFBS sequence for a combination of NTFs selected from the group consisting of CREB1, PPARA, ONECUT1, HNF4A, and PPARA.
13. The DTS is (a) GGTTAATAATTAACAGATTACTACTGATAACCTCTTCTCTGTGGGTGACCAGCGTCCTAAGATTACTACTGATAAACTAG GTCAAAGGTCAAAGATTACTACTGATAAGTATGGTTAATGATCTACAGAGATTACTACTGATACAAAACTAGGTCAAGGTCA, (b)TCGAGCGCTGGGCAAAGGTCACCTGCAGATTACTACTGATAGCCCCCAGGGCTGAGTGACAGAAAAACAGAGATTACTACTGATAAGAACAAAATGTTCTAGATTACTACTGATAAAGAACATTTTGTACGTAGATTACTACTGATAAGGTCAAAGTCCA、 (c)TGTTTACTTTAGATTACTACTGATATCGAGCGCAGGTCAAAGGTCACCTGCAGATTACTACTGATAGGTTAATAATTAACAGATTACTACTGATATCGAGCGCTGGGCAAAGGTCACCTGCAGATTACTACTGATAAGTATGGTTAATGATCTACAG、 (d)AAGAACAAAATGTTCTTAGATTACTACTGATATGGTATGATTTTGTAATGGGGTAGGAAGATTACTACTGATAAGTATGGTTAATGATCTACAGAGATTACTACTGATATGGTATGATTTTGTAATGGGGTAGGAAGATTACTACTGATAAACTAGGTCAAAGGTCAAAG、 (e)ACAGGAAGTAGATTACTACTGATAAGAACATTTTGTACGAGATTACTACTGATAAACCCGGAAGTGAGATTACTACTGATATGGTATGATTTTGTAATGGGGTAGGAAGATTACTACTGATAAATTATGGTTCTGGGTGATTCAAGTAACA、 (f)CCTCTTCTCTGTGGGTGACCAGCGTCCTAAGATTACTACTGATAAGGTCAAAGTCCAAGATTACTACTGATATGACGTCAAGATTACTACTGATAGCCCCCAGGGCTGAGTGACAGAAAAACAGAGATTACTACTGATAACAGGAAGT、 (g)AGGTCAAAGTCCAAGATTACTACTGATAAACTAGGTCAAAGGTCAAAGAGATTACTACTGATATGGTATGATTTTGTAATGGGGTAGGAAGATTACTACTGATATGGTTAATATTCACCAGCAGATTACTACTGATAAGAACATTTTGTACG、 (h) GGGACTTTCCAGATTACTACTGATAGGGACTTTCCAGATTACTACTGATAGGGACTTTCCAGATTACTACTGATAGGGACTTTCCAGATTACTACTGATAGGGACTTTCC; (i) ATCACGTGATAGATTACTACTGATAATCACGTGATTATCACGTGATAGATTACTACTGATATGATAGCCAACTGCAGC TAATAATAAACCAAAGATTACTACTGATAACCATGAACTTTGAGATTACTACTGATAGTCTGCTAAGTCAATAATCAGAAT, 11. The method of claim 10, comprising a sequence selected from the group consisting of (j) CTGACGTCAGAGATTACTACTGATACAAAACTAGGTCAAAGGTCAAGATTACTACTGATAGTCTGCTAAGTCAATAATCAGAATAGATTACTACTGATACGCCCCAGCACACATGATCAGAAGATTACTACTGATAAGGTCAAAGGTCA.
14. The method of any one of claims 1 to 8, wherein the method further comprises contacting the cell with the NTF.
15. 15. The method of claim 14, wherein the NTF is provided to the cell as mRNA.
16. 16. The method of claim 14 or 15, wherein said contacting with NT DNA occurs before said contacting with NTF.
17. 16. The method of claim 14 or 15, wherein said contacting with NT DNA occurs after said contacting with NTF.
18. 16. The method of claim 14 or 15, wherein said contacting with NT DNA occurs simultaneously with said contacting with NTF.
19. 20. The method of claim 18, wherein said contacting comprises contacting with a composition comprising said NT DNA and said NTF.
20. 20. The method of claim 19, wherein the composition comprises lipid nanoparticles (LNPs) co-formulated with the NTDNA and the NTF.
21. 20. The method of claim 19, wherein the pharmaceutical composition comprises a first lipid nanoparticle (LNP) formulated with the NTF and a second LNP formulated with the NTF.
22. The method according to any one of claims 1 to 21, wherein the NT DNA has a nucleic acid structure that can be maintained in bacteria.
23. The method according to any one of claims 1 to 21, wherein the NT DNA has a structure that cannot be maintained in bacteria.
24. The method of any one of claims 1 to 23, wherein the cells are non-dividing cells.
25. The method of any one of claims 1 to 23, wherein the cell is a dividing cell.
26. The method of any one of claims 1 to 25, wherein the cells are hepatocytes.
27. 27. The method of any one of claims 1 to 26, wherein the method further comprises contacting the cell with an inducing agent that activates the NTF to which the DTS binds, thereby mediating nuclear entry of the NT DNA.
28. 28. The method of claim 27, wherein the cells are contacted with the NT DNA and the inducing agent simultaneously.
29. 28. The method of claim 27, wherein the NT DNA and the inducing agent are contacted with the cells sequentially.
30. 30. The method of claim 29, wherein the NT DNA is contacted with the cells before the inducing agent.
31. 30. The method of claim 29, wherein the NT DNA is contacted with the cell after the inducing agent.
32. The method of any one of claims 1 to 31, wherein the method is in vivo.
33. The method of any one of claims 1 to 31, wherein the method is in vitro.
34. 34. The method of claim 32 or 33, wherein the cell is a non-dividing cell.
35. A non-naturally occurring nuclear-targeted deoxyribonucleic acid (NTDNA), said NTDNA comprising: a DNA nuclear targeting sequence (DTS) comprising a targeting factor binding sequence (TFBS) for a nuclear targeting factor (NTF); and a cargo nucleic acid sequence heterologous to said DTS.
36. 36. The NT DNA of claim 35, wherein the cargo nucleic acid sequence comprises an expression cassette.
37. 36. The NT DNA of claim 35, wherein the cargo nucleic acid comprises a promoter heterologous to the DTS.
38. 36. The NT DNA of claim 35, wherein the cargo nucleic acid comprises a coding sequence heterologous to the DTS.
39. 39. The NT DNA of any one of claims 35 to 38, wherein the cargo nucleic acid sequence is flanked by sequences that are homologous to genomic sequences of the cell in which the NT DNA is configured to be used.
40. The NT DNA of any one of claims 35 to 38, wherein the DTS comprises two or more different TFBSs.
41. The NT DNA of any one of claims 35 to 38, wherein the DTS comprises two or more copies of the same DTS.
42. The NT DNA of claim 40 or 41, wherein the DTS comprises 5 to 100 nucleotides between each TFBS.
43. The NT DNA of any one of claims 35 to 42, wherein the DTS comprises one or more TFBSs selected from Table 1 or Table 2.
44. The NT DNA of any one of claims 35 to 42, wherein the DTS comprises a TFBS for an NTF selected from the group consisting of HNF1A, PPARA, HNF4A, CEBPA, NR3C1, ONECUT1, TBP, NFkB, a TetR protein, a ZF-CCR5 protein, a TALE protein, a GAL4 protein, an Arc protein, an Mnt protein, a PurR protein, a Bac434 protein, a GCN4 protein, a LacR protein, an I-SceI D44A protein, and a Cas9 protein.
45. (a) when the TFBS is for HNF1A, the sequence is selected from GGTTAATAATTAAC and AGTATGGTTAATGATCTACAG; (b) when the TFBS is for PPARA, the sequence is selected from AACTAGGTCAAAGGTCA, CAAAACTAGGTCAAAGGTCA, and AACTAGGTCAAAGGTCAAAG; (c) when the TFBS is for HNF4A, the sequence is selected from TCGAGCGCTGGGCAAAGGTCACCTGC, TCGAGCGCAGGTCAAAGGTCACCTGC, and AGGTCAAAGTCCA; (d) when the TFBS is for CEBPA, the sequence is TGGTATGATTTTGTAATGGGGTAGGA; (e) when the TFBS is for NR3C1, the sequence is selected from AGAACAAAATGTTCT, AAGAACAAAATGTTCTT, AGAACATTTTGTACG, and AAGAACATTTTGTACGT; (f) when the TFBS is for ONECUT1, the sequence is GTCTGCTAAGTCAATAATCAGAAT; (g) when the TFBS is for TBP, the sequence is TATAAAA; (h) if the TFBS is for NFkB, the sequence is GGGACTTTCC; (i) when the TFBS is for a TetR protein, the sequence comprises TCCCTATCAGTGATAGAGA; (j) when the TFBS is for a ZF-CCR5 protein, the sequence comprises AAACTGCAAAG; (k) if the TFBS is for a TALE protein, the sequence comprises TTCATTACACCTGCAGCT, ATAAACCCCCTCCAA, or TCGAGTTTACTCCCTATCAGTGATAGAGAACG; (l) When the TFBS is for a GAL4 protein, the sequence is CGG-N 11 - includes CCG, (m) when the TFBS is for an Arc protein, the sequence comprises RYRVTAGANNNNNNTCTABYRY; (n) when the TFBS is for an Mnt protein, the sequence comprises GGNCCACNGTGGNCC; (o) when the TFBS is for a PurR protein, the sequence comprises ACGCAAACGTTTTCGT; (p) when the TFBS is for a Bac434 protein, the sequence comprises ACAAGAAAGTTTGT, ACAAGATACATTGT, or ACAAGAAAACTGT; (q) if the TFBS is for a GCN4 protein, the sequence comprises TGACTC; (r) when the TFBS is for a LacR protein, the sequence comprises TTGTTATCCGCTCACAA; (s) when the TFBS is for an I-SceI protein, the sequence comprises TAGGGATAACAGGGTAAT.
46. The combination of NTFs is (a) HNF1A, NR1I3, PPARA, HNF1A, and PPARA; (b) HNF4A, NR1I3, NR3C1, NR3C1, and HNF4A; (c) FOXA1, HNF4A, HNF1A, HNF4A, and HNF1A; (d) NR3C1, CEBPA, HNF1A, CEBPA, and PPARA; (e) ETS1, NR3C1, ELF5, CEBPA, and NR1I3; (f) NR1I3, HNF4A, CREB1, NR1I3, and ETS1; (g) HNF4A, PPARA, CEBPA, HNF1A, and NR3C1; (h) NFKB1, NFKB1, NFKB1, NFKB1, and NFKB1; (i) SREBF1, MLXIPL, HNF1A, CREB3L3, and ONECUT1; and (j) The NT DNA of claim 44, selected from the group consisting of CREB1, PPARA, ONECUT1, HNF4A, and PPARA.
47. The DTS is (a) GGTTAATAATTAACAGATTACTACTGATAACCTCTTCTCTGTGGGTGACCAGCGTCCTAAGATTACTACTGATAAACTAG GTCAAAGGTCAAAGATTACTACTGATAAGTATGGTTAATGATCTACAGAGATTACTACTGATACAAAACTAGGTCAAGGTCA, (b) TCGAGCGCTGGGCAAAGGTCACCTGCAGATTACTACTGATAGCCCCCAGGGCTGAGTGACAGAAAAAACAGAGATTACTA CTGATAAGAACAAAATGTTCTAGATTACTACTGATAAAAGAACATTTTGTACGTAGATTACTACTGATAAGGTCAAGTCCA, (c) TGTTTACTTTAGATTACTACTGATATCGAGCGCAGGTCAAGGTCACCTGCAGATTACTACTGATAGGTTAATAATT AACAGATTACTACTGATATCGAGCGCTGGGCAAAGGTCACCTGCAGATTACTACTGATAAGTATGGTTAATGATCTACAG, (d)AAGAAAAAATGTTCTTAGATTOOOATATGTATGATTTTGTAATGTTTAGAAATTTTTOA - (e)アァGGAAGTAGATTTATGATAAGAGATTTTGTAGAGATTTTAGATAAGATTAT ACT (f)CCTCTTCTCTGTGGCGGGGAGCGTCCTAAGATTTTEGATAGGTCAAAGTCCAAGATTAT - (g) - - (i)ATCGTGATAGATTTATGATAATCGGTGATTATCATTGATAGATTTAGGATATGATAACA The 45. The NT DNA of claim 44, comprising a sequence selected from the group consisting of (j) CTGACGTCAGAGATTACTACTGATACAAAACTAGGTCAAAGGTCAAGATTACTACTGATAGTCTGCTAAGTCAATAATCAGAATAGATTACTACTGATACGCCCCAGCACACATGATCAGAAGATTACTACTGATAAGGTCAAAGGTCA.
48. The NT DNA according to any one of claims 35 to 47, wherein the NT DNA has a structure that allows it to be maintained in bacteria.
49. The NT DNA according to any one of claims 35 to 47, wherein the NT DNA has a structure that cannot be maintained in bacteria.
50. 1. A composition comprising: A composition comprising a cytosolic delivery vehicle comprising the NT DNA of any one of claims 35 to 49.
51. 51. The composition of claim 50, wherein the cytosolic delivery vehicle comprises a non-viral delivery vehicle.
52. 52. The composition of claim 51, wherein the non-viral delivery vehicle comprises a lipid nanoparticle (LNP).
53. 51. The composition of claim 50, wherein the composition further comprises one or more ribonucleic acids packaged within a delivery vehicle.
54. 54. The composition of claim 53, wherein the delivery vehicle comprising the NT DNA and the delivery vehicle comprising the ribonucleic acid are the same delivery vehicle.
55. 54. The composition of claim 53, wherein the delivery vehicle comprising the NT DNA and the delivery vehicle comprising the ribonucleic acid are different delivery vehicles.
56. 56. The composition of any one of claims 53 to 55, wherein the one or more ribonucleic acids are selected from the group consisting of mRNA and guide RNA (gRNA).
57. 48. The composition of claim 47, wherein the one or more ribonucleic acids is mRNA.
58. 58. The composition of claim 57, wherein the mRNA encodes an endonuclease.
59. 58. The composition of claim 57, wherein the mRNA encodes a nuclear targeting factor (NTF) that binds to the DTS.
60. 1. A system for delivery of a cargo nucleic acid sequence to the nucleus of a cell, said system comprising: NT DNA according to any one of claims 35 to 49 or a composition according to any one of claims 50 to 59; an inducer.
61. 61. The system of claim 60, wherein the iDTS comprises a TFBS selected from the sequences in Table 2, and the inducer is a cognate inducer in Table 2.
62. 1. A method for delivering a cargo nucleic acid sequence to the nucleus of a cell, said method comprising: contacting the cells with the NT DNA of any one of claims 35 to 49 or the composition of any one of claims 50 to 59, delivering said cargo nucleic acid to the nucleus of said cell.
63. 63. The method of claim 62, further comprising contacting the cell with an inducer that activates a nuclear targeting factor to which the DTS binds.
64. 1. A method for integrating a cargo nucleic acid sequence into the genome of a cell, said method comprising: The cells NT DNA according to any one of claims 35 to 49 or a composition according to any one of claims 50 to 59, and contacting the nucleic acid sequence with an endonuclease and optionally a guide RNA; integrating said cargo nucleic acid sequence into the genome of said cell.
65. 65. The method of claim 64, further comprising contacting the cell with an inducer that activates a nuclear targeting factor to which the DTS binds.
66. 66. The method of claim 64 or 65, wherein the endonuclease is encoded by mRNA.
67. 66. The method of claim 64 or 65, wherein the endonuclease is encoded by DNA.
68. 68. The method of any one of claims 64 to 67, wherein the endonuclease is selected from the group consisting of a CRISPR-Cas endonuclease, a TALEN, a zinc finger nuclease, a homing endonuclease, and an integrase.
69. 69. The method of any one of claims 64 to 68, wherein the cargo nucleic acid sequence comprises a coding sequence.
70. 69. The method of any one of claims 64 to 68, wherein the cargo nucleic acid comprises a promoter sequence.