Novel transcription factors
Patent Information
- Application Number
- JP2024516730
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-16
- Filing Date
- 2022-09-15
- Publication Date
- 2025-09-24
AI Technical Summary
Existing adeno-associated virus (AAV) vectors are limited by their packaging capacity, making it difficult to deliver large therapeutic genes, and there is a need for transcription factors that provide reduced immunogenicity, off-target effects, increased specificity, and enhanced therapeutic efficacy for gene therapy.
Development of a novel transcription factor comprising a DNA binding domain and multiple transcriptional activation domains (TADs), such as VP16, VP64, and zinc finger proteins, which are designed to be compact and highly active, allowing for efficient gene regulation and expression.
The novel transcription factor effectively modulates gene expression, overcoming AAV vector size limitations and providing improved specificity and efficacy in gene therapy applications.
Smart Images

Figure 00000087_0000 
Figure 00000087_0001 
Figure 00000087_0002
Abstract
Description
[Technical field]
[0001] The present invention relates to novel transcription factors for regulating the expression of a gene of interest by fusing it to a DNA binding domain that targets the gene of interest. [Background technology]
[0002] Genomic alterations that result in reduced transcription or activity of one or more genes or gene products are causative factors of a variety of mammalian diseases. One such genomic alteration is haploinsufficiency. In haploinsufficiency, only one functional copy of a gene is present, and that single copy does not produce enough gene product to produce a wild-type phenotype. Another type of disease is caused by genomic alterations in one or both copies of a gene that alter the gene product such that the gene product exhibits reduced activity, rather than abolished activity. In yet another type of disease, genomic alterations reduce the transcription or reduce the stability of the transcript of one or both copies of a gene such that insufficient gene product is present to produce a wild-type phenotype.
[0003] Many approaches have been attempted to treat such diseases by increasing the amount or activity of one or more genes whose transcription or activity is reduced. One such approach is the delivery of wild-type copies of target genes to patients by adeno-associated virus (AAV) vectors that produce functional proteins. Adeno-associated virus (AAV) is a small, replication-deficient, non-enveloped virus that infects humans and some other primate species. It is not known whether AAV causes human disease or induces a mild immune response, and several properties of AAV, such as the ability of AAV vectors to infect both dividing and dormant cells without integrating into the host cell genome, make this virus an attractive vehicle for the delivery of therapeutic proteins by gene therapy. However, AAV gene therapy vectors have several drawbacks in practice. In particular, the cloning capacity of AAV vectors is limited as a result of the packaging capacity of viral DNA. The single-stranded DNA genome of wild-type AAV is approximately 4.7 kilobases (kb). In practice, it is believed that up to about 5.0 kb of the AAV genome is completely, i.e., full length, purged into AAV viral particles. Due to the requirement that the nucleic acid genome in an AAV vector must have two AAV inverted terminal repeats (ITRs) of approximately 145 bases, the DNA packaging capacity of the AAV vector is such that the protein coding sequence is a maximum of approximately 4.4 kb.
[0004] Due to this size limitation, large therapeutic genes, such as those that exceed about 4.4 kb in length, are generally not suitable for use in AAV vectors.One approach to overcome the size limitation of AAV is to increase the target gene transcription of a gene by delivering a transcription factor that targets the promoter region of the target gene instead of delivering the target gene itself.Therefore, there is a need for improved transcription factors, particularly small in size and strong in activity, that are suitable for delivery by AAV vector and can regulate the expression of any endogenous gene to help reverse the effects of disease or disorder, and in particular therapies that provide reduced immunogenicity, reduced off-target effects, increased specificity for target genes and / or increased therapeutic efficacy. Summary of the Invention
[0005] In one aspect, provided herein is a transcription factor comprising a DNA binding domain and at least three transcription activation domains (TADs), which can be the same or different.
[0006] In some embodiments, the TAD comprises i) at least one acidic TAD and at least one Q-rich TAD, or ii) at least one acidic TAD and at least one P-rich TAD. In some embodiments, the acidic TAD comprises one or more copies of a TAD selected from VP16, VP64, VP7, ATF6, TFE3, ATF6-11, or fragments thereof. In some embodiments, the Q-rich TAD comprises one or more copies of a TAD selected from Oct2-Q, SP1-Q1, or fragments thereof. In some embodiments, the P-rich TAD comprises one or more copies of a TAD selected from TFAP2-P, Oct2-P, or fragments thereof.
[0007] In some embodiments, the TAD is selected from the group consisting of full length VP64, one or more repeats of VP16, full length or partial p65, full length RTA, full length VP7, full length or partial TEF3, full length or partial ATF6-11, full length or partial ATF6 acidic, full length or partial SP1-rich Q, full length or partial Oct2-rich Q, full length or partial Oct2-rich P, full length or partial TFAP2-rich P, an active fragment of VP64, a transcriptionally active fragment of p65, a transcriptionally active fragment of RTA, an active fragment of VP7, an active fragment of TEF3, an active fragment of ATF6-11, an active fragment of ATF6 acidic, an active fragment of SP1-rich Q, an active fragment of Oct2-rich Q, an active fragment of Oct2-rich P, an active fragment of TFAP2-rich P, and any combination thereof.
[0008] In some embodiments, the total length of the TADs is less than 2000 aa, less than 1500 aa, less than 1000 aa, less than 750 aa, less than 500 aa, less than 300 aa, less than 250 aa, less than 200 aa, or less than 150 aa in length.
[0009] In some embodiments, the transcription factor comprises a sequence of SEQ ID NOs: 1-28 or a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity thereto. In certain embodiments, the transcription factor comprises a sequence of any one of SEQ ID NOs: 1-18 or a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity thereto.
[0010] In some embodiments, the nucleic acid sequence encoding the transcription factor comprises a sequence of SEQ ID NO: 30-47 or a sequence having at least 80%, at least 85%, at least 90%, or at least 95% sequence identity thereto.
[0011] In some embodiments, the DNA binding domain is linked to any one of the transcription factors with or without a linker. In certain embodiments, the DNA binding domain is linked to the N-terminus of the transcription factor directly or via a linker. In certain embodiments, the DNA binding domain is linked to the C-terminus of the transcription factor directly or via a linker. In certain embodiments, the linker comprises or consists of GGSGGGSG (SEQ ID NO: 59) or GGSGGGSGGGSGGGSG (SEQ ID NO: 60).
[0012] In some embodiments, the DNA binding domain binds to a genomic region of a gene of interest. In certain embodiments, the DNA binding domain is a gRNA / Cas complex, a transcription activator-like (TAL) effector, or a zinc finger protein. In certain embodiments, the Cas molecule in the gRNA / Cas complex is or is derived from S. pyogenes Cas9, C. jejune Cas9, S. aureus Cas9, or Deltaproteobacteria (Dpb) CasX.
[0013] In some embodiments, the Cas molecule lacks one or more activities, and optionally, the one or more activities is a cleavage activity.
[0014] In some embodiments, the Cas molecule is encoded by a nucleic acid molecule comprising fewer than 4,000 nucleotides, optionally the Cas molecule is encoded by a nucleic acid molecule comprising fewer than 3,500 nucleotides or fewer than 2,500 nucleotides. In certain embodiments, the Cas molecule is or is derived from S. aureus Cas9.
[0015] In some embodiments, the Cas molecule comprises one or more amino acid deletions compared to the wild-type Cas molecule sequence.
[0016] In some embodiments, the DNA binding domain is a zinc finger protein. In some embodiments, the zinc finger protein comprises 6-9 zinc finger domains. In some embodiments, the zinc finger protein binds to a promoter region of a gene of interest. In certain embodiments, the zinc finger protein binds to a 6, 9, 12, 15, 18, 21, 24, or 27 bp sequence of the promoter region of a gene of interest.
[0017] In some embodiments, the zinc finger protein comprises a sequence having at least 85%, at least 90%, or at least 95% identity to SEQ ID NO:61, SEQ ID NO:70, SEQ ID NO:71, or SEQ ID NO:73.
[0018] In another aspect, the application provides a nucleic acid molecule encoding one or more components, optionally all of the components, of a transcription factor according to any one of the above embodiments.
[0019] In another aspect, the present application provides a vector comprising the nucleic acid molecule. In some embodiments, the vector comprises a first promoter operably linked to a nucleic acid sequence encoding a DNA binding domain.
[0020] In some embodiments, the vector comprises a promoter that drives transcription in a human cell.
[0021] In some embodiments, the first promoter is operable in a neuron. In certain embodiments, the neuron is a GABAergic neuron or an inhibitory neuron or an inhibitory interneuron. In certain embodiments, the neuron is a parvalbumin-positive GABAergic neuron, or a somatostatin-positive GABAergic neuron, or a vasoactive intestinal peptide-positive GABAergic neuron.
[0022] In some embodiments, the vector further comprises one or more of: a. a polyA sequence; b. an intron sequence; or c. an enhancer sequence.
[0023] In some embodiments, the vector further comprises regulatory elements that control the production and / or degradation of the transcription factor.
[0024] In some embodiments, the regulatory element is a minigene linked to a transcription factor, wherein the minigene contains a splice regulator binding sequence, and in the presence of the splice regulator, the minigene undergoes splicing resulting in increased or decreased expression of the transcription factor.
[0025] In some embodiments, the regulatory element is a destabilizing domain, and in the presence of a small molecule that specifically binds to the destabilizing domain, expression of the transcription factor is decreased.
[0026] In some embodiments, the vector is an adenoviral vector, an adeno-associated viral (AAV) vector, or a lentiviral vector, or an adenoviral vector or a herpes simplex viral vector.
[0027] In certain embodiments, the vector is an AAV vector. In certain embodiments, the vector is an AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9 or AAV.rh10 vector.
[0028] In another aspect, the present application provides a cell comprising the vector according to any one of the above embodiments.
[0029] In another aspect, the application provides a method of selective expression of a transgene in a subject, comprising administering to the subject a viral vector described in any of the above embodiments.
[0030] In another aspect, the application provides a method of treating a disease associated with a mutation in a gene of interest, comprising administering to a subject in need thereof a viral vector described in any of the above embodiments.
[0031] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief description of the drawings]
[0032] [Figure 1] The activity of TAD1, TAD2+ (TAD2 with three nuclear localization signal (NLS) sequences), TAD2- (TAD2 with two NLS sequences) and TAD3 compared to VPR3 for activating the TRE promoter is shown. TADs activate TRE-containing promoters when linked to mini-Sa-Cas9. HEK293T cells were transfected with a plasmid carrying an sgRNA targeting the TRE sequence, mini-Sa-dCas9-activator, TRE promoter-luciferase and GPK-renilla luciferase plasmids. After 2 days, luciferase activity was read out. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. Relative promoter activity triggered by Sa-dCas9-activator was determined by calculating the fold change from the control guide RNA sample. The 1xTRE promoter contains one TRE site; the 3xTRE promoter carries three repeats of the TRE site. VPR3 (Ma et al. 2018) was also linked to mini-Sa-Cas9 and used as a positive transcription activator control. The sequence of the sgRNA targeting the TRE is listed in SEQ ID NO:64. [Diagram 2]The activity of TAD1-TAD7 compared to VPR3 for activation of the mouse SCN1A promoter is shown. HEK293T cells were transfected with miniSa-dCas9-activators (TAD1-TAD7), mouse scn1a promoter-luciferase, gRNA targeting the mouse scn1a promoter, and GPK-renilla luciferase plasmids. Luciferase activity was read out after 2 days. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. Each miniSa-dCas9-activator was transfected at 20, 6.7, and 2.2 ng / well, and the results showed a dose-dependent effect. The sequence of sgRNA3 corresponds to SEQ ID NO: 65. [Diagram 3] The activity of TAD7-TAD10 compared to VPR3 is shown for activation of mouse SCN1A promoter. HEK293T cells were transfected with miniSa-dCas9-activator (TAD7-TAD10), mouse scn1a promoter-luciferase, gRNA targeting mouse scn1a promoter and GPK-Renilla luciferase plasmids. After 2 days, luciferase activity was read. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. When each miniSa-dCas9-activator was transfected at 20, 6.7 and 2.2 ng / well and co-transfected with active guide RNA sgRNA3, the results showed dose-dependent activation. When miniSa-dCas9 was co-transfected with inactive guide RNA sgRNA11, the promoter was not activated. The sequences of sgRNA3 and sgRNA11 correspond to SEQ ID NO: 65 and SEQ ID NO: 66, respectively. [Figure 4]The activity of TAD2 and TAD9 compared to VPR for activating the TRE and SCN1a promoters when paired with sgRNA is shown. The mini-Sa-dCas9-TAD driven by the ubiquitin C promoter activates the mouse scn1a promoter and TRE-containing promoters when paired with active sgRNA. HEK293T cells were transfected with various amounts of mini-Sa-dCas9-activator (TAD2, TAD9 and VPR), mouse scn1a promoter-luciferase or TRE-containing promoter, TRE sgRNA for TRE-containing promoters respectively; sgRNA3 and GPK-Renilla luciferase for mSCN1A promoter. Luciferase activity was read out after 2 days. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. [Diagram 5] The activity of TAD9, TAD11, TAD14, and TAD16 compared to VP64 and full VPR for activation of mouse SCN1A and human SCN1A promoters is shown. HEK293T cells were transfected with miniSa-dCas9-activator, mouse scn1a promoter-luciferase or human scn1a promoter-luciferase, gRNA targeting the scn1a promoter, and GPK-Renilla luciferase plasmids. Luciferase activity was read out after 2 days. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. The newly created transcription activation domains (TAD9, TAD11, TAD14, and TAD16) have activity comparable to or greater than VPR (Chavez et al. 2015) and stronger than VP64. [Figure 6]The activity of TAD9, TAD11, TAD12, TAD13, TAD14, and TAD15 is shown compared to VP64 and full VPR for activating TRE and human SCN1A promoter. HEK293T cells were transfected with miniSa-dCas9-activator, TRE-containing promoter or human scn1a promoter-luciferase, gRNA targeting scn1a promoter, and GPK-Renilla luciferase plasmids. Luciferase activity was read out after 2 days. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. The newly created transcription activation domain has activity equal to or greater than VPR (Chavez et al. 2015) and stronger than VP64. [Figure 7] The activity of TAD9, TAD11, TAD12, TAD13, TAD14, and TAD15 compared to VP64 and full VPR in activating the mouse SCN1A promoter with different activity sgRNAs. HEK293T cells were transfected with miniSa-dCas9-activator, mouse scn1a promoter-luciferase, sgRNA targeting mouse scn1a promoter, and GPK-renilla luciferase plasmids. Luciferase activity was read after 2 days. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. The sequence of sgRNA42 corresponds to SEQ ID NO:67. [Figure 8]The activity of TAD8a, TAD9 and TAD15 compared to VP64 for activating a TRE reporter gene in U2OS cell line is shown. U2OS cells were transfected with miniSa-dCas9-activator, TRE-containing promoter or human scn1a promoter-luciferase, gRNA sequence targeting TRE and GPK-Renilla luciferase plasmids. Luciferase activity was read after 2 days. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. TAD8a, TAD9 and TAD15 show higher activity than VP64, with TAD15 showing the highest activity of all. [Figure 9] Figure 1 shows the activity of TAD15 linked to zinc finger proteins to enhance gene transcription. Zinc finger protein ZFP-C (SEQ ID NO: 70, "6CAG repeat" is disclosed as SEQ ID NO: 93) targeting 6 CAG repeats and zinc finger protein E2C (SEQ ID NO: 71) targeting human erb2 gene promoter were linked to TAD15. Each transcription activator was co-transfected with a reporter carrying two E2C elements (E2C 2x) or a reporter carrying 8 CAG repeats (CAG24) (SEQ ID NO: 94) and GPK-Renilla luciferase was transfected into HEK293T cells. Luciferase activity was read out after 2 days. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. Transcription factors specifically activated the promoters. That is, E2C-TAD15 activated only E2C element-containing promoters; ZFP-C-TAD15 activated only CAG repeat-containing promoters. [Figure 10]Figure 1 shows the activity of TAD15 on the activation of endogenous gene promoters. The zinc finger protein E2C (SEQ ID NO: 71) recognizes an element in the human erb2 gene promoter (Beerli et al 1998). Erb2 encodes the receptor protein Her2. E2C-TAD15 was transfected into HEK293T cells. Transfected cells were identified by staining for hemagglutinin (HA) tagged on E2C-TAD15. Her2 expression levels were quantified by immunostaining for Her2 on the cell surface. Cells with E2C-TAD15 (HA positive) showed more than three-fold higher Her2 levels compared to non-transfected cells (HA negative). [Figure 11] Activity of TAD15 in activating endogenous neuronal genes in mouse GABAergic neurons is shown. TAD13, TAD14 and TAD15 were each ligated with the zinc finger protein Nav-ZF2 (SEQ ID NO: 61) targeting the SCN1A promoter. They were purged into lentivirus and used to infect cultured neurons, achieving 70-90% transduction rates. After 4 days, cells were lysed and RNA was harvested. SCN1A and MAP2 transcripts were analyzed by qRT-PCR. Increase in SCN1A transcripts by transcriptional activators is measured by the fold increase in normalized SCN1A mRNA levels relative to normalized SCN1A mRNA levels in cells without virus. [Figure 12] A schematic diagram, not to scale, showing the arrangement of activation domains in each TAD molecule is shown. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0033] definition In order that the present invention may be more readily understood, certain terms are defined throughout the detailed description. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention pertains.
[0034] Unless otherwise stated, the following terms and phrases used herein are intended to have the following meanings:
[0035] As used herein, "mammal" includes any animal classified as a mammal, including, but not limited to, humans, domestic animals, farm animals, companion animals, and the like.
[0036] As used herein, the term "subject" or "patient" refers to, but is not limited to, humans and non-human mammals, including primates, rabbits, pigs, horses, dogs, cats, sheep and cows. Preferably, the subject or patient is a human.
[0037] The terms "a," "an," or "the" refer to one or more of the grammatical object of the article. The term may refer to "one," "one or more," "at least one," or "one or more." By way of example, "an element" means one element or more than one element. The term "or" refers to "and / or" unless otherwise stated. The terms "comprising" or "containing" are not limiting.
[0038] The term "about" when referring to a measurable value, e.g., amount, duration in time, and the like, is meant to encompass variations of ±20%, or in some cases ±10%, or in some cases ±5%, or in some cases ±1%, or in some cases ±0.1% from the specifically indicated value, where such variations are appropriate for practicing the methods of the present disclosure.
[0039] The term "conservative sequence modification" refers to an amino acid modification that does not significantly affect or alter the binding properties of the Cas9 molecule containing the amino acid sequence. Such conservative modifications include amino acid substitutions, additions, and deletions. Modifications can be introduced into the Cas9 molecules or fragments described herein by standard techniques known in the art, such as site-directed mutagenesis and PCR-mediated mutagenesis. A conservative amino acid substitution is one in which an amino acid residue is replaced with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art. These families include amino acids with basic side chains (e.g., lysine, arginine, histidine), acidic side chains (e.g., aspartic acid, glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, cysteine, tryptophan), nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine), beta-branched side chains (e.g., threonine, valine, isoleucine), and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, histidine). Thus, one or more amino acid residues in a Cas9 molecule can be replaced with other amino acid residues from the same side chain family, and the altered Cas9 molecule can be tested using the functional assays described herein. Similarly, "conservative sequence modifications" refer to amino acid modifications that do not significantly affect or alter the binding properties of the transcription factors described herein.
[0040] The term "encode" refers to the inherent property of a specific sequence of nucleotides in a polynucleotide, such as a gene, cDNA, mRNA, etc., to serve as a template for the synthesis in biological processes of other polymers and macromolecules having either a defined sequence of nucleotides (e.g., rRNA, tRNA, and mRNA) or a defined sequence of amino acids and biological properties resulting therefrom. Thus, a gene, cDNA, or RNA encodes a protein if transcription and translation of the mRNA corresponding to that gene produces the protein in a cell or other biological system. Both the coding strand, whose nucleotide sequence is identical to the mRNA sequence and is usually provided in a sequence listing, used as a template for transcription of the gene or cDNA, and the non-coding strand, can be said to encode the protein or other product of that gene or cDNA.
[0041] Unless otherwise specified, a "nucleotide sequence encoding an amino acid sequence" includes all nucleotide sequences that are degenerate versions of each other and that encode the same amino acid sequence. The phrase nucleotide sequence encoding a protein or RNA can also include introns, to the extent that a nucleotide sequence encoding a protein may in some versions contain introns.
[0042] The terms "effective amount" or "therapeutically effective amount" are used interchangeably herein and refer to an amount of a compound, formulation, material or composition described herein effective to achieve a particular biological result.
[0043] The term "endogenous" refers to any material that originates or is produced within an organism, cell, tissue, or system.
[0044] The term "exogenous" refers to any substance that is introduced from or produced outside an organism, cell, tissue or system.
[0045] The term "expression" refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcript) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as the "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.
[0046] The term "transfer vector" refers to a composition of matter that contains an isolated nucleic acid and can be used to deliver the isolated nucleic acid into a cell. Many vectors are known in the art, including, but not limited to, linear polynucleotides, polynucleotides associated with ionic or amphipathic compounds, plasmids, and viruses. Thus, the term "transfer vector" includes autonomously replicating plasmids or viruses. The term should also be construed to further include non-plasmid and non-viral compounds that facilitate the transfer of nucleic acid into a cell, such as polylysine compounds, liposomes, and the like. Examples of viral transfer vectors include, but are not limited to, adenoviral vectors, adeno-associated viral vectors, retroviral vectors, lentiviral vectors, and the like.
[0047] The term "expression vector" refers to a vector that contains a recombinant polynucleotide comprising an expression control sequence operably linked to a nucleotide sequence to be expressed. An expression vector contains sufficient cis-acting elements for expression; other elements for expression can be supplied by the host cell or in an in vitro expression system. Expression vectors include all those known in the art, including cosmids, plasmids (e.g., naked or contained in liposomes) and viruses (e.g., lentiviruses, retroviruses, adenoviruses and adeno-associated viruses) into which the recombinant polynucleotide has been incorporated.
[0048] The terms "homologous," "homology," or "identity" refer to the subunit sequence identity between two polymeric molecules, e.g., between two nucleic acid molecules such as two DNA molecules or two RNA molecules, or between two polypeptide molecules. When both subunit positions of the two molecules are occupied by the same monomer subunit; for example, if a position in each of the two DNA molecules is occupied by an adenine, then they are homologous or identical at that position. The homology between two sequences is a direct function of the number of matching or homologous positions; for example, if half of the positions in the two sequences (e.g., 5 positions in a polymer 10 subunits long) are homologous, then the two sequences are 50% homologous, and if 90% of the positions (e.g., 9 out of 10) are matched or homologous, then the two sequences are 90% homologous.
[0049] The term "isolated" means altered or removed from the natural state. For example, a nucleic acid or peptide that is naturally present in a living animal is not "isolated," but the same nucleic acid or peptide partially or completely separated from the coexisting materials of its natural state is "isolated." An isolated nucleic acid or protein can exist in a substantially purified form, or can exist in a non-native environment, such as a host cell.
[0050] The term "operably linked" or "transcriptional control" refers to the functional linkage of a regulatory sequence with a heterologous nucleic acid sequence resulting in expression of the latter. For example, a first nucleic acid sequence is operably linked to a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. By way of example, a promoter is operably linked to a coding sequence if it affects the transcription or expression of the coding sequence. A promoter or regulatory sequence can be a cis-acting or trans-acting element. Operably linked DNA sequences can be contiguous with each other and, for example, when necessary to join two protein coding regions, are in the same reading frame.
[0051] The term "nucleic acid" or "polynucleotide" refers to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) and polymers thereof in either single- or double-stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogs of natural nucleotides that have similar bonds as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences, as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994)).
[0052] The terms "peptide", "polypeptide" and "protein" are used interchangeably and refer to compounds composed of amino acid residues covalently linked by peptide bonds. A protein or peptide must contain at least two amino acids, and no limit is placed on the maximum number of amino acids that may make up a protein or peptide sequence. A polypeptide includes any peptide or protein that contains two or more amino acids linked together by peptide bonds. As used herein, the term refers to both short chains, also commonly referred to in the art as peptides, oligopeptides and oligomers, for example, and longer chains, commonly referred to in the art as proteins, of which there are many types. "Polypeptide" includes, inter alia, for example, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, mutants of polypeptides, modified polypeptides, derivatives, analogs, fusion proteins. A polypeptide includes natural peptides, recombinant peptides, or combinations thereof.
[0053] As used herein, the term "promoter" or "promoter sequence" refers to a DNA regulatory sequence capable of promoting transcription (e.g., causing detectable levels of transcription and / or increasing detectable levels of transcription above levels provided in the absence of the promoter) of an operably linked coding or non-coding sequence, e.g., downstream (3' direction) coding or non-coding sequence, e.g., via binding of RNA polymerase. In some embodiments, a promoter sequence is bounded at its 3' end by a transcription initiation site and extends upstream (5' direction) to include a minimum number of bases or elements to initiate transcription at a level detectable above background. In some embodiments, a promoter sequence may include a transcription initiation site as well as a protein binding domain involved in binding of RNA polymerase. In addition to sequences sufficient to initiate transcription, a promoter may also include sequences of other regulatory elements involved in regulating transcription (e.g., enhancers, Kozak sequences, and introns). A variety of promoters, including inducible and constitutive promoters, may be used to drive the vectors disclosed herein. Examples of promoters known in the art that can be used in some embodiments, such as the viral vectors disclosed herein, include CMV promoter, CBA promoter, smCBA promoter, and promoters derived from immunoglobulin genes, SV40, or other tissue-specific genes (e.g., RLBP1, RPE, VMD2). Additionally, standard techniques are known in the art for creating functional promoters by mixing and matching known regulatory elements. Fragments of promoters, such as those that retain at least a minimum number of bases or elements to initiate transcription at a level detectable above background, can also be used.
[0054] In some embodiments, the promoter can be a constitutively active promoter (i.e., a promoter that drives expression constitutively in any cell type and / or under any condition). In other embodiments, the promoter can be a constitutively active promoter in the context of a particular tissue, e.g., neurons, cardiac cells, etc. In other embodiments, the promoter can be an inducible promoter (i.e., a promoter whose activity is controlled by an external stimulus, e.g., a particular temperature, the presence of a compound or protein). In some embodiments, the promoter can be a spatially restricted promoter that can drive activity or not depending on the physical context in which the promoter is found. Non-limiting examples of spatially restricted promoters include tissue-specific promoters, cell type-specific promoters, etc. In some embodiments, the promoter can be a temporally restricted promoter that drives expression depending on the temporal context in which the promoter is found. For example, a temporally restricted promoter can drive expression only at a particular stage of embryonic development or at a particular stage of a biological process. A non-limiting example of a temporally restricted promoter includes the mouse hair follicle cycle promoter.
[0055] In some embodiments, the promoter is tissue specific such that in a multicellular organism, the promoter drives expression only in a specific subset of cells. For example, tissue specific promoters include, but are not limited to, neuron specific promoters, adipocyte specific promoters, cardiomyocyte specific promoters, smooth muscle specific promoters, photoreceptor specific promoters, and the like. A neuron specific promoter refers to a promoter that preferentially drives or regulates the expression of an operably linked heterologous nucleic acid, such as one encoding a protein or peptide of interest or shRNA, in a neuron, when delivered to a neuronal cell, including, for example, peripherally, directly administered to the central nervous system (CNS), or in vitro, ex vivo, or in vivo, compared to expression in a non-neuronal cell. As used herein, the terms "treat", "treatment" and "treating" refer to a partial or complete reduction or amelioration of the progression, severity and / or duration of a seizure disorder, or the amelioration of one or more symptoms (preferably one or more discernible symptoms) of a seizure disorder resulting from the administration of one or more therapies.
[0056] The terms "DNA regulatory sequence," "control element," and "regulatory element," as used interchangeably herein, refer to transcriptional and translational control sequences, such as promoters, enhancers, silencers, polyadenylation signals, terminators, protein degradation signals, etc., that provide and / or regulate the transcription of a non-coding sequence (e.g., a short hairpin RNA) or a coding sequence (e.g., a PGRN) and / or regulate the translation of an encoded polypeptide.
[0057] In specific embodiments, the terms "treat", "treatment" and "treating" refer to the amelioration of at least one measurable physical parameter of a seizure disorder, a parameter not necessarily discernible by the patient. In other embodiments, the terms "treat", "treatment" and "treating" refer to the inhibition of progression of a seizure disorder, such as the stabilization of symptoms discernible physiologically, such as by the stabilization of a physical parameter, or both.
[0058] As used herein, "therapy" refers to treatment. The therapeutic effect is achieved by partial or complete reduction, suppression, amelioration, or eradication of a disease state or symptom.
[0059] The terms "transfect" or "transformation" or "transduction" refer to the process by which exogenous nucleic acid is transferred or introduced into a host cell. A "transfected" or "transformed" or "transduced" cell is one that has been transfected, transformed or transduced with exogenous nucleic acid. The cell includes the primary subject cell and its progeny.
[0060] The term "specifically binds" refers to a molecule that recognizes and binds to a binding partner (eg, a protein or nucleic acid) preferentially over other molecules present in a sample.
[0061] Ranges: Throughout this disclosure, various aspects can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity, and should not be construed as an inflexible limitation on the scope of the disclosure. Thus, the description of a range should be considered to have specifically disclosed all possible subranges and individual numerical values within that range. For example, the description of a range such as 1-6 should be considered to have specifically disclosed subranges such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, etc., as well as individual numbers within that range, such as 1, 2, 2.7, 3, 4, 5, 5.3, and 6. As another example, a range such as 95-99% identity includes those with 95%, 96%, 97%, 98% or 99% identity, as well as subranges such as 96-99%, 96-98%, 96-97%, 97-99%, 97-98% and 98-99% identity. This is true regardless of the breadth of the range: all specifically recited ranges include the endpoints unless otherwise stated.
[0062] The term "gene editing system" or "genome editing system" refers to a system of one or more molecules that includes at least a nuclease (or nuclease domain) and a programmable nucleotide binding domain necessary and sufficient to direct and effect a nucleic acid modification (e.g., a single-stranded or double-stranded break) at a target sequence by the nuclease (or nuclease domain). In an embodiment, the gene editing system is a CRISPR system. In an embodiment, the gene editing system is a zinc finger nuclease (ZFN) system. In an embodiment, the gene editing system is a TALEN system. In an embodiment, the gene editing system is a meganuclease system. In an embodiment, the gene editing system modifies the expression of a first target, i.e., SCN1A. In an embodiment, the gene editing system further comprises a template nucleic acid as described herein. In an embodiment, one or more components of the gene editing system may be introduced into a cell as a nucleic acid encoding said one or more components. Without being bound by theory, expression of said one or more components constitutes, for example, a gene editing system in a cell.
[0063] The term "sequence-specific transcriptional regulation system" refers to a system of one or more molecules that includes at least a component that specifically binds to a DNA sequence and a component that is capable of regulating, e.g., increasing or decreasing, transcription. In an exemplary system, the sequence-specific transcriptional regulation system is derived from a genome editing system, and one or more activities of the genome editing system, such as one or more nuclease activities, are altered (e.g., abolished). In an exemplary sequence-specific transcriptional regulation system, the component that specifically binds to a DNA sequence is a zinc finger molecule. In an exemplary sequence-specific transcriptional regulation system, the component that specifically binds to a DNA sequence is a TALEN molecule. In an exemplary sequence-specific transcriptional regulation system, the component that specifically binds to a DNA sequence is a meganuclease. However, it will be appreciated that any molecule that can be engineered to bind to a target sequence of a sequence-specific molecule is useful as a component that specifically binds to a DNA sequence.
[0064] The term "CRISPR system", "Cas system" or "CRISPR / Cas system" refers to a set of molecules including an RNA-guided nuclease or other effector molecule (or a molecule derived from such a molecule) and a guide RNA molecule, which together are necessary and sufficient to direct and effect modification of a nucleic acid at a target sequence by an RNA-guided nuclease or other effector molecule. In one embodiment, the CRISPR system includes a guide RNA molecule and a Cas protein, such as a Cas9 protein. Such a system including Cas9 or a modified Cas9 molecule is referred to herein as a "Cas9 system" or a "CRISPR / Cas9 system". In one example, the guide RNA molecule and the Cas molecule can be complexed to form a ribonucleoprotein (RNP) complex. In some examples, the effector molecule can be modified to alter one or more activities of the wild-type molecule, such as one or more nuclease activities.
[0065] The terms "guide RNA", "guide RNA molecule", "gRNA molecule" or "gRNA" are used interchangeably and refer to a set of nucleic acid molecules that facilitate the specific orientation of an RNA-guided nuclease or other effector molecule (typically in a complex with a gRNA molecule) to a target sequence. As explained more fully below, a gRNA molecule may have several domains. In some embodiments, a gRNA molecule includes a targeting domain and interacts with a Cas molecule, such as Cas9, or another RNA-guided endonuclease, such as Cpf1. In some embodiments, a gRNA molecule includes a crRNA domain (including a targeting domain) and a tracr, for interacting with a Cas molecule, such as Cas9. In some embodiments, directing nuclease binding is achieved through hybridization of a portion of the gRNA to DNA (e.g., via the gRNA targeting domain) and binding of a portion of the gRNA to the RNA-guided nuclease or other effector molecule (e.g., via at least the gRNA tracr). In embodiments, the crRNA and tracr are provided on a single contiguous polynucleotide molecule, referred to herein as "single guide RNA," "sgRNA," or "single molecule DNA-targeting RNA," etc. In other embodiments, the crRNA and tracr are provided on separate polynucleotide molecules, referred to herein as "dual guide RNA," "dgRNA," "double molecule DNA-targeting RNA," etc., that are capable of associating with themselves, typically via hybridization. In some embodiments of the dgRNA, the crRNA and tracr are linked by a non-nucleotide chemical linker.
[0066] The term "targeting domain" as used herein in connection with gRNA is the portion of a gRNA molecule that recognizes, e.g., is complementary to, a target sequence, e.g., a target sequence within a nucleic acid of a cell.
[0067] The term "crRNA" as used herein in relation to a gRNA molecule is the portion of the gRNA that contains the targeting domain. In embodiments, the crRNA contains a region that interacts with tracr to form a flagpole region. In some embodiments, the crRNA can directly interact with an RNA-guided endonuclease, such as a Cas protein (e.g., Cpf1), without the tracrRNA.
[0068] The term "target sequence" refers to a nucleic acid sequence that is complementary, e.g., fully complementary, to a gRNA targeting domain, and also refers to a nucleic acid sequence that is recognized (e.g., bound) by a DNA-binding component of a sequence-specific transcriptional regulatory system, such as a sequence recognized by a zinc finger motif or a sequence recognized by a TALEN or homing endonuclease. In an embodiment, the target sequence is disposed on genomic DNA. In one embodiment, particularly in embodiments involving an RNA-guided endonuclease system, the target sequence is adjacent to a protospacer adjacent motif (PAM) sequence (either on the same or complementary strand of DNA) that is recognized by a protein with nuclease activity or other effector activity, such as the PAM sequence recognized by Cas9. The sequence and length of the PAM may depend on the Cas9 protein used. Non-limiting examples of PAM sequences include 5'-NGG-3', 5'-NGGNG-3', 5'-NG-3', 5'-NAAAAN-3', 5'-NNAAAAW-3', 5'-NNNNACA-3', 5'-GNNNCNNA-3', 5'-NNGRRT-3', 5'-NNGRRN-3', and 5'-NNNNGATT-3' (where N represents any nucleotide, R represents A or G, and W represents A or T).
[0069] In embodiments, the target sequence is within an exon sequence of SCN1A. In other embodiments, the target sequence is within an intron sequence of SCN1A. In yet other embodiments, the target sequence is in a nucleic acid region flanking the SCN1A gene. In still further embodiments, the target sequence overlaps one or more of an intron sequence of SCN1A, an exon sequence of SCN1A, and a nucleic acid region flanking SCN1A.
[0070] The term "flagpole" as used herein in reference to a gRNA molecule refers to the portion of the gRNA where the crRNA and tracr bind or hybridize to each other.
[0071] The term "tracr" or "tracrRNA" as used herein in connection with a gRNA molecule refers to the portion of the gRNA that binds to a nuclease molecule or other effector molecule. In embodiments, tracr comprises a nucleic acid sequence that specifically binds to Cas9. In embodiments, tracr comprises a nucleic acid sequence that forms part of a flagpole.
[0072] The term "Cas" refers to the RNA-guided nuclease of CRISPR system, which is necessary and sufficient to direct and carry out the modification of nucleic acid at target sequence together with guide RNA molecule.A non-limiting example is the Cas molecule from type II CRISPR system, such as Cas9 molecule.Another non-limiting example is the Cas molecule from type V CRISPR system, such as Cpf1 molecule.
[0073] The terms "Cas9" and "Cas9 molecule" refer to an enzyme from a bacterial type II CRISPR / Cas system involved in DNA cleavage. In embodiments, Cas9 includes wild-type proteins, mutant proteins (including non-catalytic proteins), and functional fragments thereof. Non-limiting examples of Cas9 sequences are known in the art and provided herein. In some embodiments, Cas9 refers to a Cas9 sequence that contains at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% homology to any Cas9 sequence, e.g., wild-type, mutant, non-catalytic, or functional fragments known in the art or disclosed herein; differs in no more than 1%, 2%, 5%, 10%, 15%, 20%, 30%, or 40% of the amino acid residues when compared to any Cas9 sequence; differs by at least 1, 2, 5, 10, or 20 amino acids from any Cas9 sequence, but not more than 100, 80, 70, 60, 50, 40, or 30 amino acids; or is identical to any Cas9 sequence. In a preferred embodiment, the Cas9 molecule is obtained from Staphylococcus aureus. In another embodiment, the Cas9 molecule is a mutant protein that has no detectable nuclease activity, hi yet another embodiment, the Cas9 molecule is a mutant protein fused to a transcriptional activator.
[0074] The terms "Cpf1" and "Cpf1 molecule" refer to an enzyme from a bacterial type V CRISPR / Cas system involved in DNA cleavage. In embodiments, Cpf1 includes wild-type proteins, mutant proteins (including non-catalytic proteins) and functional fragments thereof. Non-limiting examples of Cpf1 sequences are known in the art. In some embodiments, Cpf1 refers to a Cpf1 sequence that contains at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% homology to any Cpf1 sequence, such as a wild-type, mutant, non-catalytic or functional fragment thereof known in the art; differs in no more than 1%, 2%, 5%, 10%, 15%, 20%, 30% or 40% of amino acid residues when compared to any Cpf1; differs in at least 1, 2, 5, 10 or 20 amino acids from any Cpf1, but not more than 100, 80, 70, 60, 50, 40 or 30 amino acids; or is identical to any Cpf1 sequence. Unlike other Cas proteins (e.g., Cas9), Cpf1 does not require tracrRNA for activity and is capable of binding and cleaving genomic target sequences with only the crRNA polynucleotide. Thus, in some embodiments utilizing Cpf1 to edit a target sequence, the gRNA may lack a tracrRNA portion. The term "complementary" as used in connection with a nucleic acid refers to base pairing, A to T or U and G to C. The term complementary can also refer to nucleic acid molecules that are fully complementary, i.e., that form A:T or U pairs and G:C pairs across the entire reference sequence, as well as molecules that are at least about 80%, 85%, 90%, 95%, or 99% complementary.
[0075] The term "gene" or "gene sequence" is intended to refer to a genetic sequence, e.g., a nucleic acid sequence. The term "gene" is intended to encompass a complete gene sequence or a partial gene sequence. The term "gene" refers to a sequence that codes for a protein or polypeptide or a sequence that does not code for a protein or polypeptide, such as a regulatory sequence, a leader sequence, a signal sequence, introns, or other non-protein coding sequence.
[0076] The term "intron" refers to an intragenic nucleic acid sequence that is non-coding for the protein expressed from said gene. Intron sequences can be transcribed from DNA into RNA, but can be removed before a protein is expressed.
[0077] The term "exon" refers to an intragenic nucleic acid sequence that codes for a protein expressed from said gene.
[0078] The term "intron-exon junction" when used in connection with a gene editing system or a sequence-specific transcriptional regulation system or a gRNA molecule refers to a sequence that includes an exon nucleotide and an intron nucleotide. In an exemplary embodiment, the intron-exon junction is a target sequence for a gRNA, and when recognized by a CRISPR system that includes a gRNA that includes a targeting domain that is complementary to an intron-exon junction target sequence, the CRISPR system modifies a target sequence between two nucleotides of an intron or nearby. In another exemplary embodiment, the intron-exon junction is a target sequence for a gRNA, and when recognized by a CRISPR system that includes a gRNA that includes a targeting domain that is complementary to an intron-exon junction target sequence, the CRISPR system modifies a target sequence between two nucleotides of an exon or nearby. In other exemplary embodiments, an intron-exon junction is a target sequence for a gRNA, and when recognized by a CRISPR system comprising a gRNA that comprises a targeting domain that is complementary to the intron-exon junction target sequence, the CRISPR system modifies at or near the target sequence between the exon nucleotides and the intron nucleotides.
[0079] The term "AAV" is an abbreviation for adeno-associated virus and may be used to refer to the virus itself or its derivatives. The term encompasses all serotypes, subtypes, both naturally occurring and recombinant forms, unless otherwise required. The abbreviation "rAAV" refers to recombinant adeno-associated virus, also called recombinant AAV vector (or "rAAV vector"). The term "AAV" includes AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, rhlO and their hybrids, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV and ovine AAV. The genomic sequences of the various serotypes of AAV and the sequences of the native terminal repeats (TRs), Rep proteins and capsid subunits are known in the art. Such sequences may be found in the literature or in public databases such as GenBank. As used herein, "rAAV vector" refers to an AAV vector that contains a polynucleotide sequence that is not of AAV origin (i.e., a polynucleotide heterologous to AAV), typically a sequence of interest for genetic transformation of a cell. Generally, the heterologous polynucleotide is flanked by at least one, and generally two, AAV inverted terminal repeats (ITRs). rAAV vectors can be either single-stranded (ssAAV) or self-complementary (scAAV). "AAV virus" or "AAV viral particle" refers to a viral particle that is composed of at least one AAV capsid protein and the encapsidated polynucleotide rAAV vector. When a particle contains a heterologous polynucleotide (i.e., a polynucleotide other than the wild-type AAV genome, such as a transgene delivered to a mammalian cell), it is typically referred to as an "rAAV vector particle" or simply an "rAAV particle". Thus, since such vectors are contained within the rAAV particle, the production of the rAAV particle necessarily includes the production of the rAAV vector.
[0080] As used herein, the terms "treat", "treatment", "therapy" and the like refer to alleviating, delaying or slowing the progression of a disease or disorder, preventing, attenuating, reducing the effects or symptoms of a disease or disorder, preventing the onset of a disease or disorder, inhibiting a disease or disorder, or reversing the onset of a disease or disorder. The methods of the present disclosure may be used in any mammal. Exemplary mammals include, but are not limited to, rats, cats, dogs, horses, cows, sheep, pigs, and more preferably humans. Therapeutic benefit includes eradication or reversal of the underlying disorder being treated. Therapeutic benefit is also achieved by observing an improvement in a subject as a result of eradication or reversal of one or more of the physiological symptoms associated with the underlying disorder, even though the subject may still suffer from the underlying disorder. In some cases, for preventative benefit, a therapeutic agent may be administered to a subject at risk of developing a particular disease or reporting one or more of the physiological symptoms of a disease, even though a diagnosis of the disease may not have been made. The methods of the present disclosure may be used in any mammal. In some cases, treatment may result in a reduction or cessation of symptoms (e.g., a reduction in the frequency, duration and / or severity of attacks). A prophylactic effect includes delaying or eliminating the appearance of a disease or condition, delaying or eliminating the onset of symptoms of a disease or condition, slowing, halting or reversing the progression of a disease or condition, or any combination thereof.
[0081] A "fragment" of a nucleotide or peptide sequence means a sequence that is shorter than a reference sequence or the "full-length" sequence.
[0082] A "variant" of a molecule refers to an allelic variation of such a sequence, ie, a sequence substantially similar in structure and biological activity to either the entire molecule or a fragment thereof.
[0083] A "functional fragment" of a DNA or protein sequence refers to a fragment that retains a biological activity (either functional or structural) substantially similar to the biological activity of the full-length DNA or protein sequence. The biological activity of a DNA sequence may be its ability to affect expression in a known manner ascribed to the full-length sequence.
[0084] The term "in vivo" refers to events that take place within the body of a subject.
[0085] The term "in vitro" refers to events that occur outside a subject's body. For example, an in vitro assay includes any assay that is performed outside a subject. An in vitro assay includes cell-based assays in which live or dead cells are employed. An in vitro assay includes cell-free assays in which intact cells are not employed.
[0086] Generally, "sequence identity" or "sequence homology", which can be used interchangeably, refers to the exact nucleotide-to-nucleotide or amino acid-to-amino acid correspondence of two polynucleotide or polypeptide sequences, respectively. Typically, techniques for determining sequence identity involve comparing two nucleotide or amino acid sequences to determine their percent identity. Sequence comparison, such as for purposes of assessing identity, can be performed by any suitable alignment algorithm, including, but not limited to, the Needleman-Wunsch algorithm (see, for example, the EMBOSS Needle aligner available at www.ebi.ac.uk / Tools / psa / emboss_needle / , optionally with default settings), the BLAST algorithm (see, for example, the BLAST alignment tool available at blast.ncbi.nlm.nih.gov / Blast.cgi, optionally with default settings), and the Smith-Waterman algorithm (see, for example, the EMBOSS Water aligner available at www.ebi.ac.uk / Tools / psa / emboss_water / , optionally with default settings). Optimal alignment can be evaluated using any suitable parameters of the selected algorithm, including default parameters. "Percentage identity", also called "percentage homology" between two sequences, can be calculated by dividing the number of exact matches between two optimally aligned sequences by the length of the reference sequence and multiplying by 100. Percentage identity can also be determined by comparing sequence information using the latest BLAST computer program, including version 2.2.9, available from the National Institutes of Health.The BLAST program is based on the alignment method of Karlin and Altschul, Proc. Natl. Acad Sci. USA 87:2264-2268 (1990), as discussed in Altschul, et al., J. Mol. Biol. 215:403-410 (1990); Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90:5873-5877 (1993); and Altschul et al., Nucleic Acids Res. 25:3389-3402 (1997). Briefly, the BLAST program defines identity by dividing the number of identical aligned symbols (i.e., nucleotides or amino acids) by the total number of symbols in the shorter of the two sequences. The program can be used to determine percent identity over the entire length of the sequences being compared. Default parameters are provided to optimize searches with short query sequences, for example using the blastp program. The program also allows for the use of a SEG filter to mask and exclude segments of the query sequence as determined by the SEG program of Wootton and Federhen, Computers and Chemistry 17:149-163 (1993). High sequence identity generally includes a range of about 80% to 100% sequence identity and integer values therebetween.
[0103] As used herein, "engineered" in reference to a protein refers to a protein that is not naturally occurring, including, but not limited to, a protein derived from a naturally occurring protein or a naturally occurring protein that has been modified or reprogrammed to have a particular property.
[0087] As used herein, "synthetic" and "artificial" are used interchangeably to refer to proteins or domains thereof that have low sequence identity (e.g., less than 50% sequence identity) to naturally occurring human proteins.
[0088] As used herein, "transcription factor" or "TF" refers to a DNA-binding protein or a non-naturally occurring transcriptional regulator that has been modified or reprogrammed to bind to a specific target binding site and / or to contain a modified or replaced transcriptional effector domain.
[0089] As used herein, a "DNA-binding domain" can be used to refer, individually or collectively, to one or more DNA-binding motifs, such as a Cas molecule, a transcription activator-like (TAL) effector, a zinc finger protein, a basic helix-loop-helix (bHLH) motif, etc., as part of a DNA-binding protein.
[0090] The terms "transcription activation domain", "transcriptional activation domain", "transactivation domain", "transactivation domain" and "TAD" are used interchangeably herein and refer to a domain of a protein that, in conjunction with a DNA-binding domain, can activate transcription from a promoter by contacting the transcriptional machinery (e.g., general transcription factors and / or RNA polymerase) either directly or through other proteins known as coactivators. "Q-rich TAD" means a glutamine-rich transactivation domain. "P-rich TAD" means a proline-rich transactivation domain. "Acidic TAD" means a transactivation domain rich in acidic amino acids.
[0091] Transcription factors Transcription factors (TFs) are proteins that bind to specific sequences in the genome to control the expression of genes. The present disclosure provides novel engineered transcription factors that contain a DNA binding domain (DBD) and at least three transcriptional activation domains (TADs). In some embodiments, the TADs can be identical. In some embodiments, the TADs can be different.
[0092] Transcriptional activation domains found in various proteins have been grouped into categories based on similar structural features. Types of transcriptional activation domains that can be used in the present application include acidic transcriptional activation domains, proline-rich (P-rich) transcriptional activation domains, and glutamine-rich (Q-rich) transcriptional activation domains and / or fragments thereof. Examples of acidic transcriptional activation domains include VP16, VP64, VP7, ATF6, TFE3, ATF6-11, or fragments thereof. Examples of proline-rich activation domains include TFAP2-P, Oct2-P, or fragments thereof. Examples of glutamine-rich activation domains include amino acid residues Oct2-Q, SP1-Q1, or fragments thereof. The amino acid sequences of each of the above-mentioned regions are disclosed in Table 1 of the present application.
[0093] In certain embodiments, the TADs comprise i) at least one acidic TAD and at least one Q-rich TAD, or ii) at least one acidic TAD and at least one P-rich TAD.
[0094] In some embodiments, the acidic TAD comprises one or more copies of a TAD selected from VP16, VP64, VP7, ATF6, TFE3, ATF6-11, or fragments thereof. In some embodiments, the Q-rich TAD comprises one or more copies of a TAD selected from Oct2-Q, SP1-Q1, or fragments thereof. In some embodiments, the P-rich TAD comprises one or more copies of a TAD selected from TFAP2-P, Oct2-P, or fragments thereof.
[0095] In some embodiments, the TAD is selected from the group consisting of full length VP64, one or more repeats of VP16, full length or partial p65, full length RTA, full length VP7, full length or partial TEF3, full length or partial ATF6-11, full length or partial ATF6 acidic, full length or partial SP1-rich Q, full length or partial Oct2-rich Q, full length or partial Oct2-rich P, full length or partial TFAP2-rich P, an active fragment of VP64, a transcriptionally active fragment of p65, a transcriptionally active fragment of RTA, an active fragment of VP7, an active fragment of TEF3, an active fragment of ATF6-11, an active fragment of ATF6 acidic, an active fragment of SP1-rich Q, an active fragment of Oct2-rich Q, an active fragment of Oct2-rich P, an active fragment of TFAP2-rich P, and any combination thereof.
[0096] In some embodiments, the combined length of the at least three TADs is less than 2000 aa, less than 1500 aa, less than 1000 aa, less than 750 aa, less than 500 aa, less than 300 aa, less than 250 aa, less than 200 aa, or less than 150 aa.
[0097] In some embodiments, the transcription factor comprises a sequence of SEQ ID NO: 1-28 or a sequence having at least 90% or at least 95% sequence identity thereto. In certain embodiments, the transcription factor comprises a sequence having any one of SEQ ID NO: 1-18.
[0098] In some embodiments, the nucleic acid sequence encoding the transcription factor comprises a sequence of SEQ ID NO: 30-47 or a sequence having at least 80%, at least 85%, at least 90%, or at least 95% sequence identity thereto.
[0099] In some embodiments, the DNA binding domain is linked to any one of the transcription factors with or without a linker. In certain embodiments, the DNA binding domain is linked to the N-terminus of the transcription factor directly or via a linker. In certain embodiments, the DNA binding domain is linked to the C-terminus of the transcription factor directly or via a linker.
[0100] In some embodiments, the linker has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 60, 70, 75, 80, 90 or 100 amino acids or 1-5, 1-10, 1-20, 1-30, 1-40, 1-50, 1-75, 1-100, 5-10, 5-20, 5-30, 5-40, 5-50, 5-75, 5-100, 10-20, 10-30, 10-40, 10-50, 10-75, 10-100, 20-30, 20-40, 20-50, 20-75 or 20-100 amino acids.
[0101] Suitable linkers can be flexible, cleavable, non-cleavable, hydrophilic and / or hydrophobic. In certain embodiments, the linker comprises multiple glycine and / or serine residues. Examples of glycine / serine peptide linkers include [GS]n, [GGGS]n (SEQ ID NO: 74), [GGGGS]n (SEQ ID NO: 75), [GGSG]n (SEQ ID NO: 76), where n is an integer equal to or greater than 1. In certain embodiments, the linker comprises or consists of GGSGGGSG (SEQ ID NO: 59) or GGSGGGSGGGSGGGSG (SEQ ID NO: 60).
[0102] Table 1 below provides the sequences of the transcription activation domains.
[0103] [Table 1]
[0104] [Table 2]
[0105] [Table 3]
[0106] [Table 4]
[0107]
Table 5
[0108]
Table 6
[0109]
Table 7
[0110]
Table 8
[0111]
Table 9
[0112]
Table 10
[0113]
Table 11
[0114]
Table 12
[0115]
Table 13
[0116]
Table 14
[0117] [Table 15]
[0118] [Table 16]
[0119] [Table 17]
[0120] [Table 18]
[0121] [Table 19]
[0122] DNA binding domain (DBD) The transcription factors provided herein include any suitable DBD that binds to a target site of interest (e.g., a target site that, when bound by a transcription factor provided herein, results in upregulation of a target gene).
[0123] TALEN In some embodiments, the DBD can be a TALEN. TALENs are artificially produced by fusing a TAL effector DNA binding domain to a DNA cleavage domain. Transcription activator-like effects (TALEs) can be engineered to bind to any desired DNA sequence. They can then be introduced into cells, where they can be used for genome editing. Boch, Nature Biotech. 29:135-6 (2011); and Boch et al. Science 326:1509-12 (2009); Moscou et al. Science 326:3501 (2009).
[0124] TALEs are proteins secreted by Xanthomonas bacteria. The DNA-binding domain contains a repeated and highly conserved 33-34 amino acid sequence, with the exception of the 12th and 13th amino acids. These two positions are highly variable and show a strong correlation with specific nucleotide recognition. Therefore, they can be engineered to bind to desired DNA sequences.
[0125] TALEs specific to the sequences described herein can be constructed using any method known in the art, including various schemes using modular components. Zhang et al. Nature Biotech. 29:149-53 (2011); Geibler et al. PLoS ONE 6:e19509 (2011); U.S. Patent No. 8,420,782; U.S. Patent No. 8,470,973 (the contents of which are incorporated herein by reference in their entirety).
[0126] CRISPR gene editing system In some embodiments, the DBD can be a gRNA / Cas complex.
[0127] As described more fully below, gRNA molecules may have several domains. In some embodiments, gRNA molecules include a targeting domain and interact with a Cas molecule, such as Cas9. In some embodiments, gRNA molecules include a crRNA domain (including a targeting domain) and a tracr. In embodiments, the crRNA and tracr are provided on a single contiguous polynucleotide molecule. In other embodiments, the crRNA and tracr are provided on separate polynucleotide molecules capable of associating with themselves, such as through non-covalent hybridization. gRNA molecules used as components of CRISPR systems are useful for modifying (e.g., modifying sequences of) DNA at or near a target site. Such modifications include, for example, insertions and or deletions that result in reduced or eliminated expression of a functional product of a gene that includes the target site. Such modifications may also include upregulation of expression of a functional product of a gene that includes a target site, for example, when the Cas9 molecule lacks nuclease activity but is fused to one or more transcription factors. In some embodiments, separate gRNA molecules and CRISPR systems are used to upregulate expression of a functional product of a gene that contains a target site, the use of which, among others, is described more fully below.
[0128] In one embodiment, the unimolecular or sgRNA preferably comprises, in the 5' to 3' direction: crRNA (comprising a targeting domain complementary to the target sequence and a region forming part of the flagpole (i.e. the crRNA flagpole region)); optionally a loop; and tracr (comprising a domain complementary to the crRNA flagpole region and a domain that additionally binds to a nuclease molecule or other effector molecule, e.g. a Cas molecule, such as a Cas9 molecule), in the following format (in the 5' to 3' direction): [targeting domain]-[crRNA flagpole region]-[optional first flagpole extension]-[optional loop]-[optional first tracr extension]-[tracr flagpole region]-[tracr nuclease binding domain] In embodiments, the tracr nuclease binding domain binds to a Cas protein, such as a Cas9 protein.
[0129] In one embodiment, the bimolecular or dgRNA comprises two polynucleotides; a first, preferably in the 5'→3' direction: crRNA (containing a targeting domain complementary to the target sequence and a region forming part of the flagpole); and a second, preferably in the 5'→3' direction: tracr (containing a domain complementary to the crRNA flagpole region and a domain that additionally binds to a nuclease molecule or other effector molecule, e.g. a Cas molecule, such as a Cas9 molecule), in the following format (in the 5'→3' direction): Polynucleotide 1 (crRNA): [targeting domain] - [crRNA flagpole region] - [optional first flagpole extension] - [optional second flagpole extension] Polynucleotide 2 (tracr): [any first tracr extension]-[tracr flagpole region]-[tracr nuclease binding domain] It is possible to take.
[0130] In embodiments, the dgRNA comprises two polynucleotides covalently linked by a non-nucleotidic linker, for example, as described in He et al., ChemBioChem 17:1809-1812 (2016). In some embodiments, a chemical reaction is used to link the two polynucleotides, for example, using a copper(I)-catalyzed alkyne-azide cycloaddition (CuAAC) reaction (see He et al., ChemBioChem 17:1809-1812 (2016)) or via strain-promoted azide-alkyne cycloaddition (SPAAC) (see U.S. Patent Application Publication No. 2016 / 0215275 A1), both of which are incorporated herein by reference in their entireties. In another embodiment, the two polynucleotides are covalently linked via a thio-ether linker, which can occur, for example, by reaction between thiol and maleimide functional groups or by reaction between other functional groups (see, for example, U.S. Patent Application Publication No. 2016 / 0215275 A1). In yet other embodiments, the non-nucleotidic linker can include carbamates, ethers, esters, amides, imines, amidines, aminotridines, hydrozones, disulfides, thioesters, phosphorothioates, phosphorodithioates, sulfonamides, sulfonates, sulfones, sulfoxides, ureas, thioureas, hydrazides, oximes, photolabile linkages or CC bond-forming groups, such as Diels-Alder cycloaddition pairs and / or ring-closing metathesis pairs and / or Michael reaction pairs (see, for example, WO 2016 / 18745 A1, which is incorporated by reference in its entirety).
[0131] In some embodiments, the flagpole, e.g., the crRNA flagpole region, comprises in the 5' to 3' direction: AGUACUCUG.
[0132] In some embodiments, the loop comprises in the 5' to 3' direction: GAAA.
[0133] In some embodiments, tracr comprises from 5' to 3' direction: AGUACUCUGGAAACAGAAUCUACUCUAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGAUUUU (SEQ ID NO: 79).
[0134] In some aspects, the gRNA may also include an additional U nucleic acid at the 3' end. For example, the gRNA may include an additional 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 U nucleic acid at the 3' end. In one embodiment, the gRNA includes an additional 4 to 5 U nucleic acid at the 3' end. In the case of a dgRNA, one or more of the polynucleotides of the dgRNA (e.g., a polynucleotide comprising a targeting domain and a polynucleotide comprising tracr) may include an additional U nucleic acid at the 3' end. For example, in the case of a dgRNA, one or more of the polynucleotides of the dgRNA (e.g., a polynucleotide comprising a targeting domain and a polynucleotide comprising tracr) may include an additional 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 U nucleic acid at the 3' end. In one embodiment, in the case of a dgRNA, one or more of the polynucleotides of the dgRNA (e.g., a polynucleotide comprising a targeting domain and a polynucleotide comprising tracr) may include an additional 4 to 5 U nucleic acid at the 3' end. In one embodiment of the dgRNA, only the polynucleotide comprising tracr comprises an additional U nucleic acid, for example, 4-5 U nucleic acid. In one embodiment of the dgRNA, only the polynucleotide comprising the targeting domain comprises an additional U nucleic acid. In one embodiment of the dgRNA, both the polynucleotide comprising the targeting domain and the polynucleotide comprising tracr comprise an additional U nucleic acid, for example, 4-5 U nucleic acid.
[0135] In some aspects, the gRNA may also include an additional A nucleic acid at the 3' end. For example, the gRNA may include an additional 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 A nucleic acid at the 3' end. In one embodiment, the gRNA includes an additional 4 A nucleic acid at the 3' end. In the case of a dgRNA, one or more of the polynucleotides of the dgRNA (e.g., a polynucleotide comprising a targeting domain and a polynucleotide comprising tracr) may include an additional A nucleic acid at the 3' end. For example, in the case of a dgRNA, one or more of the polynucleotides of the dgRNA (e.g., a polynucleotide comprising a targeting domain and a polynucleotide comprising tracr) may include an additional 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 A nucleic acid at the 3' end. In one embodiment, in the case of a dgRNA, one or more of the polynucleotides of the dgRNA (e.g., a polynucleotide comprising a targeting domain and a polynucleotide comprising tracr) may include an additional 4 A nucleic acid at the 3' end. In one embodiment of the dgRNA, only the polynucleotide comprising tracr comprises an additional A nucleic acid, for example, 4A nucleic acid. In one embodiment of the dgRNA, only the polynucleotide comprising the targeting domain comprises an additional A nucleic acid. In one embodiment of the dgRNA, both the polynucleotide comprising the targeting domain and the polynucleotide comprising tracr comprise an additional U nucleic acid, for example, 4A nucleic acid.
[0136] In embodiments, one or more of the polynucleotides of the gRNA molecule may include a cap at the 5' end.
[0137] In one embodiment, the unimolecular or sgRNA comprises, preferably in the 5' to 3' direction: crRNA (a targeting domain complementary to the target sequence; crRNA flagpole region; a first flagpole extension); a loop; a first tracr extension (containing a domain complementary to at least a portion of the first flagpole extension); and tracr (containing a domain complementary to the crRNA flagpole region and a domain additionally binding to a Cas9 molecule). In some aspects, the targeting domain comprises a targeting domain sequence as described herein or comprises a targeting domain comprising or consisting of 17, 18, 19, 20 (preferably 20) consecutive nucleotides of the targeting domain sequence, such as 17, 18, 19 or 20 (preferably 20) consecutive nucleotides 3' of the targeting domain sequence. In an embodiment, the 17, 18, 19, 20 (preferably 20) consecutive nucleotides of the targeting domain sequence are 17, 18, 19, 20 (preferably 20) consecutive nucleotides 3' of the targeting domain sequence. In embodiments, the 17, 18, 19, 20 (preferably 20) contiguous nucleotides of the targeting domain sequence are 17, 18, 19, 20 (preferably 20) contiguous nucleotides 5' to the targeting domain sequence.
[0138] In the embodiment comprising the first flagpole extension and / or the first tracr extension, the flagpole, loop and tracr sequence can be as described above. In general, any first flagpole extension and first tracr extension can be employed as long as they are complementary. In an embodiment, the first flagpole extension and the first tracr extension are composed of 3, 4, 5, 6, 7, 8, 9, 10 or more complementary nucleotides.
[0139] In some embodiments, the first flag pole extension comprises in the 5' to 3' direction: UGCUG. In some embodiments, the first flag pole extension consists of SEQ ID NO:80.
[0140] In some embodiments, the first tracr extension comprises in the 5' to 3' direction: CAGCA. In some embodiments, the first tracr extension consists of SEQ ID NO:81.
[0141] In one embodiment, the dgRNA comprises two nucleic acid molecules. In some aspects, the dgRNA comprises, preferably in the 5'→3' direction: a first nucleic acid containing a targeting domain complementary to a target sequence; a crRNA flagpole region; optionally a first flagpole extension; and optionally a second flagpole extension; and preferably in the 5'→3' direction: a second nucleic acid (which may be referred to herein as tracr and includes at least a domain that binds to a Cas molecule, such as a Cas9 molecule) comprising a tracr (containing a domain complementary to the crRNA flagpole region and a domain that additionally binds to a Cas molecule, such as Cas9). The second nucleic acid may additionally include an additional U nucleic acid at the 3' end (e.g., 3' side of tracr). For example, tracr may include an additional 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 U nucleic acid at the 3' end (e.g., 3' side of tracr). The second nucleic acid may additionally or alternatively include an additional A nucleic acid at the 3' end (e.g., 3' to tracr). For example, tracr may include an additional 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 A nucleic acid at the 3' end (e.g., 3' to tracr).
[0142] In embodiments involving dgRNA, the crRNA flagpole region, the optional first flagpole extension, the optional first tracr extension and the tracr sequence may be as described above.
[0143] In some aspects, the optional second flagpole extension comprises in the 5' to 3' direction: UUUUG. In embodiments, the 3' 1, 2, 3, 4 or 5 nucleotides, the 5' 1, 2, 3, 4 or 5 nucleotides, or both the 3' and 5' 1, 2, 3, 4 or 5 nucleotides of the gRNA molecule (and in the case of a dgRNA molecule, the polynucleotides that make up the targeting domain and / or the polynucleotides that make up the tracr) are modified nucleic acids, as described more fully below.
[0144] Domains are briefly discussed below.
[0145] Guidance on the selection of targeting domains can be found, for example, in Fu Y el al. NAT BIOTECHNOL (doi:10.1038 / nbt.2808) (2014) and Sternberg SH el al. NATURE (doi:10.1038 / naturel3011) (2014).
[0146] The targeting domain comprises a nucleotide sequence that is complementary, e.g., at least 80, 85, 90, 95, or 99% complementary, or e.g., completely complementary, to a target sequence on a target nucleic acid. The targeting domain is a part of an RNA molecule and therefore will contain the base uracil (U), whereas any DNA encoding a gRNA molecule will contain the base thymine (T). Without wishing to be bound by theory, it is believed that the complementarity of the targeting domain with the target sequence contributes to the specificity of the interaction of the gRNA molecule / Cas9 molecule complex with the target nucleic acid. In the targeting domain and target sequence pair, it is understood that the uracil base in the targeting domain pairs with the adenine base in the target sequence.
[0147] In one embodiment, the targeting domain is 5 to 50, such as 10 to 40, such as 10 to 30, such as 15 to 30, such as 15 to 25 nucleotides in length. In one embodiment, the targeting domain is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotides in length. In one embodiment, the targeting domain is 16 nucleotides in length. In one embodiment, the targeting domain is 17 nucleotides in length. In one embodiment, the targeting domain is 18 nucleotides in length. In one embodiment, the targeting domain is 19 nucleotides in length. In one embodiment, the targeting domain is 20 nucleotides in length. In one embodiment, the targeting domain is 21 nucleotides in length. In one embodiment, the targeting domain is 22 nucleotides in length. In one embodiment, the targeting domain is 23 nucleotides in length. In one embodiment, the targeting domain is 24 nucleotides in length. In one embodiment, the targeting domain is 25 nucleotides in length. In embodiments, said 16, 17, 18, 19 or 20 nucleotides comprise 16, 17, 18, 19 or 20 nucleotides 5' from a targeting domain listed in Table 2. In embodiments, said 16, 17, 18, 19 or 20 nucleotides comprise 16, 17, 18, 19 or 20 nucleotides 3' from the targeting domain. In embodiments, said 16, 17, 18, 19 or 20 nucleotides consist of 16, 17, 18, 19 or 20 nucleotides 3' from the targeting domain.
[0148] In some embodiments, the Cas molecule in the gRNA / Cas complex is a class 1 Cas nuclease. In some embodiments, the Cas molecule is a class 2 Cas nuclease. See, e.g., Makarova et al. Nat Rev Microbiol, 13(11):722-36(2015); Shmakov et al. Molecular Cell, 60:385-397(2015). Class 2 Cas molecules can be single protein endonucleases. In some embodiments, class 2 Cas molecules are from type II, V or VI CRISPR / Cas systems and can be single protein endonucleases. Non-limiting examples of class 2 Cas molecules include Cas9, Cpf1, C2c1, C2c2 and C2c3 proteins. See, e.g., Yang et al. Cell, 167(7):1814-28 (2016); Zetsche et al. Cell, 163:1-13 (2015). In some embodiments, the Cas molecule is a S. aureus Cas9 molecule.
[0149] In some embodiments, the Cas molecule is a Cas9 molecule or a fragment or variant thereof, such as a non-catalytic variant. Cas9 molecules of various species can be used in the methods and compositions described herein. Although the S. aureus Cas9 molecule is the subject of much of the disclosure herein, Cas9 molecules derived from or based on the Cas9 proteins of other species listed herein can be used as well. In other words, other Cas9 molecules, such as S. thermophilus, Staphylococcus pyrogenes, and / or Neisseria meningitidis Cas9 molecules, can be used in the systems, methods, and compositions described herein.
[0150] In some embodiments, the Cas9 molecule is a high-fidelity mutant that carries modifications designed to reduce non-specific DNA contacts. See, e.g., Kleinstiver et al. Nature 529(7587):490-95 (2016); Slaymaker et al. Science, 351(6268):84-88 (2016); Tsai et al. Nat. Biotech. 32:569-577 (2014). In some embodiments, the high-fidelity Cas9 retains on-target activity comparable to wild-type Cas9. In some embodiments, the high-fidelity Cas9 reduces off-target activity by at least about 50%, 60%, 70%, 80%, 90%, 95%, or 99% compared to wild-type Cas9, e.g., as measured by genome-wide cleavage capture and targeted sequencing. In some embodiments, the high-fidelity Cas9 has undetectable off-target activity as measured by genome-wide cleavage capture and targeted sequencing, hi some embodiments, the high-fidelity Cas9 is Streptococcus aureus Cas9.
[0151] Additional Cas9 species include Acidovorax avenae, Actinobacillus pleuropneumoniae, Actinobacillus succinogenes, Actinobacillus suis, Actinomyces species, cycliphilus denitrificans, Aminomonas paucivorans, Bacillus cereus, Bacillus smithii, Bacillus thuringiensis, and Actinobacillus spp. thuringiensis, Bacteroides spp., Blastopyrellula marina, Bradyrhizobium spp., Brevibacillus latemsporus, Campylobacter coli, Campylobacter jejuni, Campylobacter lari, Candidatus Puniceispirillum, Clostridium cellulolyticum, Clostridium perfringens, Corynebacterium acoruscens accolens, Corynebacterium diphtheria, Corynebacterium matruchotii, Dinoroseobacter sliibae, Eubacterium dolichum, Gammaproteobacteriaproteobacterium, Gluconacetobacter diazotrophicus, Haemophilus parainfluenzae, Haemophilus sputorum, Helicobacter canadensis, Helicobacter cinaedi, Helicobacter mustelae, Ilyobacler polytropus, Kingella kingae, Lactobacillus crispatus, Listeria ivanovii, Listeria monocytogenes monocytogenes, Listeriaceae bacteria, Methylocystis spp., Methylosinus trichosporium, Mobiluncus mulieris, Neisseria bacilliformis, Neisseria cinerea, Neisseria flavescens, Neisseria lactamica, Neisseria spp., Neisseria wadsworthii, Nitrosomonas spp. and Parvibaculum lavamentivorans, Pasteurella multocida multocida, Phascolarctobacterium succinatutens, Ralstonia syzygii, Rhodopseudomonas palustrispalustris, Rhodovulum spp., Simonsiella muelleri, Sphingomonas spp., Sporolactobacillus vineae, Staphylococcus lugdunensis, Streptococcus spp., Subdoligranulum spp., Tislrella mobilis, Treponema spp. or Verminephrobacter eiseniae.
[0152] A Cas9 molecule, as that term is used herein, refers to a molecule that is capable of interacting with a gRNA molecule (e.g., a tracr domain sequence) and localizes (e.g., targets or homes to) a site that includes a target sequence and a PAM sequence in concert with the gRNA molecule.
[0153] In one embodiment, the ability of an active Cas9 molecule to interact with a target nucleic acid is PAM sequence dependent. A PAM sequence is a sequence in a target nucleic acid. Active Cas9 molecules from different bacterial species can recognize different sequence motifs (e.g., PAM sequences). In one embodiment, a Cas9 molecule from S. aureus recognizes the sequence motif NGRR (R=A or G) and directs cleavage of the target nucleic acid sequence 1-10, e.g., 3-5 base pairs upstream from that sequence. See, e.g., Ran F. et al., NATURE 520:186-191 (2015). The ability of a Cas9 molecule to recognize a PAM sequence can be determined, e.g., using a transformation assay as described in Jinek et al., SCIENCE 337:816 (2012). Some Cas9 molecules have the ability to interact with and home to (e.g., target or localize to) a core target domain in cooperation with a gRNA molecule, but are incapable of cleaving a target nucleic acid or are incapable of cleaving at an efficient rate. Cas9 molecules that have no or substantially no cleavage activity may be referred to herein as inactive Cas9 (enzymatically inactive Cas9), dead Cas9, or dCas9 molecules. For example, an inactive Cas9 molecule may lack cleavage activity or have substantially less cleavage activity than a reference Cas9 molecule, e.g., less than 20, 10, 51, or 0.1%, as measured by the assays described herein.
[0154] Exemplary naturally occurring Cas9 molecules that can be used in the methods provided herein are described in Chylinski et al., RNA Biology 10(5):727-737 (2013). Such Cas9 molecules include those from Cluster 1 Bacteria, Cluster 2 Bacteria, Cluster 3 Bacteria, Cluster 4 Bacteria, Cluster 5 Bacteria, Cluster 6 Bacteria, Cluster 7 Bacteria, Cluster 8 Bacteria, Cluster 9 Bacteria, Cluster 10 Bacteria, Cluster 11 Bacteria, Cluster 12 Bacteria, Cluster 13 Bacteria, Cluster 14 Bacteria, Cluster 1 Bacteria, Cluster 16 Bacteria, Cluster 17 Bacteria, Cluster 18 Bacteria, Cluster 19 Bacteria, Cluster 20 Bacteria, Cluster 21 Bacteria, Cluster 22 Bacteria, Cluster 23 Bacteria, Cluster 24 Bacteria, Cluster 25 Bacteria, Cluster 26 Bacteria, Cluster 27 Bacteria, Cluster 28 Bacteria, Cluster 29 Bacteria, Cluster 30 Bacteria, Cluster 31 Bacteria, Cluster 32 Bacteria, Cluster 33 Bacteria, Cluster 34 Bacteria, Cluster 35 Bacteria, Cluster 36 Bacteria, Cluster 37 Bacteria, Cluster 38 Bacteria, Cluster 39 Bacteria, Cluster 40 Bacteria, Cluster 41 Bacteria, Cluster 42 Bacteria, Cluster 43 Bacteria, Cluster 44 Bacteria, Cluster 45 Bacteria, Cluster 46 Bacteria, Cluster 47 Bacteria, Cluster 48 Bacteria, Cluster 49 Bacteria, Cluster 50 Bacteria, Cluster 51 Bacteria, Cluster 52 Bacteria, Cluster 53 Bacteria, Cluster 54 Bacteria, Cluster 55 Bacteria, Cluster 56 Bacteria, Cluster 57 Bacteria, Cluster 58 Bacteria, Cluster 59 Bacteria, Cluster 60 Bacteria, Cluster 61 Bacteria, Cluster 62 Bacteria, Cluster 63 Bacteria, Cluster 64 Bacteria, Cluster Cas9 molecules of the family Cluster 40, Cluster 41, Cluster 42, Cluster 43, Cluster 44, Cluster 45, Cluster 46, Cluster 47, Cluster 48, Cluster 49, Cluster 50, Cluster 51, Cluster 52, Cluster 53, Cluster 54, Cluster 55, Cluster 56, Cluster 57, Cluster 58, Cluster 59, Cluster 60, Cluster 61, Cluster 62, Cluster 63, Cluster 64, Cluster 65, Cluster 66, Cluster 67, Cluster 68, Cluster 69, Cluster 70, Cluster 71, Cluster 72, Cluster 73, Cluster 74, Cluster 75, Cluster 76, Cluster 77, or Cluster 78.
[0155] Exemplary naturally occurring Cas9 molecules include those from the cluster 1 bacterial family, such as S. pyogenes (e.g., strains SF370, MGAS10270, MGAS10750, MGAS2096, MGAS315, MGAS5005, MGAS6180, MGAS9429, NZ131, and SSI-1), S. thermophilus (e.g., strain LMD-9), S. pseudoporcinus (e.g., strain SPIN20026), S. mutans (e.g., strains UA159, NN2025), S. macacae (e.g., strain NCTC11558), S. gallolylicus (e.g., strain UCN34, ATCC 61261, and SCI-1), S. pyogenes (e.g., strain SF370, MGAS10270, MGAS10750, MGAS2096, MGAS315, MGAS5005, MGAS6180, MGAS9429, NZ131, and SSI-1), S. thermophilus (e.g., strain LMD-9), S. pseudoporcinus (e.g., strain SPIN20026), S. mutans (e.g., strains UA159, NN2025), S. macacae (e.g., strain NCTC11558), S. gallolylicus (e.g., strain UCN34, ATCC 61261, and SCI-1), S. pyogenes (e.g., strain SF370, MGAS10270, MGAS10750, MGAS2096, MGAS315, MGAS5005, MGAS6180, BAA-2069), S. equines (e.g., strain ATCC9812, MGCS124), S. dysdalactiae (e.g., strain GGS124), S. bovis (e.g., strain ATCC700338), S. anginosus (e.g., strain F0211), S. agalactia (e.g., strain NEM316, A909), Listeria monocytogenes (e.g., strain F6854), Listeria innocua (e.g., strain Clip11262), Enterococcus italicus (e.g., strain DSM15952), Enterococcus faecium (e.g., strain L15952), L15952 (e.g., strain L25952), L25952 (e.g., strain L25952), L15952 (e.g., strain L1 ... faecium (e.g., strains 1,231, 408), C. jejune, or Deltaproteobacteria (Dbp). Additional exemplary Cas9 molecules are the Neisseria meningitidis Cas9 molecule (Hou et al. PNAS Early Edition 1-6 (2013)) and the S. aureus Cas9 molecule.
[0156] In one embodiment, the Cas9 molecule, e.g., a non-active Cas9 molecule, is any Cas9 molecule sequence described herein or a naturally occurring Cas9 molecule sequence, e.g., any of the Cas9 molecule sequences listed herein or described in Chylinski et al., RNA Biology 10:5 (2013), or Hou et al. PNAS Early Edition. 1-6 (2013); differs in no more than 1%, 2%, 5%, 10%, 15%, 20%, 30% or 40% of the amino acid residues when compared to said Cas9 sequence; differs in at least 1, 2, 5, 10 or 20 amino acids but not more than 100, 80, 70, 60, 50, 40 or 30 amino acids from said Cas9 sequence; or is identical to said Cas9 sequence.
[0157] In one embodiment, the Cas9 molecule comprises an amino acid sequence that has 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% homology to S. aureus Cas9; differs in no more than 1%, 2%, 5%, 10%, 15%, 20%, 30% or 40% of the amino acid residues when compared to the S. aureus Cas9 sequence; differs by at least 1, 2, 5, 10 or 20 amino acids from the S. aureus Cas9 sequence but not more than 100, 80, 70, 60, 50, 40 or 30 amino acids; or is identical to S. aureus Cas9.
[0158] Various types of Cas molecules can be used herein. In some embodiments, Cas molecules of type II Cas systems are used. In other embodiments, Cas molecules of other Cas systems are used. For example, type I or type III Cas molecules can be used. Exemplary Cas molecules (and Cas systems) are described, for example, in Haft et al., PLoS COMPUTATIONAL BIOLOGY 1(6):e60 (2005) and Makarova et al., NATURE REVIEW MICROBIOLOGY 9:467-477 (2011) (the contents of both references are incorporated herein by reference in their entirety).
[0159] Modified Cas9 molecule Naturally occurring Cas9 molecules can have several properties, including nickase activity, nuclease activity (e.g., endonuclease and / or exonuclease activity); helicase activity; the ability to functionally associate with a gRNA molecule; and the ability to target (or localize to) a site on a nucleic acid (e.g., PAM recognition and specificity). In one embodiment, the Cas9 molecule used in the methods disclosed herein can include all or a subset of these properties. In an exemplary embodiment, the Cas9 molecule has the ability to interact with a gRNA molecule and localize to a site in a nucleic acid in concert with the gRNA molecule. Other activities, such as PAM specificity, cleavage activity, or helicase activity, can be more widely variable by the Cas9 molecule.
[0160] Cas9 molecules with desired properties can be generated in several ways, for example, by modification of a parent Cas9, such as a naturally occurring Cas9 molecule, to provide a modified Cas9 molecule with desired properties. For example, one or more mutations or differences can be introduced into the parent Cas9 molecule. Such mutations and differences can include substitutions (e.g., conservative substitutions or substitutions of non-essential amino acids); insertions; or deletions. In one embodiment, the Cas9 molecule can include one or more mutations or differences, e.g., at least 1, 2, 3, 4, 5, 10, 15, 20, 30, 300, 40, or 50 mutations, but not more than 200, 100, or 80 mutations, compared to a reference Cas9 molecule, while retaining or enhancing one or more activities of the reference Cas9 molecule. For example, in one embodiment, the Cas molecule includes one or more amino acid deletions compared to a wild-type Cas molecule sequence.
[0161] In one embodiment, the one or more mutations do not substantially affect Cas9 activity (e.g., PAM recognition and specificity) targeted to (or localized to) one side of a nucleic acid, hi one embodiment, the one or more mutations substantially affect Cas9 activity, such as nickase activity, nuclease activity, and helicase activity.
[0162] For example, substitutions, whether or not specific to a sequence, can affect one or more activities, such as targeting activity, cleavage activity, etc., and can be assessed or predicted, for example, by assessing whether the mutation is conservative or by the methods described above. In one embodiment, a "non-essential" amino acid residue used in connection with a Cas9 molecule is a residue that can be altered from the wild-type sequence of a Cas9 molecule, e.g., a naturally occurring Cas9 molecule, e.g., an active Cas9 molecule, without eliminating, or more preferably without substantially altering, a Cas9 activity (e.g., targeting ability), whereas altering an "essential" amino acid residue results in substantial loss of activity (e.g., targeting ability).
[0163] Cas9 molecules with altered PAM recognition Naturally occurring Cas9 molecules can recognize specific PAM sequences, such as the PAM recognition sequences described above for S. pyogenes, S. thermophilus, S. mutans, S. aureus, and N. meningitidis.
[0164] In one embodiment, the Cas9 molecule has the same PAM specificity as a naturally occurring Cas9 molecule. In other embodiments, the Cas9 molecule has a PAM specificity that is not associated with a naturally occurring Cas9 molecule or with a naturally occurring Cas9 molecule with the closest sequence homology. For example, a naturally occurring Cas9 molecule can be modified, for example, to alter PAM recognition, for example, to reduce off-target sites and / or improve specificity by changing the PAM sequence that the Cas9 molecule recognizes; or to eliminate the PAM recognition requirement. In one embodiment, the Cas9 molecule can be modified to reduce off-target sites and increase specificity, for example, by increasing the length of the PAM recognition sequence and / or improving Cas9 specificity to a high identity level. In one embodiment, the length of the PAM recognition sequence is at least 4, 5, 6, 7, 8, 9, 10, or 15 amino acids in length. Cas9 molecules that recognize different PAM sequences and / or have reduced off-target activity can be generated using directed evolution. Exemplary methods and systems that can be used for the directed evolution of Cas9 molecules are described, for example, in Esvelt et al, Nature 472(7344):499-503 (2011). Candidate Cas9 molecules can be evaluated, for example, by the methods described herein.
[0165] Uncleaved Cas9 molecule In one embodiment, the Cas9 molecule comprises a cleavage that is different from a naturally occurring Cas9 molecule, e.g., different from the closest homologous naturally occurring Cas9 molecule. For example, the Cas9 molecule can differ from a naturally occurring Cas9 molecule, such as a Cas9 molecule from S. aureus, in that it has, for example, a reduced ability to modulate double-strand break cleavage (endonuclease and / or exonuclease activity) compared to a naturally occurring Cas9 molecule (e.g., a Cas9 molecule from S. aureus); a reduced ability to modulate cleavage of a single strand of a nucleic acid, e.g., a non-complementary strand of a nucleic acid molecule or a complementary strand of a nucleic acid molecule (nickase activity) compared to a naturally occurring Cas9 molecule (e.g., a Cas9 molecule from S. aureus); or can eliminate the ability to cleave a nucleic acid molecule, e.g., a double-stranded or single-stranded nucleic acid molecule.
[0166] Uncleaved inactive Cas9 molecule In one embodiment, the modified Cas9 molecule is an inactive Cas9 molecule that does not cleave a nucleic acid molecule (either a double-stranded or single-stranded nucleic acid molecule) or cleaves a nucleic acid molecule with a significantly lower efficiency, e.g., less than 20, 10, 5, 1, or 0.1% of the cleavage activity of a reference Cas9 molecule, as measured by the assays described herein. The reference Cas9 molecule can be a naturally occurring unmodified Cas9 molecule, e.g., a naturally occurring Cas9 molecule, such as a Cas9 molecule of S. pyogenes, S. thermophilus, S. aureus, or N. meningitidis. In one embodiment, the reference Cas9 molecule is the naturally occurring Cas9 molecule with the closest sequence identity or homology. In one embodiment, the inactive Cas9 molecule lacks substantial cleavage activity associated with the N-terminal RuvC-like domain and cleavage activity associated with the HNH-like domain. In one embodiment, the Cas9 molecule is dCas9. See, e.g., Tsai et al. Nat. Biotech. 32:569-577 (2014).
[0167] A catalytically inactive Cas9 molecule can be fused to a transcriptional activator. An inactive Cas9 fusion protein complexes with a gRNA and localizes to the DNA sequence specifically indicated by the targeting domain of the gRNA, but unlike active Cas9, it does not attempt to cleave the target DNA. Fusion of an effector domain, such as a transcriptional activation domain, to an inactive Cas9 allows for the recruitment of the effector to any DNA site specifically indicated by the gRNA. Site-specific targeting of a Cas9 fusion protein to the promoter region of a gene can induce or affect polymerase binding to the promoter region. For example, Cas9 fusion with a transcription factor (e.g., a transcriptional activator) and / or a transcriptional enhancer increases transcriptional activation upon binding to the nucleic acid. In one embodiment of the present invention, the transcriptional activator or a domain thereof is encoded by a nucleic acid molecule comprising less than 1,650 nucleotides. In another embodiment of the present invention, the transcriptional repressor or a domain thereof is encoded by a nucleic acid molecule comprising less than 1,650 nucleotides.
[0168] Table 2 below provides the sequences of gRNAs for SCN1A.
[0169] [Table 20]
[0170] Zinc Finger Proteins In some embodiments, the DBD is a zinc finger protein. Zinc finger proteins can be used interchangeably as "zinc finger domains" and "zinc finger motifs." Zinc fingers contain one or more zinc ions (Zn 2+) coordination. Zinc finger (Znf) domains are relatively small protein motifs that contain multiple finger-like protrusions that contact DNA target sites in tandem. The modular nature of zinc finger motifs allows for numerous combinations of DNA sequences to be bound with high affinity and specificity, making them ideally suited to engineer proteins capable of targeting and binding to specific DNA sequences. Many engineered zinc finger arrays are based on the zinc finger domain of the mouse transcription factor Zif268. Zif268 has three individual zinc finger motifs that collectively bind to a 9 bp sequence with high affinity. A variety of zinc finger proteins have been identified and characterized into various types based on their structures as further described herein. Any such zinc finger proteins are useful in conjunction with the DBDs described herein.
[0171] Various methods are available for designing zinc finger proteins. For example, methods for designing zinc finger proteins to bind to target DNA sequences of interest have been described, see, e.g., Liu Q, et al., Design of polydactyl zinc-finger proteins for unique addressing within complex genomes, Proc Natl Acad Sci USA. 94(11):5525-30(1997); Wright DA et al., Standardized reagents and protocols for engineering zinc finger nucleases by modular assembly, Nat Protoc. Nat Protoc. 2006; 1(3):1637-52; and CA Gersbach and T Gaj, Synthetic Zinc Finger Proteins: The Advent of Targeted Gene Regulation and Genome Modification Technologies, Am Chem Soc 47:2309-2318(2014). Additionally, various web-based tools are publicly available for designing zinc finger proteins to bind to a DNA target sequence of interest, see, e.g., the Zinc Finger Nuclease Design Software Tools and Genome Engineering Data Analysis website from OmicX, available on the World Wide Web at omictools.com / zfns-category; and the Zinc Finger Tools design website from Scripps, available on the World Wide Web at scripps.edu / barbas / zfdesign / zfdesignhome.php.Additionally, various commercial services are available for designing zinc finger proteins to bind to a DNA target sequence of interest, see for example the commercial services or kits provided by Creative Biolabs (world wide web at creative-biolabs.com / Design-and-Synthesis-of-Artificial-Zinc-Finger-Proteins.html), the Zinc Finger Consortium Modular Assembly Kit available from Addgene (world wide web at addgene.org / kits / zfc-modular-assembly / ) or the CompoZr Custom ZFN Service from Sigma Aldrich (world wide web at sigmaaldrich.com / life-science / zinc-fmger-nuclease-technology / custom-zfn.html).
[0172] In certain embodiments, the transcription factors provided herein comprise a DBD that includes one or more zinc fingers or are derived from a DBD of a zinc finger protein. In some cases, the DBD includes multiple zinc fingers, with each zinc finger linked at either its N-terminus or C-terminus or both to another zinc finger or another domain via an amino acid linker. In some cases, the DBDs provided herein include multiple zinc finger structures or motifs or multiple zinc fingers having one or more of SEQ ID NOs: 70-73 and 61 set forth in Table 3, or any combination thereof.
[0173] In certain embodiments, the DBD comprises X-[ZF-X]n and / or [X-ZF]nX, where ZF is a zinc finger domain, X is an amino acid linker comprising 1-50 amino acids, and n is an integer from 1-15, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15, each ZF can independently have the same or a different sequence from other ZF sequences in the DBD, and each linker X can independently have the same or a different sequence from other X sequences in the DBD. Each zinc finger can be linked to another sequence, zinc finger or domain at its C-terminus, N-terminus, or both. In a DBD, each linker X can be identical in sequence, length, and / or properties (e.g., flexibility or charge) or can differ in sequence, length, and / or properties. In some cases, the two or more linkers may be identical, while other linkers are different. In exemplary embodiments, the linker may be derived or obtained from a sequence connecting zinc fingers found in one or more naturally occurring zinc finger proteins. In other embodiments, suitable linker sequences include, for example, linkers of 5 amino acids or more in length. For exemplary linker sequences of 6 amino acids or more in length, see also U.S. Pat. Nos. 6,479,626; 6,903,185; and 7,153,949, each of which is incorporated herein in its entirety. The DBD proteins provided herein may include any combination of suitable linkers between individual zinc fingers of the protein. The DBD proteins described herein may include any combination of suitable linkers between individual zinc fingers of the protein.
[0174] In certain embodiments, the transcription factors provided herein comprise a DBD that includes one or more classical zinc fingers. A classical C2H2 zinc finger has two cysteine residues on one strand and two histidine residues on the other strand coordinated by a zinc ion. A classical zinc finger domain has two beta sheets and an alpha helix, which interacts with a DNA molecule and forms the basis of the DBD that binds to a target site, and may be referred to as a "recognition helix." In an exemplary embodiment, the recognition helix of the zinc finger contains at least one amino acid substitution at positions 1, 2, 3, or 6 to alter the binding specificity of the zinc finger domain. In other embodiments, the DBDs provided herein comprise one or more non-classical zinc fingers, such as C2-H2, C2-CH, and C2-C2.
[0175] In another embodiment, a transcription factor provided herein comprises a DBD that comprises a zinc finger motif having the following structure: LEPGEKP-[YKCPECGKSFSXHQRTHTGEKP]n-YKCPECGKSFSXHQRTH-TGKKTS, where "LEPGEKPYKCPECGKSFS" is disclosed as SEQ ID NO: 91, "HQRTHTGEKPYKCPECGKSFS" is disclosed as SEQ ID NO: 92, and "HQRTHTGKKTS" is disclosed as SEQ ID NO: 83, and n is an integer from 1 to 15, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15, and each X is independently a recognition sequence (e.g., a recognition helix) capable of binding 3 bp of a target sequence. In an exemplary embodiment, n is 3, 6, or 9. In a particularly preferred embodiment, n is 6. In various embodiments, each X can independently have the same amino acid sequence or a different amino acid sequence compared to the sequences of other X in the DBD. In an exemplary embodiment, each X is a sequence comprising 7 amino acids designed to interact with 3bp of a target binding site of interest using the Zinger Finger Design Tool from Scripps located on the world wide web at scripps.edu / barbas / zfdesign / zfdesignhome.php.
[0176] Since each zinc finger in a DBD recognizes 3 bp, the number of zinc fingers contained in the DBD informs the length of the binding site recognized by the DBD, e.g., a DBD with 1 zinc finger recognizes a target binding site with 3 bp, a DBD with 2 zinc fingers recognizes a target binding site with 6 bp, a DBD with 3 zinc fingers recognizes a target binding site with 9 bp, a DBD with 4 zinc fingers recognizes a target binding site with 12 bp, a DBD with 5 zinc fingers recognizes a target binding site with 15 bp, a DBD with 6 zinc fingers recognizes a target binding site with 18 bp, a DBD with 9 zinc fingers recognizes a target binding site with 27 bp, etc. In general, DBDs that recognize longer target binding sites will exhibit greater binding specificity (e.g., less off-target or non-specific binding).
[0177] In other embodiments, the transcription factors provided herein comprise a DBD derived from a naturally occurring zinc finger protein by making one or more amino acid substitutions in one or more of the recognition helices of the zinc finger domain to alter the binding specificity of the DBD (e.g., to alter the target site recognized by the DBD). The DBDs provided herein can be derived from any naturally occurring zinc finger protein.
[0178] In various embodiments, such DBDs can be derived from zinc finger proteins of any species, e.g., mouse, rat, human, etc. In exemplary embodiments, the DBDs provided herein are derived from human zinc finger proteins. In certain embodiments, the DBDs provided herein are derived from naturally occurring proteins listed in Table 9. In exemplary embodiments, the DBD proteins provided herein are derived from human EGR zinc finger proteins, e.g., EGR1, EGR2, EGR3, or EGR4.
[0179] In certain embodiments, the transcription factors provided herein that upregulate SCN1A comprise a DBD derived from a naturally occurring protein by modifying the DBD to increase the number of zinc finger domains in the DBD protein by repeating one or more zinc fingers in the DBD of the naturally occurring protein. In certain embodiments, such modifications include duplication, triplication, quadruplicate or further multiplication of zinc fingers in the DBD of the naturally occurring protein. In some cases, one zinc finger from the DBD of a human protein is multiplexed, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more copies of the same zinc finger motif are repeated in the DBD of the transcription factor. In some cases, a set of zinc fingers from the DBD of a naturally occurring protein is multiplexed. For example, a set of 3 zinc fingers from the DBD of a naturally occurring protein is duplicated to generate a transcription factor with a 6 zinc finger DBD, triplicated to generate a transcription factor with a 9 zinc finger DBD, or quadruplicated to generate a transcription factor with a 12 zinc finger DBD, etc. In some cases, a set of zinc fingers from the DBD of a naturally occurring protein is partially duplicated to form a transcription factor DBD with a larger number of zinc fingers. For example, if the zinc fingers present in the DBD of a transcription factor one copy of the first zinc finger, one copy of the second zinc finger, and two copies of the third zinc finger from a naturally occurring protein, totaling four zinc fingers, the DBD of the transcription factor contains four zinc fingers. Such a DBD is then further modified by making one or more amino acid substitutions in one or more of the recognition helices of the zinc finger domain to change the binding specificity of the DBD (e.g., to change the target site recognized by the DBD). In exemplary embodiments, the DBD is derived from a naturally occurring human protein, such as a human EGR zinc finger protein, such as EGR1, EGR2, EGR3, or EGR4.
[0180] Human EGR1 and EGR3 are characterized by a three-fingered C2H2 zinc finger DBD. General binding rules for zinc fingers dictate that all three fingers interact with their cognate DNA sequence in a similar geometry, using identical amino acids in the alpha helix of each zinc finger to determine the specificity or recognition of the target binding site sequence. Such binding rules allow the DBD of EGR1 or EGR3 to be modified to engineer a DBD that recognizes a desired target binding site. In some cases, the seven amino acid DNA recognition helix in the zinc finger motif of EGR1 or EGR3 is modified according to published zinc finger design rules. In certain embodiments, each zinc finger of the three-fingered DBD of EGR1 or EGR3 is modified, for example, by changing the sequence of one or more recognition helices and / or by increasing the number of zinc fingers of the DBD. In certain embodiments, EGR1 or EGR3 is reprogrammed to recognize a target binding site that is at least 9, 12, 15, 18, 21, 24, 27, 30, 33, 36 or more base pairs of a desired target site. In certain embodiments, such DBDs derived from ERG1 or EGR3 contain at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more zinc fingers. In exemplary embodiments, one or more of the zinc fingers in the DBD contains at least one amino acid substitution at position 1, 2, 3 or 6 of the recognition helix.
[0181] Table 3 below provides the amino acid sequences of exemplary zinc finger proteins used in the examples.
[0182] [Table 21]
[0183] Table 4 below provides other sequences for use in transcription factors.
[0184] [Table 22]
[0185] [Table 23]
[0186] [Table 24]
[0187] [Table 25]
[0188] [Table 26]
[0189] [Table 27]
[0190] Upregulation of genes of interest The transcription factors provided herein can be designed to recognize any target site (e.g., promoter region of a gene of interest) that results in upregulation of the gene of interest. In an exemplary embodiment, the DBD is designed to recognize the genomic location of the endogenous gene of interest and upregulate its expression when the transcription factor binds. The binding site capable of regulating the expression of the endogenous gene of interest when the transcription factor provided herein binds can be located at any location in the genome that results in modulation of gene expression of the gene of interest. In various embodiments, the binding site can be located on a different chromosome than the gene of interest, on the same chromosome as the gene of interest, upstream of the transcription start site (TSS) of the gene of interest, downstream of the TSS of the gene of interest, proximal to the TSS of the gene of interest, distal to the gene of interest, within the coding region of the gene of interest, within an intron of the gene of interest, downstream of the polyA tail of the gene of interest, within the promoter sequence that regulates the gene of interest, or within an enhancer sequence that regulates the gene of interest.
[0191] The DBD can be designed to bind to a target binding site of any length, for example, so long as it provides specific recognition of the target binding site sequence by the DBD with minimal or no off-target binding. In certain embodiments, the target binding site, when bound by a trans-acting factor, can regulate expression of the gene of interest at least 2-fold, 5-fold, 10-fold, 20-fold, 50-fold, 75-fold, 100-fold, 150-fold, 200-fold, 250-fold, 500-fold, or more, compared to all other genes. In certain embodiments, the target binding site, when bound by a trans-acting factor, can regulate expression of the gene of interest at least 2-fold, 5-fold, 10-fold, 20-fold, 50-fold, 75-fold, 100-fold, 150-fold, 200-fold, 250-fold, 500-fold, or more, compared to the 40 nearest neighbor genes (e.g., the 40 genes located closest to the coding sequence of the gene of interest on the chromosome, either upstream or downstream). In certain embodiments, the target binding site may be at least 5 bp, 10 bp, 15 bp, 20 bp, 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, or 50 bp or more. The specific length of the binding site will be characterized by the type of DBD in the transcription factor. In general, the longer the length of the binding site, the greater the specificity of the binding and the modulation of gene expression (e.g., the longer the length of the binding site, the fewer off-target effects). In certain embodiments, a transcription factor with a DBD that recognizes a longer target binding site will have fewer off-target effects associated with non-specific binding (e.g., modulation of expression of off-target genes or genes other than the gene of interest) compared to the off-target effects observed with a transcription factor with a DBD that binds to a shorter target site. In some cases, the reduction in off-target binding will be at least 1.2, 1.3, 1.4, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, or 10-fold compared to a comparable transcription factor with a DBD that recognizes a shorter target binding site.
[0192] In certain embodiments, the DBDs provided herein can be modified to increase binding affinity and bind to the target binding site for a longer period of time, such that a TAD conjugated to the DBD can recruit more transcription factors and / or can recruit such transcription factors for a longer period of time to exert a greater effect on the expression level of an endogenous gene of interest. In certain embodiments, the DBD can be modified to increase its specific binding (or on-target binding) to a desired target site and / or to decrease its non-specific or off-target binding.
[0193] In various embodiments, binding of the DBD or transcription factor to the target binding site can be determined using various methods. In certain embodiments, specific binding of the DBD or transcription factor to the target binding site can be determined using a mobility shift assay, a DNase protection assay, or any other in vitro method known in the art for assaying protein-DNA binding. In other embodiments, specific binding of the transcription factor to the target binding site can be determined using a functional assay, for example, by measuring expression (RNA or protein) of a gene of interest upon binding of the transcription factor to the target binding site. For example, the target binding site can be located upstream of a reporter gene (e.g., eGFP, etc.) or gene of interest on a vector contained in the cell, or can be integrated into the genome of the cell if the cell expresses the transcription factor. Alternatively, a vector expressing the transcription factor can be introduced into a cell type that naturally contains the gene of interest. A greater level of expression of the reporter gene (or gene of interest) in the presence of the transcription factor compared to a control (e.g., no transcription factor or a transcription factor that recognizes a different target site) indicates that the DBD of the transcription factor binds to the target site. Similarly, suitable in vitro (e.g., non-cell-based) transcription and translation systems may be used. In certain embodiments, a transcription factor that binds to a target site may increase expression of a reporter gene or gene of interest by at least 2-fold, 3-fold, 5-fold, 10-fold, 15-fold, 20-fold, 30-fold, 50-fold, 75-fold, 100-fold, 150-fold or more compared to a control (e.g., no transcription factor or a transcription factor that recognizes a different target site).
[0194] In certain embodiments, the transcription factors disclosed herein that upregulate a gene of interest have a size of at least 9 bp, 12 bp, 15 bp, 18 bp, 21 bp, 24 bp, 27 bp, 30 bp, 33 bp, or 36 bp; more than 9 bp, 12 bp, 15 bp, 18 bp, 21 bp, 24 bp, 27 bp, or 30 bp; or 9-33 bp, 9-30 bp, 9-27 bp, 9-24 bp, 9-21 bp, 9-18 bp, 9-15 bp, 9-12 bp, 12-33 bp, 12-30 bp, The target binding site is 12-27bp, 12-24bp, 12-21bp, 12-18bp, 12-15bp, 15-33bp, 15-30bp, 15-27bp, 15-24bp, 15-21bp, 15-18bp, 18-33bp, 18-30bp, 18-27bp, 18-24bp, 18-21bp, 21-33bp, 21-30bp, 21-27bp, 21-24bp, 24-33bp, 24-30bp, 24-27bp, 27-33bp, 27-30bp, or 30-33bp. In an exemplary embodiment, the transcription factor disclosed herein that upregulates a gene of interest recognizes a target binding site that is 18-27bp, 18bp, or 27bp.
[0195] In certain embodiments, the transcription factors disclosed herein that upregulate a gene of interest recognize a target binding site located on chromosome 2. In certain embodiments, the transcription factors disclosed herein that upregulate a gene of interest recognize a target binding site located within 110 kb, 100 kb, 90 kb, 80 kb, 70 kb, 60 kb, 50 kb, 40 kb, 30 kb, 20 kb, 10 kb, 5 kb, 4 kb, 3 kb, 2 kb, or 1 kb upstream or downstream of the TSS of gene of interest A on the chromosome.
[0196] The gene of interest may be any gene in which genomic alterations result in reduced transcription or activity of one or more genes or whose gene product is a causative agent of a mammalian disease. For example, the genomic alteration may be haploinsufficient, in which only one functional copy of the gene is present and the single copy does not produce sufficient gene product to produce a wild-type phenotype. Other diseases are caused by genomic alterations of one or both copies of the gene that alter the gene product to exhibit reduced activity but not eliminated. In still other diseases, the genomic alteration reduces the transcription or reduces transcript stability of one or both genes such that insufficient gene product is present to produce a wild-type phenotype.
[0197] In some embodiments, the increase in expression of the gene of interest is measured by an increase in the number of RNA transcripts of the transgene sequence. In some embodiments, the increase in expression of the transgene sequence is measured by PCR. In some embodiments, the increase in expression of the transgene sequence is measured by RT-PCR. In some embodiments, the increase in expression of the transgene sequence is measured by qPCR. In some embodiments, the increase in expression of the transgene sequence is measured by qRT-PCR. In some embodiments, the increase in expression of the transgene sequence is measured by sequencing. In some embodiments, the increase in expression of the transgene sequence is measured by Northern blot analysis. In some embodiments, the increase in expression of the transgene sequence is measured by single molecule fluorescent in-situ hybridization (FISH). In some embodiments, the increase in expression of the transgene sequence is measured by an increase in the production of a protein encoded by the transgene. In some embodiments, the increase in expression of the transgene sequence is measured by enzyme-linked immunosorbent assay (ELISA). In some embodiments, the increase in expression of the transgene sequence is measured by Western blot analysis. In some embodiments, the increase in expression of the transgene sequence is measured by immunostaining. In some embodiments, increased expression of the transgene sequence is measured by more than one of the methods listed above.
[0198] promoter All cells in an animal or human body contain the same DNA, but different cells in different tissues express a common set of genes on the one hand and a set of genes that varies depending on the tissue type and stage of development on the other hand. Without being bound by theory, any promoter that does not contain introns can be used in the various aspects and embodiments (e.g., nucleic acid molecules) described herein. Exemplary promoters that can be used in the various aspects and embodiments described herein include, but are not limited to, promoters that can act in GABAergic neurons, promoters that can act in inhibitory neurons, cytomegalovirus (CMV) promoter, CAG promoter, SV40 promoter, JeT promoter, PGK promoter, and chicken beta-actin promoter (CBA) promoter. In embodiments, the promoter is active in multiple cell types. In other embodiments, the promoter is active in one cell type (e.g., cell-specific) or in one tissue, such as, for example, central nervous tissue (e.g., brain tissue), (e.g., tissue-specific). In embodiments, the promoter is neuron-specific. Examples of neuron-specific promoters that can be used in various aspects and embodiments described herein include isolated or synthetic neuron-specific promoters and functional fragments thereof used in vectors and other nucleic acids to drive expression of operably linked minigenes and transgenes, such as the promoter derived from neuron-specific enolase (NSE) (see, e.g., EMBL HSENO2, X51956); the aromatic amino acid decarboxylase (AADC) promoter; the neurofilament promoter (see, e.g., GenBank HUMNFL, L04147); the synapsin promoter (see, e.g., GenBank HUMSYNIB, M55301); the thy-1 promoter (see, e.g., Chen et al., (1987) Cell, 51:7-19; Llewellyn et al. (2010) Nat. Med., 16(10):1161-1166); serotonin receptor promoter (see, e.g., GenBank S62283); tyrosine hydroxylase promoter (TH) (see, e.g., Oh et al., (2009) Gene Ther., 16:437; Sasaoka et al., (1992) Mol. Brain Res., 16:274; Boundy et al., (1998) J. Neurosci., 18:9989; and Kaneda et al., (1991) Neuron, 6:583-594); GnRH promoter (see, e.g., Radovick et al., (1991) Proc. Natl. Acad. Sci. USA, 88:3402-3406); L7 promoter (see, e.g., Oberdick et al., (1991) Proc. Natl. Acad. Sci. USA, 88:3402-3406); al., (1990) Science, 248:223-226); DNMT promoter (see, e.g., Bartge et al., (1988) Proc. Natl. Acad. Sci. USA, 85:3648-3652); enkephalin promoter (see, e.g., Comb et al., (1988) EMBO J., 17:3793-3805); myelin basic protein (MBP) promoter; Ca2+-calmodulin-dependent protein kinase II-alpha (CamKIM) promoter (see, e.g., Mayford et al., (1996) Proc. Natl. Acad. Sci. USA, 93:13250; and Casanova et al., (2001) Genesis, 31:37); CMV enhancer / platelet-derived growth factor-p promoter (see, e.g., Liu et al., (1999) Proc. Natl. Acad. Sci. USA, 85:3648-3652); al., (2004) Gene Ther., 11:52-60); and the like. In some embodiments, part or all of the minimal human synapsin 1 promoter (SYN) is used. Kugler et al., (2003) Gene Ther., 10(4):337-47; Thiel et al, (1991) Proc. Natl. Acad. Sci. USA, 88(8)3431-5; Castle et al., (2016) Methods Mol. Biol.,1382:133-49;McLean et al.,(2014)Neurosci.Lett.,576:73-78;Kugler et al.,(2003)Virology,311(1):89-95。.
[0199] In some embodiments, the tissue or cell specific promoter is configured to provide higher expression of the operably linked minigene and / or transgene in neural cells or tissues compared to non-neuronal cells. In some embodiments, the neural specific promoter is configured to provide higher expression of the operably linked minigene and / or transgene in neural cells compared to non-neuronal cells. Examples of neural cells or tissues include neurons and those including Schwann cells, glial cells, astrocytes, and the like. Examples of non-neuronal cells include, but are not limited to, hepatocytes, cardiomyocytes, erythrocytes, epithelial cells, and the like. Higher levels of expression of the operably linked minigene and / or transgene can include an increase in the number of RNA transcripts produced from transcription of the minigene and / or transgene. In some embodiments, the number of RNA transcripts produced can be measured by PCR. In some other embodiments, the number of RNA transcripts produced can be measured by RT-PCR, e.g., qPCR. In some embodiments, the number of RNA transcripts produced can be measured by sequencing. In some embodiments, the number of RNA transcripts produced may be measured by single molecule fluorescent in situ hybridization (FISH). In some embodiments, the number of RNA transcripts produced may be measured by Northern blot analysis. If the minigene and / or transgene encodes a protein of interest, a higher level of expression of the operably linked minigene and / or transgene may alternatively or additionally include an increased amount of protein produced. In some embodiments, the amount of protein produced may be measured by enzyme-linked immunosorbent assay (ELISA). In some embodiments, the amount of protein produced may be measured by Western blot analysis. In some embodiments, the amount of protein produced may be measured by immunostaining. In some embodiments, the amount of protein produced may be measured by time-resolved Förster resonance energy transfer (TR-FRET). In some embodiments, the amount of protein produced may be measured by immunohistochemistry (IHC).In some embodiments, the level of expression is measured by more than one of these or other methods.
[0200] Poly A signal sequence In various embodiments, the nucleic acids, vectors and other compositions disclosed herein may include one or more polyadenylation (polyA) signal sequences. The polyadenylation signal sequence may include a central sequence (e.g., AAUAAA) flanked by auxiliary sequence elements. Without being bound by theory, the sequence may signal the end of the transcript and serve as the site at which a homopolymeric A sequence is added to the 3' end by polyadenylate polymerase.
[0201] Polyadenylation signal sequences known in the art are contemplated, including, but not limited to, SV40 poly A, human growth hormone (HGH) poly A, bovine growth hormone (BGH) poly A, beta globin poly A, alpha globin poly A, ovalbumin poly A, kappa light chain poly A, and synthetic poly A. Poly A signal sequences can be used in the nucleic acids and other compositions disclosed herein.
[0202] Post-transcriptional regulatory elements In various embodiments, the nucleic acids, transgenes and other compositions disclosed herein may include one or more post-transcriptional regulatory elements (PREs), such as those that can enhance or otherwise improve expression of the transgene. Without being bound by theory, PREs may enhance expression by enabling mRNA stability and 3' end formation, and / or may facilitate transport of unspliced mRNA to the nucleocytoplasm. PREs may also include binding sites for RNA-binding proteins (RBPs) or microRNAs.
[0203] Exemplary PREs include, but are not limited to, PREs from Hepatitis B virus (HPRE), bat virus (BPRE), ground squirrel virus (GSPRE), arctic squirrel virus (ASPRE), duck virus (DPRE), chimpanzee virus (CPRE), woolly monkey virus (WMPRE), or woodchuck virus (WPRE). In some embodiments, the nucleic acid or transgene comprises a PRE. In certain embodiments, the PRE comprises a HPRE. In some embodiments, a synthetic PRE is used.
[0204] Viral Vectors Also disclosed herein are vectors that include the nucleic acids discussed herein (e.g., minigenes, transgenes, promoters, other nucleic acid components such as PRE and polyA, and combinations thereof). In some embodiments, the vectors can be useful for delivering a transgene to a target cell and / or for increasing expression of that transgene in a target cell. In various embodiments, the vectors can be used to regulate expression of proteins, antibodies or functional binding fragments, enzymes, etc., and / or nucleic acids, such as shRNAs, siRNAs, gRNAs, etc., for use in CRISPR, etc., by use in combination with splice regulators.
[0205] As an example, the vector may include a transcription factor described herein that increases expression of a gene of interest. The vector may be useful for transferring genetic information to another cell. The vector may be used for cloning, for example, as a cloning vector or a plasmid. The vector may also be specifically designed to drive expression, for example, therapeutic protein and / or RNA expression, for other purposes, such as cell infection, for example, in human neuronal cells. In some embodiments, vectors comprising the nucleic acids disclosed herein are contemplated. The vector may be a DNA vector, a circular vector, or a plasmid. In some embodiments, the vector is double-stranded. In other embodiments, the vector is single-stranded.
[0206] In some embodiments, the vector is a viral vector. In some embodiments, the vector is a viral vector used to deliver transgene sequences to neural cells or tissues. Examples of viruses used in vectors include, but are not limited to, retroviruses, adenoviruses, lentiviruses, adeno-associated viruses, and other hybrid viruses. In some embodiments, the viral vector is an adeno-associated virus (AAV) vector, a chimeric AAV vector, an adenovirus vector, a retrovirus vector, a lentivirus vector, a DNA virus vector, a herpes simplex virus vector, a baculovirus vector, or any mutant or derivative thereof.
[0207] Without being bound by theory, the viral vectors disclosed herein can insert their genome into the host cells they infect to deliver their nucleic acid sequences to the host. The inserted viral genome can be episomal or integrated into the host cell chromosome at a site that can be random or targeted. In one embodiment, the vector is a viral vector used to deliver transgene sequences to cells. Examples of viruses used in vectors include, but are not limited to, retroviruses, adenoviruses, lentiviruses, adeno-associated viruses, and other hybrid viruses. Warnock et al., (2011) Methods Mol. Biol., 737:1-25. Lentiviruses are a genus of retroviruses that can integrate significant amounts of viral DNA into host cells, making them an efficient method of gene delivery. Adenoviruses, on the other hand, introduce genetic material that is not integrated into the host cell chromosome, thus reducing the risk of destroying the host cell. In some embodiments, the viral vector is an adeno-associated viral (AAV) vector, a chimeric AAV vector, an adenoviral vector, a retroviral vector, a lentiviral vector, a DNA viral vector, a herpes simplex viral vector, a baculoviral vector, or any mutant or derivative thereof.
[0208] In some embodiments, the vector containing the transgene is or is derived from an adeno-associated virus (AAV). In some embodiments, the vector is a recombinant adeno-associated virus vector (rAAV). The rAAV genome may contain one or more AAV ITRs adjacent to the sequence encoding the transcription factor. In embodiments, the vector further contains other transcription control elements such as those disclosed herein, such as promoters, enhancers, PREs and / or polyA sequences that are functional in target cells to drive the expression of the transgene sequence. The transgene sequence may also contain intron sequences that facilitate the processing of RNA transcripts when expressed in mammalian cells.
[0209] In various embodiments, the AAV vector, e.g., rAAV vector, is a self-complementary AAV vector (scAAV). As used herein, "self-complementary" means that the coding region is designed to form an intramolecular double-stranded template, e.g., at one or more inverted terminal repeats (ITRs). Without being bound by theory, the rate-limiting step of the AAV genome often involves second-strand synthesis, since a typical AAV genome is a single-stranded DNA template. Ferrari et al, (1996) J. Virology, 70(5): 3227-34; Fisher et al, (1996) J. Virology, 70(1): 520-32. However, in the case of the scAAV genome, upon infection, the two complementary halves of the scAAV can assemble to form one double-stranded DNA (dsDNA) unit that is ready for replication and transcription, rather than waiting for cell-mediated synthesis of the second strand. In some embodiments, the rAAV vectors disclosed herein are scAAV vectors, providing faster and / or increased expression.
[0210] In some embodiments, the rAAV vectors disclosed herein lack one or more (e.g., all) AAV rep and / or cap genes. The AAV vector can include nucleic acid sequences (e.g., DNA) from any suitable AAV serotype (e.g., in its ITRs). Suitable AAV serotypes include, but are not limited to, AAV serotypes AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, AAV-9, AAV-10, AAV-11, AAV-12, AAVrh8, AAVrh10, AAV.Anc80, AAV.Anc80L65, AAV-DJ and AAV-DJ / 8, AAVrh37, AAV-DJ, AAV-DJ / 8, AAV-PHP.B, AAV-PHP.B2, AAV-PHP.B3, AAV-PHP.A, AAV-PHP.eB and AAV-PHP.S. For example, an AAV vector, such as an scAAV vector, can include a nucleic acid sequence from AAV2, such as an ITR sequence from AAV2. An AAV vector, such as an scAAV vector, can also include nucleic acids from multiple serotypes. The nucleotide sequences of the genomes of AAV serotypes are known in the art.For example, the entire genome of AAV1 is provided under GenBank Accession No. NC_002077; the entire genome of AAV2 is provided under GenBank Accession No. NC 001401 and Srivastava et al., Virol., 45:555-564 {1983); the entire genome of AAV3 is provided under GenBank Accession No. NC_1829; the entire genome of AAV4 is provided under GenBank Accession No. NC_001829; the AAV5 genome is provided under GenBank Accession No. AF085716; the entire genome of AAV-6 is provided under GenBank Accession No. NC_001862; at least portions of the AAV7 and AAV8 genomes are provided under GenBank Accession Nos. AX753246 and AX753249, respectively; and the AAV9 genome is provided by Gao et al. al., J. Virol., 78:6381-6388 (2004); the AAV10 genome is provided in Williams, (2006) Mol. Ther., 13(1):67-76; the AAV11 genome is provided in Mori et al., (2004) Virology, 330(2):375-383.
[0211] In some embodiments, functional inverted terminal repeat (ITR) sequences may be used, for example, to support rescue, replication, and packaging of AAV virions. Thus, the AAV vectors disclosed herein may include sequences that provide viral replication and packaging in cis (e.g., functional ITRs). The ITRs may be, but need not be, wild-type nucleotide sequences and may be altered, for example, by insertion, deletion, or substitution of nucleotides, so long as the sequences provide functional rescue, replication, and packaging. The ITRs may be from any AAV serotype from which a recombinant virus can be derived, including, but not limited to, AAV serotypes AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, AAV-9, AAV-10, and AAV-11. The nucleotide sequences of the genomes of AAV serotypes are known in the art. For example, the complete genome of AAV-1 is provided under GenBank Accession No. NC_002077; the complete genome of AAV-2 is provided under GenBank Accession No. NC 001401 and Srivastava et al., Virol., 45:555-564 {1983); the complete genome of AAV-3 is provided under GenBank Accession No. NC_1829; the complete genome of AAV-4 is provided under GenBank Accession No. NC_001829; the AAV-5 genome is provided under GenBank Accession No. AF085716; the complete genome of AAV-6 is provided under GenBank Accession No. NC_001862; at least portions of the AAV-7 and AAV-8 genomes are provided under GenBank Accession Nos. AX753246 and AX753249, respectively; and the AAV-9 genome is provided by Gao et al. al., (2004) J. Virol., 78:6381-6388; the AAV-10 genome is provided in Williams, (2006) Mol. Ther., 13(1):67-76; the AAV-11 genome is provided in Mori et al., (2004) Virology, 330(2):375-383. In one embodiment, the vector is an AAV-9 vector having ITRs from AAV-2.
[0212] In some embodiments, the rAAV vectors disclosed herein include one or more ITRs, e.g., two ITRs, one upstream and one downstream of the transgene (e.g., encoding hPGRN) and / or other nucleic acid elements discussed above. In some embodiments, e.g., in scAAV vectors, the nucleic acids disclosed herein include a first ITR located 5' and a second ITR located 3' to the promoter, minigene, transgene, post-transcriptional regulatory element, and / or polyA, e.g., the ITRs are independently 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, 150, 200, 250 nucleotides 5' and / or 3' of the other element. The ITR sequence can be wild-type or can include one or more mutations, e.g., so long as it retains one or more functions of the wild-type ITR. In some embodiments, the wild-type ITR may be modified to include a deletion of a terminal resolution site. In some embodiments, the scAAV disclosed herein may include two ITR sequences, where both are wild-type, mutant, or modified AAV ITR sequences. In some embodiments, at least one ITR sequence is a wild-type, mutant, or modified AAV ITR sequence. In some embodiments, both of the two ITR sequences are wild-type, mutant, or modified AAV ITR sequences. In some embodiments, the "left" or 5'-ITR is a modified AAV ITR sequence that allows for the production of a self-complementary genome, and the "right" or 3'-ITR is a wild-type AAV ITR sequence. In some embodiments, the "right" or 3'-ITR is a modified AAV ITR sequence that allows for the production of a self-complementary genome, and the "left" or 5'-ITR is a wild-type AAV ITR sequence. In some embodiments, the ITR sequences are wild-type, mutant, or modified AAV2 ITR sequences. In some embodiments, at least one ITR sequence is a wild-type, mutant or modified AAV2 ITR sequence, hi some embodiments, both of the two ITR sequences are wild-type, mutant or modified AAV2 ITR sequences.In some embodiments, the "left" or 5'-ITR is a modified AAV2 ITR sequence that allows for the production of a self-complementary genome, and the "right" or 3'-ITR is a wild-type AAV2 ITR sequence. In some embodiments, the "right" or 3'-ITR is a modified AAV2 ITR sequence that allows for the production of a self-complementary genome, and the "left" or 5'-ITR is a wild-type AAV2 ITR sequence. Exemplary sequences that may be used for one or more ITRs are described herein. The embodiments of AAV ITRs provided in WO 2019 / 094253 (PCT / US2018 / 058744), which is incorporated herein by reference in its entirety, may also be used for any of the AAV ITRs disclosed herein.
[0213] In some embodiments, the vector is a rAAV. In some embodiments, the rAAV vector lacks one or more (e.g., all) AAV rep and / or cap genes. The AAV vector can include (e.g., in its ITRs) a nucleic acid sequence (e.g., DNA) from any suitable AAV serotype. Suitable AAV serotypes include, but are not limited to, AAV serotypes AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, AAV-9, AAV-10, and AAV-11. For example, an AAV vector, such as an scAAV vector, can include a nucleic acid sequence from AAV-2, such as an ITR sequence from AAV-2. An AAV vector, such as an scAAV vector, can also include nucleic acids from multiple serotypes. GenBank Accession No. NC 001401 and Srivastava et al., Virol., 45:555-564 {1983); GenBank Accession No. NC_1829; GenBank Accession No. NC_001829; GenBank Accession No. AF085716; GenBank Accession No. NC_001862; GenBank Accession Nos. AX753246 and AX753249; Gao et al., J. Virol., 78:6381-6388 (2004); Williams, (2006) Mol. Ther., 13(1):67-76; and Mori et al., (2004) Virology, 330(2):375-383.
[0214] In some embodiments, functional inverted terminal repeat (ITR) sequences in viral vectors can be used to support, for example, rescue, replication and packaging of AAV virions. Thus, the AAV vectors disclosed herein can include sequences that provide viral replication and packaging in cis (e.g., functional ITRs). The ITRs do not have to be wild-type nucleotide sequences and can be altered, for example, by nucleotide insertion, deletion or substitution, so long as the sequences provide functional rescue, replication and packaging. The ITRs can be from any AAV serotype from which a recombinant virus can be derived, including, but not limited to, AAV serotypes AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, AAV-9, AAV-10 and AAV-11. GenBank Accession No. NC_002077; GenBank Accession No. NC 001401 and Srivastava et al., Virol., 45:555-564 {1983); GenBank Accession No. NC_1829; GenBank Accession No. NC_001829; GenBank Accession No. AF085716; GenBank Accession No. NC_001862; GenBank Accession Nos. AX753246 and AX753249, respectively; Gao et al., (2004) J. Virol., 78:6381-6388; Williams, (2006) Mol. Ther., 13(1):67-76; and Mori et al., (2004) Virology, 330(2):375-383. In one embodiment, the vector is an AAV-9 vector with ITRs from AAV-2.
[0215] In some embodiments, the vectors or nucleic acid sequences disclosed herein form cloning or expression vectors. In such embodiments, the vectors may contain other components that facilitate replication or maintenance of the vector. In some embodiments, the vectors further comprise a selection marker for clone selection. In some embodiments, the selection marker in the vector comprises a prokaryotic or eukaryotic antibiotic resistance gene. In some embodiments, the selection marker in the vector comprises a kanamycin resistance gene. In some embodiments, the selection marker in the vector comprises an ampicillin resistance gene. In some embodiments, the vector further comprises a puromycin resistance gene. In some embodiments, the selection marker in the vector comprises a hygromycin resistance gene.
[0216] Recombinant viruses In various embodiments, the nucleic acids and vectors discussed herein may be present in one or more viral particles, such as recombinant viral particles. Recombinant viruses are viruses that have been produced by recombinant means. A variety of different virus types may be used, such as retroviruses, adenoviruses, lentiviruses, AAV, murine leukemia viruses, and the like. Without being bound by theory, vectors delivered from retroviruses, such as lentiviruses, may provide long-term gene transfer and low immunogenicity, as they allow long-term stable integration and transmission of the transgene to daughter cells. Other suitable retroviruses include gamma retroviruses. Exemplary gamma retroviral vectors include murine leukemia virus (MLV), spleen-limited focus-forming virus (SFFV) and myeloproliferative sarcoma virus (MPSV) and vectors derived therefrom. Other gamma retroviral vectors are described, for example, in Tobias Maetzig et al., "Gamma retroviral Vectors: Biology, Technology and Application" Viruses. 2011 Jun;3(6):677-713. In some embodiments, the virus is a recombinant adenovirus comprising a nucleic acid or vector disclosed herein. In some embodiments, the virus is a recombinant AAV comprising a nucleic acid or vector disclosed herein.
[0217] In some embodiments, the nucleic acid or vector disclosed herein is for use in the manufacture of a recombinant virus. In some embodiments, the nucleic acid or vector disclosed herein is for use in the manufacture of a rAAV. Thus, in various embodiments, a viral composition (also called a virion), such as a rAAV viral composition comprising a viral vector or a nucleic acid disclosed above, is also disclosed herein. In some embodiments, the recombinant virus is an adeno-associated virus (AAV) or any mutant or derivative thereof. In some embodiments, the recombinant virus is a chimeric AAV or any mutant or derivative thereof. In some embodiments, the recombinant virus is an adenovirus or any mutant or derivative thereof. In some embodiments, the recombinant virus is a retrovirus or any mutant or derivative thereof. In some embodiments, the recombinant virus is a lentivirus or any mutant or derivative thereof. In some embodiments, the recombinant virus is a DNA virus or any mutant or derivative thereof. In some embodiments, the recombinant virus is a herpes simplex virus or any mutant or derivative thereof. In some embodiments, the recombinant virus is a baculovirus or any mutant or derivative thereof.
[0218] In some embodiments, the AAV disclosed herein may comprise one or more AAV capsid proteins. The AAV capsid proteins may be from any AAV serotype from which a recombinant virus can be derived, including, but not limited to, AAV serotypes AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, AAV-9, AAV-10, AAV-11, AAV-12, AAVrh8, AAVfh10, AAV-DJ, AAV-DJ / 8, AAV-PHP.B, AAV-PHP.B2, AAV-PHP.B3, AAV-PHP.A, AAV-PHP.eB, and AAV-PHP.S. In some embodiments, the one or more capsid proteins in the AAV are from AAV-9. Without being bound by theory, typically in AAV, three capsid proteins, VP1, VP2 and VP3, multimerize to form a capsid. The polypeptide sequences of capsid proteins are known in the art and can be derived from the genome of AAV. These can be used as exemplary capsids in the AAV virus compositions disclosed herein.For example, the entire genome of AAV-1 is provided under GenBank Accession No. NC_002077; the entire genome of AAV-2 is provided under GenBank Accession No. NC 001401 and Srivastava et al., Virol., 45:555-564 {1983); the entire genome of AAV-3 is provided under GenBank Accession No. NC_1829; the entire genome of AAV-4 is provided under GenBank Accession No. NC_001829; the AAV-5 genome is provided under GenBank Accession No. AF085716; the entire genome of AAV-6 is provided under GenBank Accession No. NC_001862; at least portions of the AAV-7 and AAV-8 genomes are provided under GenBank Accession Nos. AX753246 and AX753249, respectively; and the AAV-9 genome is provided by Gao et al. al., J. Virol., 78:6381-6388 (2004); the AAV-10 genome is provided in Williams, (2006) Mol. Ther., 13(1):67-76; the AAV-11 genome is provided in Mori et al., (2004) Virology, 330(2):375-383. The capsid proteins AAV-PHP.B, AAV-PHP.B2, AAV-PHP.B3, AAV-PHP.A, AAV-PHP.eB or AAV-PHP.S are provided in Deverman et al., (2016) Nat. Biotech., 34:204-209 and Chan et al., (2017) Nat. Neurosci., 20:1172-1179. In some embodiments, the recombinant virus is an AAV that comprises one or more of the AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, and AAV11, AAV12, AAVrh8, AAVrh10, AAV-DJ, AAV-DJ / 8, AAV-PHP.B, AAV-PHP.B2, AAV-PHP.B3, AAV-PHP.A, AAV-PHP.eB, or AAV-PHP.S capsid serotypes or functional variants thereof. In some embodiments, the recombinant virus is an AAV that comprises a combination of capsids from multiple AAV serotypes.
[0219] In some embodiments, the AAV compositions disclosed herein contain one or more cis-acting sequences that direct viral DNA replication (rep), encapsidation / packaging, and host cell chromosomal integration, and are contained within the ITRs. In some embodiments, one or more of these sequences may also be present in trans, rather than in cis, e.g., on a separate plasmid during the virus production process in the host cell. Typically, three AAV promoters (named p5, p19, and p40 by their relative map positions) drive the expression of two AAV internal open reading frames that encode the rep and cap genes of wild-type virus. In some embodiments, one or more of these promoters and / or open reading frames are present in cis in the AAV vectors and / or AAV virions disclosed herein, or are present on a separate plasmid during the AAV virus production process, e.g., in the host cell that produces the virus. Two rep promoters (p5 and p19), coupled with specific splicing of a single AAV intron (nucleotides 2107 and 2227), can result in the production of four rep proteins (rep78, rep68, rep52, and rep40) from the rep gene. The rep proteins have multiple enzymatic properties that are ultimately responsible for the replication of the viral genome. The cap gene is typically expressed from the p40 promoter and encodes the three capsid proteins VP1, VP2, and VP3. Alternative splicing and non-consensus translation initiation sites are responsible for the production of the three related capsid proteins. A single consensus polyadenylation site is present at map position 95 of the AAV genome. The life cycle and genetics of AAV are reviewed in Muzyczka, (1992) Curr. Topics Microbiol. Imm., 158:97-129.
[0220] In some embodiments, the AAV capsid proteins VP1, VP2, VP3 used in the AAV disclosed herein are encoded by or comprise the following sequences: VP1 nucleic acid (SEQ ID NO:84): [ka] VP2 nucleic acid (SEQ ID NO:85): [ka] VP3 nucleic acid (SEQ ID NO:86): [ka] VP1 protein (SEQ ID NO:87): [ka] VP2 protein (SEQ ID NO:88): [ka] VP3 protein (SEQ ID NO:89): [ka]
[0221] In one embodiment, the recombinant virus is an AAV comprising the AAV9 capsid serotype or any mutant or derivative thereof. In some embodiments, the recombinant virus comprises the AAV9 capsid proteins VP1, VP2 and VP3. In some embodiments, the recombinant virus is a scAAV.
[0222] In various embodiments, the target cells of the present disclosure can be any mammalian cell type. In some aspects of the present disclosure, the nucleic acids and vectors modulate expression in neural tissue or body fluids or cells. In some embodiments, the neural tissue is the brain. In some embodiments, the neural tissue is the frontal lobe of the brain. In some embodiments, the neural tissue is the temporal lobe of the brain. In some embodiments, the neural tissue is the central nervous system. In some embodiments, the neural tissue is the spinal cord. In some embodiments, the neural cells are human neural cells. In some embodiments, the neural cells are neurons. In some embodiments, the neural cells are astrocytes. In some embodiments, the neural fluid is cerebrospinal fluid. In some embodiments, the non-neural tissue is liver. In some embodiments, the non-neural body fluid is plasma. In some embodiments, the non-neural cell is a hepatocyte. In some embodiments, the non-neural cell is an astrocyte adipose tissue cell. In some embodiments, the non-neural cell is a Kupffer cell. In some embodiments, the non-neural cell is a liver endothelial cell. In some embodiments, the non-neural body fluid is plasma. In some embodiments, the non-neural body fluid is serum. In some embodiments, the non-neural body fluid is blood.
[0223] Methods for producing recombinant viruses In various embodiments, methods of producing recombinant viruses comprising neuron-specific promoters are also disclosed herein. In some embodiments, nucleic acid sequences, e.g., plasmids, encoding AAV or other viral genomes are used to produce recombinant viruses. In some embodiments, nucleic acid sequences, e.g., plasmids, comprising AAV rep and / or AAV cap genes are also used in preparing AAV or other viruses. Nucleic acid sequences, e.g., plasmids, comprising adenovirus helper function genes are also disclosed herein. In some embodiments, nucleic acids encoding AAV rep, AAV cap and / or adenovirus helper genes can be present in the same structure, e.g., a single plasmid, or they can be present in separate structures. In some embodiments, one or more plasmids are co-transfected with nucleic acid encoding an AAV vector into competent cells, and the cells are cultured to produce recombinant viruses. Optionally, plasmids encoding the AAV viral genome and AAV rep and / or cap genes are introduced into cells permissive for infection by a helper virus for AAV (e.g., adenovirus, E1-deleted adenovirus or herpesvirus). In some embodiments, the rAAV genome assembles into an infectious viral particle with AAV capsid proteins within the cell after transfection. Techniques for producing rAAV particles, in which the packaged AAV genome, rep and cap genes, and helper virus functions are provided to the cell, are known in the art and may include, for example, electroporation. In some embodiments, the production of rAAV includes the following components present within a single cell (referred to herein as a packaging cell): the rAAV vector, the AAV rep and cap genes separate from (i.e., not within) the rAAV vector, and helper virus functions. The generation of pseudotyped rAAV is disclosed, for example, in WO 01 / 83692, which is incorporated herein by reference in its entirety. In various embodiments, the AAV capsid proteins may be modified to enhance delivery of the recombinant vector. Modification of capsid proteins is generally known in the art.See, for example, US Patent Application Publication No. 2005 / 0053922 and US Patent Application Publication No. 2009 / 0202490, the disclosures of which are incorporated herein by reference in their entireties.
[0224] In various embodiments, the general principles of viral vector production can be used to produce the vectors and viruses disclosed herein, such as rAAV. Carter, (1992) Curr. Opinions Biotech., 1533-539; Muzyczka, (1992) Curr. Topics Microbial. Immunol., 158:97-129. Various approaches have been proposed by Ratschin et al.,(1984)Mol.Cell.Biol.,4:2072;Hennonat et al.,(1984)Proc.Natl.Acad.Sci.USA,81:6466;Tratschin et al.,(1985)Mol.Cell.Biol.,5:3251;McLaughlin et al. al.,(1988)J.Virol.,62:1963;Lebkowski et al.,(1988)Mol.Cell.Biol.,7:349;Samulski et al. al. (1989) J. Virol., 63:3822-3828; U.S. Pat. No. 5,173,414; WO 95 / 13365 and corresponding U.S. Pat. No. 5,658,776; WO 95 / 13392; WO 96 / 17947; PCT / US98 / 18600; WO 97 / 09441 (PCT / US96 / 14423); WO 97 / 08298 (PCT / US96 / 13872); WO 97 / 21825 (PCT / US96 / 20777); WO 97 / 06243 (PCT / FR96 / 01064); WO 99 / 11764; Perrin et al. al., (1995) Vaccine, 13:1244-1250; Paul et al., (1993) Hum. Gene Ther., 4:609-615; Clark et al. (1996) Gene Therapy, 3:1124-1132; U.S. Patent No. 5,786,211; U.S. Patent No. 5,871,982; and U.S. Patent No. 6,258,595.The aforementioned documents are incorporated herein by reference in their entireties, with particular emphasis on the sections of the documents relevant to rAAV production.
[0225] An exemplary method for producing packaging cells is to create a cell line that stably expresses all components required for the production of AAV particles. For example, a plasmid (or multiple plasmids) encoding a rAAV vector lacking the AAV rep and cap genes, a selection marker such as the AAV rep and cap genes separate from the rAAV vector, and a neomycin resistance gene is integrated into the genome of the cell. The AAV genome has been introduced into a bacterial plasmid by procedures such as GC tailing (Samulski et al., (1982) Proc. Natl. Acad. Sci. USA, 79:2077-2081), adding a synthetic linker containing a restriction endonuclease cleavage site (Laughlin et al., (1983) Gene, 23:65-73), or by direct blunt-end ligation (Senapathy et al., (1984) J. Biol. Chem., 259:4661-4666). The packaging cell line is then infected with a helper virus, such as adenovirus, and / or a plasmid encoding the helper virus. The advantage of this method is that the cells are selectable and are suitable for large-scale production of rAAV. Another example of a suitable method uses adenovirus or baculovirus, rather than plasmids, to introduce the rAAV vector and / or the rep and cap genes into the packaging cells.
[0226] In some embodiments, the method of producing a recombinant virus includes providing a nucleic acid to be packaged. In some embodiments, the nucleic acid is a plasmid. In other embodiments, the nucleic acid includes a transgene sequence inserted between a first AAV terminal repeat and a second AAV terminal repeat. In some embodiments, the transgene encodes human progranulin (hPGRN). In some embodiments, the method of producing a recombinant virus includes providing one or more additional nucleic acids. In some embodiments, the one or more additional nucleic acids include an AAV rep gene and / or an AAV cap gene. In some embodiments, the one or more additional nucleic acids include an AAV rep gene from AAV serotype 1, AAV serotype 2, AAV serotype 3, AAV serotype 4, AAV serotype 5, AAV serotype 6, AAV serotype 7, AAV serotype 8, or AAV serotype 9. In some embodiments, the one or more additional nucleic acids comprise an AAV cap gene from AAV serotype 1, AAV serotype 2, AAV serotype 3, AAV serotype 4, AAV serotype 5, AAV serotype 6, AAV serotype 7, AAV serotype 8, or AAV serotype 9. In some embodiments, the one or more additional nucleic acids comprise one or more adenovirus helper function genes.
[0227] In some embodiments, the nucleic acid is co-transfected into competent or packaging cells. Methods of co-transfection are known in the art and include, but are not limited to, lipofectamine, electroporation, and polyethylenimine transfection. Competent or packaging cells can be non-adherent cells cultured in suspension or adherent cells. In one embodiment, any suitable packaging cell line can be used, such as HeLa cells, HEK293 cells, and PerC.6 cells (a homologous 293 line). In one embodiment, the packaging cells are human cells. In one embodiment, the packaging cells are HEK293 cells. In one embodiment, the packaging cells are insect cells. In one embodiment, the packaging cells are Sf9 cells. In some embodiments, the method includes culturing the transfected cells to produce recombinant virus. In some embodiments, the method includes recovering the recombinant virus. Methods of recovering recombinant virus include, for example, those disclosed in U.S. Pat. No. 6,143,548 and U.S. Pat. No. 9,408,904. In some embodiments, the recombinant virus is secreted into the cell culture medium and purified from the medium. In some embodiments, the packaging cells are lysed and the contents purified to recover the recombinant virus. In some embodiments, the virus is recovered from the packaging cells by filtration or centrifugation. In some embodiments, the virus is recovered from the packaging cells by chromatography.
[0228] In various embodiments, disclosed herein are cells comprising a nucleic acid disclosed herein, a vector disclosed herein, or a virus disclosed herein. The cells comprising a nucleic acid disclosed herein, a vector disclosed herein, or a virus disclosed herein can be human cells. The cells comprising a nucleic acid disclosed herein, a vector disclosed herein, or a virus disclosed herein can also be insect cells. In some embodiments, the cells comprising a nucleic acid disclosed herein, a vector disclosed herein, or a virus disclosed herein are HEK293 cells. In some other embodiments, the cells comprising a nucleic acid disclosed herein, a vector disclosed herein, or a virus disclosed herein are Sf9 cells.
[0229] In some embodiments, the method of producing a recombinant virus comprises transfecting an insect cell. In some embodiments, the method comprises transfecting an insect cell with a baculovirus comprising a nucleic acid as disclosed herein. In some embodiments, the method comprises transfecting an insect cell with a baculovirus comprising a nucleic acid comprising a transgene sequence inserted between a first AAV terminal repeat and a second AAV terminal repeat. In some embodiments, the method comprises transfecting an insect cell with a baculovirus comprising one or more additional nucleic acids. In some embodiments, the one or more additional nucleic acids comprise an AAV rep gene and / or an AAV cap gene. In some embodiments, the one or more additional nucleic acids comprise an AAV rep gene from AAV serotype 1, AAV serotype 2, AAV serotype 3, AAV serotype 4, AAV serotype 5, AAV serotype 6, AAV serotype 7, AAV serotype 8, or AAV serotype 9. In some embodiments, the one or more additional nucleic acids comprise an AAV cap gene from AAV serotype 1, AAV serotype 2, AAV serotype 3, AAV serotype 4, AAV serotype 5, AAV serotype 6, AAV serotype 7, AAV serotype 8, or AAV serotype 9.c. In some embodiments, the one or more additional nucleic acids comprise one or more adenovirus helper function genes. In some embodiments, the insect cells are cultured under conditions suitable for producing recombinant virus. In some embodiments, the virus is recovered from the insect cells. In some embodiments, the virus is recovered from the insect cells by filtration or centrifugation. In some embodiments, the virus is recovered from the insect cells by chromatography.
[0230] Pharmaceutical Compositions In various embodiments, pharmaceutical compositions are disclosed. In some embodiments, the pharmaceutical compositions comprise one or more of the nucleic acids, vectors and / or viruses disclosed herein. In some embodiments, the pharmaceutical compositions comprise a pharma- ceutically acceptable carrier.
[0231] The nucleic acid, vector and / or recombinant virus (e.g., viral particle) according to the present disclosure can be formulated to prepare a pharma- ceutically useful composition. Exemplary formulations include those disclosed, for example, in U.S. Pat. Nos. 9,051,542 and 6,703,237, which are incorporated by reference in their entirety. The compositions of the present disclosure can be formulated for administration to a mammalian subject, for example, a human. In some embodiments, the delivery system can be formulated for intramuscular, intradermal, mucosal, subcutaneous, intravenous, intrathecal, injectable depot-type device, or topical administration.
[0232] In some embodiments, when the delivery system is formulated as a solution or suspension, the delivery system is in an acceptable carrier, for example, an aqueous carrier. A variety of aqueous carriers can be used, for example, water, buffered water, 0.8% saline, 0.3% glycine, hyaluronic acid, etc. These compositions can be sterilized and / or sterile filtered. The resulting aqueous solution can be packaged for immediate use or lyophilized. In some embodiments, lyophilized preparations are combined with a sterile solution before administration.
[0233] In some embodiments, the compositions, e.g., pharmaceutical compositions, may contain pharma- ceutically acceptable auxiliary substances to approximate physiological conditions, such as pH adjusting and buffering agents, tonicity adjusting agents, wetting agents, etc., e.g., sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, etc. In some embodiments, the pharmaceutical composition contains a preservative. In other embodiments, the pharmaceutical composition does not contain a preservative.
[0234] Methods of Use and Treatment Without being bound by theory, the nucleic acids and other embodiments described herein are used in a method of conditionally expressing a molecule (e.g., a gene of interest) comprising administering to a subject in need thereof an expression system, e.g., a cell comprising a nucleic acid molecule described herein, a vector described herein, wherein a) expression of the gene of interest is increased, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, or 100 fold compared to the expression level of the gene of interest in a subject not receiving such treatment.
[0235] In one embodiment, the term "treating" includes administering an effective dose or effective doses of a composition comprising a nucleic acid, vector, recombinant virus, or pharmaceutical composition disclosed herein to an animal (including a human) in need thereof. If the dose is administered before the onset of the disorder / disease, the administration is prophylactic. If the dose is administered after the onset of the disorder / disease, the administration is therapeutic. In embodiments, an effective dose is a dose that detectably alleviates (eliminates or reduces) at least one symptom associated with the disorder / disease state being treated, delays or prevents progression to the disorder / disease state, delays or prevents progression of the disorder / disease state, reduces the extent of the disease, results in remission (partial or complete) of the disease, and / or extends survival. The term encompasses, but does not require, complete treatment (i.e., cure) and / or prevention. In some embodiments, an effective dose is greater than or equal to 1×10 per milliliter of the virus disclosed herein. 10 ~1×10 15 In some embodiments, an effective dose comprises 1×10 vector genomes per milliliter of virus disclosed herein (vg / ml). 6 ~1×10 10 In some embodiments, an effective dose comprises 1×10 plaque forming units (pfu / ml) per milliliter of virus disclosed herein. 6 ~1×10 9 Examples of conditions contemplated for treatment are described herein.
[0236] In some embodiments, the disease to be treated is caused by a mutation in a gene of interest. In some embodiments, the mutation in the gene of interest is a deletion mutation. In some embodiments, the mutation in the gene of interest is a null mutation. In some embodiments, the mutation in the gene of interest is an indel. In some embodiments, the mutation in the gene of interest is a loss-of-function mutation. In some embodiments, the mutation in the gene of interest is a knockout mutation. In some embodiments, the mutation in the gene of interest results in a loss of protein expression and / or function. In some embodiments, patients in need of treatment with the nucleic acids, vectors and / or viruses disclosed herein are identified by screening for mutations prior to administration. In some embodiments, screening involves obtaining a cell or tissue sample from the subject and sequencing or genotyping one or more loci in the sample to confirm the presence of the mutation. In some embodiments, screening is performed on genetic material from a sample, such as, but not limited to, saliva, blood and / or skin cells.
[0237] In some embodiments, the nucleic acids, vectors, recombinant viruses or pharmaceutical compositions disclosed herein are used in the manufacture of a medicament for treating a subject in need thereof, in embodiments, the subject is afflicted with a disorder caused by one or more mutations in a gene of interest.
[0238] In various embodiments, the nucleic acid, vector, recombinant virus, or pharmaceutical composition disclosed herein may be delivered to a subject in need thereof by intravenous administration, direct brain administration (e.g., intrathecal, intracerebral, and / or intraventricular administration), intranasal administration, intraaural administration, or intraocular route of administration, or any combination thereof. In some embodiments, the nucleic acid, vector, recombinant virus, or pharmaceutical composition is delivered by intrathecal administration. In some embodiments, the nucleic acid, vector, recombinant virus, or pharmaceutical composition is delivered by intracerebral or intraventricular administration route. In some embodiments, the administered nucleic acid, vector, recombinant virus, or pharmaceutical composition is ultimately delivered to the brain, spinal cord, peripheral nervous system, and / or CNS, either directly or by transfer following administration to another tissue or bodily fluid (e.g., blood).
[0239] Without being bound by theory, in some embodiments, the methods disclosed herein may rescue cells carrying a mutation in a gene encoding a polypeptide that results in a non-functional polypeptide. In some embodiments, the method of expressing a molecule, e.g., a protein or a ribonucleic acid (e.g., siRNA), comprises delivering a nucleic acid, viral vector, virus, or pharmaceutical composition disclosed herein to a cell. In some embodiments, the cell is a neuronal cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the neuronal cell is a neuron. In some embodiments, the delivery is in vitro. In some embodiments, the delivery is ex vivo. In some embodiments, the delivery is by systemic administration. In some embodiments, the delivery is localized. In some embodiments, the delivery is by direct application to a target tissue. In some embodiments, the target tissue is the brain. In some embodiments, the delivery is by injection into the brain. In some embodiments, the delivery is by intrathecal administration. Without being bound by theory, the methods disclosed herein may reduce lipofuscin deposition, astrocyte and microglial activation and / or inflammation in the brains of humans or mice with mutations in a gene of interest, thus providing a potential benefit to subjects in need thereof.
[0240] In various embodiments, the nucleic acids, vectors, viruses and pharmaceutical compositions disclosed herein can be used to treat a disorder. In some embodiments, the nucleic acids, vectors, viruses and / or pharmaceutical compositions disclosed herein can be used in the manufacture of a pharmaceutical agent for treating a disorder. In some embodiments, the disorder is caused by one or more mutations in a gene of interest.
[0241] Also provided herein is a kit comprising a nucleic acid molecule described herein, a vector described herein, a recombinant virus described herein, a cell described herein, or a pharmaceutical composition described herein.
[0242] The details of one or more embodiments of the present disclosure are set forth in the accompanying description above. Although any methods and materials similar or equivalent to those described herein can be used to practice or test the present disclosure, the preferred methods and materials are described below. Other features, objects, and advantages of the present disclosure will be apparent from the description and claims. In this specification and the appended claims, the singular form includes plural references unless otherwise clearly indicated by the context. Unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by those skilled in the art to which the disclosure belongs. All patents and publications cited herein are incorporated by reference where applicable, unless otherwise indicated. The following examples are presented to more particularly illustrate preferred embodiments of the present disclosure. These examples should not be construed in any way as limiting the scope of the disclosed subject matter, which is defined by the appended claims. EXAMPLES
[0243] Example 1: TRE reporter All of the TRE reporter constructs were modified from pLVX-TetOne-Puro (TaKaRa) by removing the hPGK promoter and the Tet-On 3G coding region and inserting a firefly luciferase coding sequence into the multiple cloning site downstream of the TRE3Gs promoter, followed by deleting multiple tetracycline response elements (TREs) and reducing the repeats from 7 to 3, 2 and 1, generating 3×TRE, 2×TRE and 1×TRE reporters, respectively.
[0244] Example 2: Human and mouse SCN1A reporters For the human SCN1A promoter reporter, a fragment of the 2673 base pair (bp) promoter + 242 bp exon sequence of the human scn1a gene, which correlates to GRCh38 / hg38chr2:166127806-166130720 (negative strand), was PCR amplified from human genomic DNA and cloned upstream of the firefly luciferase reporter gene. For the mouse SCN1A promoter reporter, a sequence containing the 3064 bp promoter and 256 bp exon sequence of the mouse scn1a gene was PCR amplified from genomic DNA of the C57BL / 6 mouse strain. The sequence correlates to the mouse genomic reference sequence GRCm38 / mm10chr2:66409862-66413181 (negative strand) and was cloned directly upstream of the firefly luciferase reporter gene in the vector.
[0245] Example 3: Reporter constructs for normalization The Renilla luciferase gene was cloned directly downstream of the hGPK promoter in a separate construct. Renilla luciferase was used to normalize firefly luciferase activity in all transient transfection reporter assays.
[0246] Example 4: Cas9 activator constructs Mininuclease-dead Staphylococcus aureus Cas9 (miniSa-dCas9) (Ma et al. 2018) was synthesized and cloned directly downstream of the elongation factor 1 alpha (EF1a) promoter in a mammalian expression vector. In some cases, it was cloned into a vector and driven by the mammalian ubiquitin C (UbC) promoter. Transcription activation domain (TAD) (Chavez et al. 2015, Ma et al. 2018) coding sequences including VP64, p65 and Rta (VPR) and VPR3 were ligated to the 3' end of the Cas9 gene. A human influenza hemagglutinin (HA) tag sequence was ligated in frame to the TAD sequence. Three nuclear localization sequences (NLS) were placed at the N-terminus of Sa-dCas9, between Cas9 and the TAD, and between the TAD and the HA tag.
[0247] Example 5: sgRNA constructs For the design of sgRNAs, potential Sa-Cas9 binding sites were identified in silico by determining the Sa-Cas9 protospacer adjacent motif (PAM) site NNGRRT (R can be either G or A) within the TRE promoter, human or mouse scn1a proximal promoter region, including the sequence after the transcription start site. The 5' 20 nucleotides of each PAM were cloned into a vector containing SaCas9 transactivating crisprRNA (tracrRNA). (Ma et al. 2018) All gRNA-tracrRNAs were driven by the U6 promoter.
[0248] Example 6: Zinc finger protein activator constructs The Sa-dCas9 coding sequence in Sa-dCas9-TAD15 was replaced with a zinc finger protein coding sequence. In E2C-TAD15, a humanized coding sequence of a zinc finger protein that binds to the E2C site of the erb2 / Her2 promoter region was synthesized according to the publication (Beerli et al. PNAS 1998) and linked to the 5' end of the TAD15 sequence. In ZFP-C-TAD15, a humanized coding sequence of ZFP-C (Zeitler et al. 2019) was synthesized and linked to the 5' end of the TAD15 sequence. According to (Beerli, PNAS 1998, Segal et al. PNAS 1999), the zinc finger protein Nav-ZF2 with the sequence 5'-GGC-GAG-GAT-GAA-GCC-GAG-3' (SEQ ID NO: 90) targeting the SCN1A gene was constructed on the Sp1C zinc finger framework and linked to the 5' end of the TAD13, TAD14 and TAD15 sequences, respectively. One nuclear localization sequence (NLS) was placed at the N-terminus of the zinc finger protein and another one between the TAD and the HA tag.
[0249] Example 7: Cells and cell transfection HEK293T cells were printed at 25,000 cells / well in 100 μl DMEM (Life Technologies, 11965092) supplemented with 10% heat-inactivated fetal bovine serum, 1× GlutaMAX (Life Technologies, 35050061) and 1× penicillin / streptomycin (Life Technologies, 10378016) onto poly-D-lysine coated 96-well black clear bottom plates at 37°C and 5% CO2. The following day, cultures were transfected using Lipofectamine 3000 (Life Technologies, L300015) according to the manufacturer's instructions. The amounts of each plasmid per transfection sample were as follows: 20 ng miniSad-Cas9-VPR, 60 ng sgRNA plasmid, 25 ng SCN1A promoter reporter plasmid and 0.6 ng hGPK-Renilla luciferase plasmid. After 2 days, luciferase activity was measured.
[0250] Example 8: Recording of luciferase activity Two days after transfection, the culture medium was removed and 50ul fresh DMEM complete medium was added to each well. Luciferase activity was measured using the DualGlo Luciferase Assay System (Promega, E2920) according to the manufacturer's instructions. Briefly, 50ul of firefly luciferase substrate was added to each well and samples were mixed. After 10 minutes at room temperature, firefly luciferase activity was recorded using Envision (Perkin Elmer). Subsequently, 50ul of Renilla luciferase substrate was added to each well and mixed. After 10 minutes at room temperature, Renilla luciferase activity was recorded using Envision (Perkin Elmer).
[0251] Example 9: Packaging of lentiviral zinc finger protein activators 293T cells were plated on poly-D-lysine coated 10 cm dishes at 4,000,000 cells / dish using 10 ml culture medium. One day later, the cultures were transfected with lentiviral vectors carrying zinc finger protein transactivators and helper plasmids (pMD2.G, psPAX2) using Lipofectamine 3000 (Life Technologies, L300015). After 46 hours, the medium was harvested and the cell debris was removed from the medium by centrifugation at 10,000×g for 5 min. The virus in the supernatant was concentrated by ultracentrifugation at 49,000×g for 90 min. The pellet was resuspended in Neurobasal Plus medium (Thermo Fisher Scientific, A3582901), aliquoted, and frozen at −80° C. until use.
[0252] Example 10: Mouse GABAergic neuronal cultures and neuronal gene activation GABAergic neurons were dissected from the medial ganglionic eminence (MGE) of mice at embryonic day 13. MGE were dissociated into single cell suspensions using a papain dissociation system (Worthington Biochemical Corporation) according to the manufacturer's protocol, and cells were plated at 15,000 neurons / well in Neurobasal Plus medium onto a glial support layer on a 96-well plate coated with poly-D-lysine. Neuronal cultures were infected with lentivirus for 4 h on day 2 or 3 in vitro (DIV). Each virus dose was chosen to achieve 70-90% infection efficiency. After infection, cultures were washed three times in plain Neurobasal medium (Thermo Fisher Scientific), then conditioned medium / fresh medium (50 / 50) was returned to the plate, and cultures were maintained as described. At DIV7, cells were lysed and RNA was harvested. SCN1A and MAP2 transcripts were analyzed by qRT-PCR. SCN1A levels in each sample were normalized to MAP2 levels in the same sample. Increase in SCN1A transcripts by transcriptional activators is measured by the fold increase in normalized SCN1A mRNA levels relative to normalized SCN1A mRNA levels in virus-free cells.
[0253] Example 11. Transcription factors for activation of the TRE promoter TAD1, TAD2+ (TAD2 with three nuclear localization signal (NLS) sequences), TAD2- (TAD2 with two NLS sequences) and TAD3 activate TRE-containing promoters when linked to mini-Sa-Cas9. Plasmids carrying sgRNA targeting the TRE sequence, mini-Sa-dCas9-activator, TRE promoter-luciferase and GPK-Renilla luciferase plasmids were transfected into HEK293T cells. Luciferase activity was read out after 2 days. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. Relative promoter activity triggered by Sa-dCas9 activator was determined by calculating the fold change from the control guide RNA sample. 1×TRE promoter contains one TRE site; 3×TRE promoter carries three repeats of the TRE site. VPR3 (Ma et al. 2018) was also linked to mini-Sa-Cas9 and used as a positive transcription activator control. The sequence of the sgRNA targeting the TRE is listed in SEQ ID NO:64.
[0254] As shown in FIG. 1, TAD1, TAD2+, and TAD2- exhibit higher activity in activating the TRE promoter compared to VPR3.
[0255] Example 12. Transcription factors for activation of the SCN1A promoter The miniSa-dCas9-activator (TAD1-TAD7), mouse scn1a promoter-luciferase, gRNA targeting mouse scn1a promoter and GPK-Renilla luciferase plasmids were transfected into HEK293T cells. After 2 days, luciferase activity was read. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. Each miniSa-dCas9-activator was transfected at 20, 6.7 and 2.2 ng / well, and the results showed a dose-dependent effect. The sequence of sgRNA3 corresponds to SEQ ID NO: 65. The results are shown in Figure 2.
[0256] The miniSa-dCas9-activator (TAD7-TAD10), mouse scn1a promoter-luciferase, gRNA targeting mouse scn1a promoter and GPK-Renilla luciferase plasmids were transfected into HEK293T cells. After 2 days, luciferase activity was read. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. Each miniSa-dCas9-activator was transfected at 20, 6.7 and 2.2 ng / well, and the results showed dose-dependent activation when co-transfected with active guide RNA sgRNA3. When miniSa-dCas9 was co-transfected with inactive guide RNA sgRNA11, the promoter was not activated. The sequences of sgRNA3 and sgRNA11 correspond to SEQ ID NO:65 and SEQ ID NO:66, respectively. The results are shown in Figure 3.
[0257] Example 13. Transcription factors for activation of the SCN1A promoter when paired with sgRNA Mini-Sa-dCas9-TAD driven by ubiquitin C promoter activates mouse scn1a promoter and TRE-containing promoter when paired with active sgRNA. HEK293T cells were transfected with various amounts of mini-Sa-dCas9-activator (TAD2, TAD9 and VPR), mouse scn1a promoter-luciferase or TRE-containing promoter, TRE sgRNA for TRE-containing promoter respectively; sgRNA3 for mSCN1A promoter and GPK-Renilla luciferase. After 2 days, luciferase activity was read.
[0258] The miniSa-dCas9-activator, mouse scn1a promoter-luciferase or human scn1a promoter-luciferase, gRNA targeting the scn1a promoter, and GPK-Renilla luciferase plasmids were transfected into HEK293T cells. After 2 days, luciferase activity was read. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. The newly created transcription activation domains (TAD9, TAD11, TAD14, and TAD16) have activity equal to or greater than VPR (Chavez et al. 2015) and stronger than VP64. The results are shown in Figure 5.
[0259] The miniSa-dCas9-activator, TRE-containing promoter or human scn1a promoter-luciferase, gRNA targeting the scn1a promoter and GPK-Renilla luciferase plasmids were transfected into HEK293T cells. After 2 days, luciferase activity was read out. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. The newly created transcription activation domains have activity equal to or greater than VPR (Chavez et al. 2015) and stronger activity than VP64. Figure 6 shows the activity of TAD9, TAD11, TAD12, TAD13, TAD14, TAD15 compared to VP64 and full VPR for activation of TRE and human SCN1A promoter.
[0260] The miniSa-dCas9-activator, mouse scn1a promoter-luciferase, sgRNA targeting mouse scn1a promoter and GPK-Renilla luciferase plasmids were transfected into HEK293T cells. Luciferase activity was read after 2 days. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. The sequence of sgRNA42 corresponds to SEQ ID NO: 67. Figure 7 shows the activity of TAD9, TAD11, TAD12, TAD13, TAD14, TAD15 compared to VP64 and full VPR in activating the mouse SCN1A promoter with different active sgRNAs.
[0261] Example 14. Transcription factors for activation of TRE reporter genes in U2OS cell lines The miniSa-dCas9-activator, TRE-containing promoter or human scn1a promoter-luciferase, gRNA sequence targeting TRE, and GPK-Renilla luciferase plasmid were transfected into U2OS cells. After 2 days, luciferase activity was read. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. As shown in Figure 8, TAD8a, TAD9, and TAD15 show higher activity than VP64, and TAD15 shows the highest activity of all.
[0262] Example 15. Activity of transcription factors containing zinc finger proteins Zinc finger protein ZFP-C (SEQ ID NO: 70, "6CAG repeat" is disclosed as SEQ ID NO: 93) targeting 6 CAG repeats and zinc finger protein E2C (SEQ ID NO: 71) targeting human erb2 gene promoter were linked to TAD15. Each transcription activator was co-transfected with a reporter carrying two E2C elements (E2C 2x) or a reporter carrying 8 CAG repeats (CAG24) (SEQ ID NO: 94) and GPK-Renilla luciferase was transfected into HEK293T cells. Luciferase activity was read out after 2 days. Promoter activity was calculated by normalizing firefly luciferase activity with Renilla luciferase activity. Transcription factors specifically activated the promoters. That is, E2C-TAD15 activated only E2C element-containing promoters; ZFP-C-TAD15 activated only CAG repeat-containing promoters. Figure 9 shows the activity of TAD15 linked to zinc finger proteins to enhance gene transcription.
[0263] The zinc finger protein E2C (SEQ ID NO: 71) recognizes an element in the human erb2 gene promoter (Beerli et al 1998). Erb2 encodes the receptor protein Her2. E2C-TAD15 was transfected into HEK293T cells. Transfected cells were identified by staining for hemagglutinin (HA) tagged on E2C-TAD15. Her2 expression levels were quantified by immunostaining for Her2 on the cell surface. Cells with E2C-TAD15 (HA positive) showed more than three-fold higher Her2 levels compared to non-transfected cells (HA negative). The results are shown in Figure 10.
[0264] TAD13, TAD14 and TAD15 were each ligated with the zinc finger protein Nav-ZF2 (SEQ ID NO: 61) targeting the SCN1A promoter. They were purged into lentivirus and used to infect cultured neurons, achieving 70-90% transduction rates. After 4 days, cells were lysed and RNA was harvested. SCN1A and MAP2 transcripts were analyzed by qRT-PCR. Increase in SCN1A transcripts by transcriptional activators is measured by the fold increase in normalized SCN1A mRNA levels relative to normalized SCN1A mRNA levels in cells without virus. Activity of TAD15 in activating endogenous neuronal genes in mouse GABAergic neurons is shown in Figure 11.
[0265] It is understood that the examples and embodiments described herein are intended to be merely illustrative and that various modifications or changes in light thereof will be suggested to those skilled in the art and are to be included within the spirit and scope of this application and the appended claims.
[0266] Furthermore, when features or aspects of the invention are described in terms of a Markush group, those skilled in the art will recognize that the invention may also be described in terms of any individual member or subgroup of members of the Markush group.
[0267] All publications, patent applications, patents, and other references mentioned herein are expressly incorporated by reference in their entirety to the same extent as if each was individually incorporated by reference (or as determined by context). In the case of conflict, the present specification, including definitions, will control.
Claims
1. A transcription factor comprising a DNA binding domain and at least three transcription activation domains (TADs), said TADs can be the same or different.
2. 2. The transcription factor of claim 1, wherein the TADs comprise: i) at least one acidic TAD and at least one Q-rich TAD; or ii) at least one acidic TAD and at least one P-rich TAD.
3. 3. The transcription factor of claim 2, wherein the acidic TAD comprises one or more copies of a TAD selected from VP16, VP64, VP7, ATF6, TFE3, ATF6-11, or fragments thereof.
4. 3. The transcription factor of claim 2, wherein the Q-rich TAD comprises one or more copies of a TAD selected from Oct2-Q, SP1-Q1, or fragments thereof.
5. 3. The transcription factor of claim 2, wherein the P-rich TAD comprises one or more copies of a TAD selected from TFAP2-P, Oct2-P, or fragments thereof.
6. 2. The transcription factor of claim 1, wherein the TAD is selected from the group consisting of full length VP64, one or more repeats of VP16, full length or partial p65, full length RTA, full length VP7, full length or partial TEF3, full length or partial ATF6-11, full length or partial ATF6 acidic, full length or partial SP1-rich Q, full length or partial Oct2-rich Q, full length or partial Oct2-rich P, full length or partial TFAP2-rich P, an active fragment of VP64, a transcriptionally active fragment of p65, a transcriptionally active fragment of RTA, an active fragment of VP7, an active fragment of TEF3, an active fragment of ATF6-11, an active fragment of ATF6 acidic, an active fragment of SP1-rich Q, an active fragment of Oct2-rich Q, an active fragment of Oct2-rich P, an active fragment of TFAP2-rich P, and any combination thereof.
7. 2. The transcription factor of claim 1, wherein the total length of the TADs is less than 2000 aa, less than 1500 aa, less than 1000 aa, less than 750 aa, less than 500 aa, less than 300 aa, less than 250 aa, less than 200 aa, or less than 150 aa.
8. 2. The transcription factor of claim 1, comprising a sequence having at least 95% sequence identity to any of SEQ ID NOs: 1 to 28 or a sequence having at least 80%, 85%, 90%, preferably 95% sequence identity to SEQ ID NOs: 1 to 28.
9. 2. The transcription factor of claim 1, comprising a sequence having any one of SEQ ID NOs: 1 to 18 or a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity thereto.
10. The transcription factor of claim 1, wherein the nucleic acid sequence encoding the transcription factor comprises a sequence of SEQ ID NO: 30 to 47 or a sequence having at least 80%, at least 85%, at least 90%, or at least 95% sequence identity thereto.
11. 2. The transcription factor of claim 1, wherein the DNA binding domain is linked to the transcription factor of claim 1 with or without a linker.
12. The transcription factor of claim 11 , wherein the DNA-binding domain is linked to the N-terminus of the transcription factor directly or via a linker.
13. The transcription factor of claim 11 , wherein the DNA-binding domain is linked to the C-terminus of the transcription factor directly or via a linker.
14. The transcription factor of claim 1 , wherein the DNA binding domain binds to a genomic region of a gene of interest.
15. 15. The transcription factor of claim 14, wherein the DNA binding domain is a gRNA / Cas complex, a transcription activator-like (TAL) effector, or a zinc finger protein.
16. 16. The transcription factor of claim 15, wherein the gRNA / Cas complex comprises a Cas molecule that is or is derived from S. pyogenes Cas9, C. jejune Cas9, S. aureus Cas9, or Deltaproteobacteria (Dpb) CasX.
17. 17. The transcription factor of claim 16, wherein the Cas molecule lacks one or more activities, and optionally the one or more activities is a cleavage activity.
18. 17. The transcription factor of claim 16, wherein the Cas molecule is encoded by a nucleic acid molecule comprising fewer than 4,000 nucleotides, optionally wherein the Cas molecule is encoded by a nucleic acid molecule comprising fewer than 3,500 nucleotides or fewer than 2,500 nucleotides.
19. 17. The transcription factor of claim 16, wherein the Cas molecule is or is derived from S. aureus Cas9.
20. 20. The transcription factor of claim 19, wherein the Cas molecule comprises one or more amino acid deletions compared to a wild-type Cas molecule sequence.
21. The transcription factor of claim 15 , wherein the DNA binding domain is a zinc finger protein.
22. 22. The transcription factor of claim 21, wherein the DNA binding domain binds to the promoter region of the gene of interest.
23. 23. The transcription factor of claim 22, wherein the DNA binding domain is a zinc finger protein that binds to a 6, 9, 12, 15, 18, 21, or 24 bp sequence in the promoter region of the gene of interest.
24. A nucleic acid molecule encoding one or more components, optionally all of the components, of a transcription factor according to any one of claims 1 to 23.
25. A vector comprising the nucleic acid molecule of claim 24.
26. 26. The vector of claim 25, comprising a first promoter operably linked to a nucleic acid sequence encoding the DNA-binding domain.
27. 27. The vector of claim 26, comprising a promoter that drives transcription in a human cell.
28. 27. The vector of claim 26, wherein the first promoter is operative in a neuron.
29. 29. The vector of claim 28, wherein the neuron is a GABAergic neuron or an inhibitory neuron or an inhibitory interneuron.
30. 30. The vector of claim 29, wherein the neuron is a parvalbumin-positive GABAergic neuron, or a somatostatin-positive GABAergic neuron, or a vasoactive intestinal peptide-positive GABAergic neuron.
31. a. polyA sequence; b. an intron sequence; or c. Enhancer sequence 26. The vector of claim 25, further comprising one or more of:
32. 26. The vector of claim 25, further comprising a regulatory element that controls the production and / or degradation of the transcription factor.
33. 33. The vector of claim 32, wherein the regulatory element is a minigene linked to the transcription factor, the minigene comprising a splice regulator binding sequence, and in the presence of the splice regulator, the minigene undergoes splicing that results in increased or decreased expression of the transcription factor.
34. 33. The vector of claim 32, wherein the regulatory element is a destabilizing domain, and in the presence of a small molecule that specifically binds to the destabilizing domain, expression of the transcription factor is decreased.
35. 26. The vector of claim 25, which is an adenoviral vector, an adeno-associated viral (AAV) vector, or a lentiviral vector, or an adenoviral vector or a herpes simplex viral vector.
36. 36. The vector of claim 35, which is an AAV vector.
37. 37. The vector of claim 36, which is an AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, or AAV.rhlO vector.
38. A cell comprising the vector of claim 25.
39. 26. A method for selective expression of a transgene in a subject, comprising administering to the subject the vector of claim 25.
40. 1. A pharmaceutical composition for use in a method of treating a disease associated with a mutation in a gene of interest, comprising: The pharmaceutical composition comprises the vector of claim 25, The method comprises administering to a subject in need thereof the vector of claim 25.