DNA-Binding Domain Transactivators and Their Uses
Recombinant AAV vectors with fusion proteins targeting regulatory regions of genes like SCN1A enhance expression, addressing limitations of traditional vectors and providing a therapeutic boost for diseases like Dravet syndrome.
Patent Information
- Application Number
- CN202080021751.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-25
- Filing Date
- 2020-02-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-02-24
AI Technical Summary
In the prior art, gene enhancement methods based on AAV vectors are difficult to effectively treat diseases caused by insufficient haploid target genes, especially symptoms such as deravir syndrome caused by defective SCN1A gene expression.
These fusion proteins containing DNA binding domains and transcriptional regulatory subdomains, such as zinc finger proteins (ZFP) and transactivator domains, are used to deliver these fusion proteins to target cells by targeting regulatory regions that bind to target genes, by targeting the regulatory regions of target genes, and by using recombinant AAV vectors.
Significantly increase the expression level of target genes, such as the expression of SCN1A gene, to achieve the effect of treating haploid deficiency-related diseases, such as deravir syndrome.
Smart Images

Figure CN113710693B_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims the benefit of the filing date of U.S. Provisional Application Serial No. 62 / 810,005, entitled "ZINC FINGER PROTEIN TRANSACTIVATORS AND USES THEREOF", filed on February 25, 2019, the entire content of which is incorporated herein by reference. Background of the Invention
[0003] Regulation of target gene expression has emerged as a major area of biomedical research. Upregulation of gene expression can correct haploinsufficient phenotypes resulting from reduced gene expression. Haploinsufficiency typically results when one or more loss-of-function mutations are present in at least one copy of a gene. AAV-based gene augmentation methods for treating diseases associated with haploinsufficiency are hindered by the packaging capacity of conventional rAAV vectors. Summary of the Invention
[0004] Aspects of the present disclosure relate to isolated nucleic acids and recombinant AAV vectors for gene delivery. The present disclosure is based, in part, on compositions (e.g., rAAV vectors and rAAVs) and methods for regulating the expression of a target gene, where the target gene is haploinsufficient, such as SCN1A. In some embodiments, the present disclosure provides a fusion protein comprising a DNA binding domain (such as a Cys2-His2 zinc finger protein (ZFP)) and a transcriptional regulator domain. In some embodiments, the compositions described herein comprise a fusion protein comprising a DNA binding domain (e.g., ZFP, transcription activator-like effector (TALE) domain, etc.) fused to a transcriptional regulator domain. In some embodiments, the fusion proteins described herein increase the expression of a target gene (e.g., SCN1A) and can thus be used to treat diseases characterized by a defect in target gene expression in a cell or subject compared to a normal cell or subject (e.g., diseases associated with haploinsufficiency of the target gene).
[0005] Accordingly, in some aspects, the present disclosure provides an isolated nucleic acid comprising a transgene configured to express at least one DNA binding domain fused to at least one transcriptional regulator domain, where the DNA binding domain binds to the target gene or a regulatory region (e.g., enhancer sequence, promoter sequence, repressor sequence, etc.) of the target gene (e.g., in a subject or cell), where the target gene encodes a voltage-gated sodium channel (e.g., Na v1.1). In some embodiments, the target gene is the SCN1A gene. In some embodiments, the transgene is flanked by adeno-associated virus (AAV) inverted terminal repeats (ITRs). In some embodiments, at least one DNA binding domain binds to the target gene (e.g., in a subject or cell), and the transcriptional regulator domain modifies (e.g., upregulates) the expression of the target gene.
[0006] In some aspects, the present disclosure provides a recombinant AAV (rAAV) comprising: a nucleic acid comprising at least one DNA binding domain fused to at least one transcriptional regulator domain, wherein the DNA binding domain binds to the target gene or a regulatory region of the target gene (e.g., in a subject or cell), wherein the target gene encodes a voltage-gated sodium channel (e.g., Nav1.1); and at least one capsid protein. In some embodiments, the target gene is the SCN1A gene. In some embodiments, the transgene is flanked by AAV inverted terminal repeats (ITRs).
[0007] In some embodiments, at least one DNA binding domain binds to the target gene (e.g., in a subject or cell), and the transcriptional regulator domain modifies (e.g., upregulates) the expression of the target gene in the subject.
[0008] In some embodiments, at least one DNA binding domain binds to the untranslated region of the target gene. In some embodiments, the DNA binding domain binds to a regulatory region of the target gene, optionally an enhancer sequence, a promoter sequence, and / or a repressor sequence.
[0009] In some embodiments, the DNA binding domain binds between 2 bp and 2000 bp upstream or upstream or downstream 2 bp and 2000 bp of the regulatory region of the target gene (e.g., enhancer sequence, promoter sequence, and / or repressor sequence, etc.).
[0010] In some embodiments, at least one DNA binding domain encodes a zinc finger protein (ZFP), a transcription activator-like effector (TALE), a dCas protein (e.g., dCas9 or dCas12a), and / or a homeodomain. In some embodiments, at least one DNA binding domain binds to the nucleic acid sequence listed in any one of SEQ ID NOs: 5-7. In some embodiments, at least one DNA binding domain is a zinc finger protein comprising a recognition helix encoded by a nucleic acid having a sequence listed in any one of SEQ ID NOs: 11-16, 23-28, or 35-40. In some embodiments, at least one DNA binding domain is a zinc finger protein comprising the amino acid sequence listed in any one of SEQ ID NOs: 17-22, 29-34, or 41-46.
[0011] In some embodiments, at least one DNA-binding domain is a zinc finger protein that includes a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 11, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 12, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 13, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 14, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 15, and / or a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 16. In some embodiments, at least one DNA-binding domain is a zinc finger protein that includes the amino acid sequence of SEQ ID NO: 57. In some embodiments, the ZFP that binds to the SCN1A gene comprises at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% sequence identity to the amino acid sequence of SEQ ID NO: 57.
[0012] In some embodiments, at least one DNA-binding domain is a zinc finger protein that includes a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 23, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 24, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 25, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 26, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 27, and / or a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 28. In some embodiments, at least one DNA-binding domain is a zinc finger protein that includes the amino acid sequence of SEQ ID NO: 59. In some embodiments, the ZFP that binds to the SCN1A gene comprises at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% sequence identity to the amino acid sequence of SEQ ID NO: 59.
[0013] In some embodiments, at least one DNA binding domain is a zinc finger protein that includes a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 35, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 36, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 37, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 38, a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 39, and / or a recognition helix encoded by a nucleic acid comprising SEQ ID NO: 40. In some embodiments, at least one DNA binding domain is a zinc finger protein that includes the amino acid sequence of SEQ ID NO: 61. In some embodiments, the ZFP that binds to the SCN1A gene has at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% sequence identity to the amino acid sequence of SEQ ID NO: 61.
[0014] In some embodiments, at least one DNA binding domain is a zinc finger protein that includes a recognition helix comprising the amino acid sequence of SEQ ID NO: 17, a recognition helix comprising the amino acid sequence of SEQ ID NO: 18, a recognition helix comprising the amino acid sequence of SEQ ID NO: 19, a recognition helix comprising the amino acid sequence of SEQ ID NO: 20, a recognition helix comprising the amino acid sequence of SEQ ID NO: 21, and / or a recognition helix comprising the amino acid sequence of SEQ ID NO: 22.
[0015] In some embodiments, at least one DNA binding domain is a zinc finger protein that includes a recognition helix comprising SEQ ID NO: 29, a recognition helix comprising SEQ ID NO: 30, a recognition helix comprising SEQ ID NO: 31, a recognition helix comprising SEQ ID NO: 32, a recognition helix comprising SEQ ID NO: 33, and / or a recognition helix comprising SEQ ID NO: 34.
[0016] In some embodiments, at least one DNA binding domain is a zinc finger protein that includes a recognition helix comprising SEQ ID NO: 41, a recognition helix comprising SEQ ID NO: 42, a recognition helix comprising SEQ ID NO: 43, a recognition helix comprising SEQ ID NO: 44, a recognition helix comprising SEQ ID NO: 45, and / or a recognition helix comprising SEQ ID NO: 46.
[0017] In some embodiments, at least one DNA-binding domain is a catalytically inactive CRISPR-associated protein (Cas protein). In some embodiments, the catalytically inactive Cas protein (or "dead Cas protein") is a dCas9 or dCas12 protein. In some embodiments, the nucleic acid or rAAV further comprises at least one guide nucleic acid (e.g., guide RNA or gRNA). In some embodiments, the guide nucleic acid comprises a spacer sequence targeting SCN1A. In some embodiments, the guide nucleic acid comprises a spacer sequence having a nucleotide sequence of any one of SEQ ID NO: 85, 86, 89, 90, 93, or 94. In some embodiments, the guide nucleic acid comprises a nucleotide sequence of any one of SEQ ID NO: 83-94. In some embodiments, the guide nucleic acid is encoded by a nucleic acid sequence listed in any one of SEQ ID NO: 83-94.
[0018] In some embodiments, at least one transcriptional regulator domain is a transactivation domain comprising a VP16 domain, a VP64 domain, an Rta domain, a p65 domain, an Hsf1 domain, or any combination thereof, such as a VPR domain (VP64 + p65 + Rta1 domain). In some embodiments, at least one transcriptional regulator domain is encoded by a nucleic acid sequence listed in SEQ ID NO: 47. In some embodiments, at least one transactivation domain comprises the amino acid sequence listed in SEQ ID NO: 48.
[0019] In some embodiments, the ITR flanking the transgene comprises an ITR selected from the group consisting of: AAV1 ITR, AAV2 ITR, AAV3 ITR, AAV4 ITR, AAV5 ITR, AAV6 ITR, AAV8 ITR, AAVrh8 ITR, AAV9 ITR, AAV10 ITR, or AAVrh10 ITR. In some embodiments, the ITR is a ΔTR or mTR.
[0020] In some embodiments, the transgene of the isolated nucleic acid is operably linked to a promoter. In some embodiments, the promoter is a tissue-specific promoter. In some embodiments, the tissue-specific promoter is a neuronal promoter, such as SST, NYP, phosphate-activated glutaminase (PAG), vesicular glutamate transporter-1 (VGLUT1), glutamate decarboxylase 65 and 57 (GAD65, GAD67), synapsin I, a-CamKII, Dock10, Prox1, parvalbumin (PV), somatostatin (SST), cholecystokinin (CCK), calretinin (CR), or neuropeptide Y (NPY).
[0021] In some embodiments, the transgenic DNA-binding domain is fused to the transcriptional regulator domain via a linker domain. In some embodiments, the linker domain is a flexible linker, such as a glycine-rich linker or a glycine-serine linker; or a cleavable linker, such as a photocleavable linker or an enzyme (e.g., protease)-cleavable linker.
[0022] In some embodiments, the isolated nucleic acid comprises a transgene encoding multiple DNA-binding domains (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 DNA-binding domains). In some embodiments, the isolated nucleic acid comprises a transgene encoding multiple transcriptional regulator domains (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 transcriptional regulator domains).
[0023] In some embodiments, the isolated nucleic acid or rAAV is expressed in a cell or subject characterized by abnormal expression or haploinsufficiency (e.g., increased or decreased expression) of a target gene relative to a normal cell or subject. In some embodiments, the isolated nucleic acid or rAAV is expressed in a cell or subject characterized by insufficient expression (e.g., decreased) of a target gene relative to a normal cell or subject. In some embodiments, the target gene of the isolated nucleic acid or rAAV is SCN1A.
[0024] In some embodiments, the AAV capsid serotype is selected from the group consisting of: AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAV9, AAV10, AAVrh10, or AAV.PHPB.
[0025] In some aspects, the present disclosure provides methods of regulating the expression of a target gene. In some embodiments, the methods of the present disclosure comprise administering to a cell or subject expressing a target gene an isolated nucleic acid or rAAV as described herein, wherein the subject is haploinsufficient for the target gene (e.g., haploinsufficient for SCN1A). For example, in some embodiments, the expression of the target gene (such as SCN1A) in the cell or subject is defective (e.g., decreased) relative to the expression of the target gene in a normal cell or subject. In some embodiments, the cell to which the isolated nucleic acid or rAAV is administered is a neuron. In some embodiments, the neuron is a GABAergic neuron.
[0026] In some embodiments, administration of the isolated nucleic acid or rAAV results in an increase in target gene expression (e.g., SCN1A expression) of at least 2-fold, at least 10-fold, at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, or at least 100-fold relative to a subject to whom the isolated nucleic acid or rAAV has not been administered. In some embodiments, administration of the isolated nucleic acid or rAAV results in an increase in target gene expression (e.g., SCN1A expression) of at least 2-fold, at least 10-fold, at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, or at least 100-fold relative to the expression of the target gene (e.g., SCN1A) in the subject prior to administration of the isolated nucleic acid or rAAV.
[0027] In some aspects, the present disclosure provides a method of modulating gene expression (e.g., expression of SCN1A) in a subject, wherein the isolated nucleic acid or rAAV as described herein is administered to a subject that expresses a target gene. In some embodiments, the expression of the target gene in the subject is abnormal (e.g., increased or decreased) relative to a healthy subject. In some embodiments, the subject is or is suspected of being haploinsufficient with respect to the expression of the target gene relative to a healthy subject.
[0028] In some embodiments, the subject has or is suspected of having a disease or condition caused by haploinsufficient expression of a target gene. For example, in some embodiments, a subject haploinsufficient for SCN1A expression has Dravet syndrome. In some embodiments, the isolated nucleic acid or rAAV is administered to the subject by intravenous injection, intramuscular injection, inhalation, subcutaneous injection, and / or intracranial injection.
[0029] In some aspects, the present disclosure provides a composition comprising the isolated nucleic acid or rAAV as described in the present disclosure. In some embodiments, the composition comprises a pharmaceutically acceptable carrier.
[0030] In some aspects, the present disclosure provides a kit comprising a container containing the isolated nucleic acid or rAAV as described in the present disclosure. In some embodiments, the kit comprises a container containing a pharmaceutically acceptable carrier. In some embodiments, the isolated nucleic acid or rAAV and the pharmaceutically acceptable carrier are contained in the same container. In some embodiments, the container is a syringe.
[0031] In some aspects, the present disclosure provides a host cell comprising a separated nucleic acid or rAAV as described in the present disclosure. In some embodiments, the host cell is a eukaryotic cell. In some embodiments, the host cell is a mammalian cell. In some embodiments, the host cell is a human cell, optionally a neuron, such as a GABAergic neuron. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 Chromatogram sequencing data showing sequence conservation between the human (HEK) and mouse (HEPG2) SCN1A genes (consensus sequence – SEQ ID NO:98; target sequence – SEQ ID NO:99; Hep-SCN1A_R4 sequence (top) – SEQ ID NO:100; Hep-SCN1A_R4 sequence (bottom) – SEQ ID NO:101
[0033] Figure 2 Shows a sequence alignment of the proximal promoter regions of the human (SEQ ID NO:1) and mouse (SEQ ID NO:2) SCN1A genes, with conserved sequences highlighted. Within this conserved sequence is the target region of interest of the zinc finger protein (ZFP) binding region, which is shown in bold (SEQ ID NO:4).
[0034] Figure 3 Is a schematic diagram showing the positions (SEQ ID NO:3) of the binding sites of three overlapping target ZFPs (ZFP-1, ZFP-2, ZFP-3) (SEQ ID NO:5-7) in the proximal promoter region of the SCN1A gene.
[0035] Figures 4A to 4D Shows an alignment of the six recognition helix sequences of the individual zinc fingers (finger 1 to finger 6; F1-F6) in ZFP-1, which will recognize separate three-base regions (DNA triplets shown in red, separated by “·”) within the proximal promoter region of the SCN1A gene (SEQ ID NO:2). Figure 4A The nucleotide sequence (SEQ ID NO:3) to which zinc fingers 1 to 6 (F1-F6) of ZFP-1 will bind is highlighted. Figure 4B Shows the three nucleotide sequences recognized by each of the recognition helices (seven amino acids) of fingers 1 to 6 of ZFP-1 (SEQ ID NO:17-22). Figure 4C Shows the amino acid sequence of ZFP-1, which contains 6 fingers, one per line, with the linkers between the fingers highlighted to designate classical (TGEKP) and non-classical (TGSQKP) linker sequences (SEQ ID NO:65-70). Figure 4DShows the nucleotide sequence of ZFP-1 (F1-F6) (SEQ ID NO: 102-107).
[0036] Figures 5A to 5D Shows the alignment of the six recognition helix sequences of individual zinc fingers (fingers 1 to 6; F1-F6) in ZFP-2, which will recognize separate three-base regions (DNA triplets shown in red, separated by "*") within the proximal promoter region of the SCN1A gene (SEQ ID NO: 3). Figure 5A Highlights the nucleotide sequence (SEQ ID NO: 3) to which zinc fingers 1 to 6 (F1-F6) of ZFP-2 will bind. Figure 5B Shows the first three nucleotides recognized by each recognition helix (seven amino acids) of fingers 1 to 6 of ZFP-2 (SEQ ID NO: 29-34). Figure 5C Shows the amino acid sequence of ZFP-2, which contains 6 fingers, one per line (SEQ ID NO: 69-74), where the linkers between the fingers are highlighted to designate classical (TGEKP) and non-classical (TGSQKP) linker sequences. Figure 5D Shows the nucleotide sequence of ZFP-2 (F1-F6) (SEQ ID NO: 108-113).
[0037] Figures 6A to 6D Shows the alignment of the six recognition helix sequences of individual zinc fingers (fingers 1 to 6; F1-F6) in ZFP-3, which will recognize separate three-base regions (DNA triplets shown in red, separated by "*") within the proximal promoter region of the SCN1A gene (SEQ ID NO: 4). Figure 6A Highlights the nucleotide sequence (SEQ ID NO: 3) to which zinc fingers 1 to 6 (F1-F6) of ZFP-3 will bind. Figure 6B Shows the first three nucleotides recognized by each recognition helix (seven amino acids) of fingers 1 to 6 of ZFP-3 (SEQ ID NO: 41-46). Figure 6C Shows the amino acid sequence of ZFP-3, which contains 6 fingers, one per line (SEQ ID NO: 75-80), where the linkers between the fingers are highlighted to designate classical (TGEKP) and non-classical (TGSQKP) linker sequences. Figure 6D Shows the nucleotide sequence of ZFP-3 (F1-F6) (SEQ ID NO: 114-119).
[0038] Figure 7Shows data demonstrating that the SCN1A-binding ZFPs described in FIGS. 4-6 increase SCN1A gene expression in HEK293T cells as measured by quantitative real-time polymerase chain reaction (qRT-PCR). These expression constructs were delivered to cells by transient transfection of expression plasmids encoding the following transcriptional regulators: Streptococcus pyogenes Cas9 + SCN1A guide RNA (SpCas9+Scn1a); nuclease-dead Cas9 (dCas9); VPR activation domain + SCN1A guide RNA (dCas9_VPR+Scn1a); VPR activation domain + ZFP1 (VPR_ZFP1); VPR activation domain + ZPF2 (VPR_ZFP2); VPR activation domain + ZFP3 (VPR_ZFP3); SpCas9 + ASCL1 guide RNA (SpCas9+Ascl1); three VPR_ZFPs (VPR_ZFP1+VPR_ZFP2+VPR_ZFP3). Expression levels were normalized to the TBP expression level determined by qRT-PCR in each sample.
[0039] Figure 8 Shows data demonstrating that the SCN1A-binding ZFPs and Cas9 + SCN1A guide RNA described in FIGS. 4-6 increase SCN1A gene expression in HEK293T cells as measured by quantitative real-time polymerase chain reaction (qRT-PCR). Detailed Description
[0040] Aspects of the present disclosure relate to methods and compositions for regulating (e.g., increasing) the expression of a target gene in a cell or subject, wherein the target gene is haploinsufficient (i.e., the target gene contains one functional copy). In some embodiments, the target gene is SCN1A.
[0041] In some embodiments, the present disclosure provides fusion proteins comprising a DNA-binding domain (such as a ZFP) and a transcriptional regulator domain. In some embodiments, the present disclosure provides fusion proteins comprising a DNA-binding domain (such as a ZFP) and a transactivation domain (e.g., a VPR domain). In some embodiments, the DNA-binding protein binds to a target gene sequence or a regulatory region of the target gene. In some embodiments, the regulatory region is an enhancer sequence, a promoter sequence, or a repressor sequence. In some embodiments, the promoter sequence can be an internal promoter (e.g., located in an intron of the target gene) or an external promoter (e.g., located upstream of the transcription start site of the target gene). In some embodiments, the DNA-binding domain of the fusion proteins described herein binds to a conserved sequence in the promoter region of the target gene (e.g., SCN1A), such that the transactivation domain increases gene expression.
[0042] In some aspects, the present disclosure relates to methods for increasing the expression of a target gene (e.g., SCN1A) in a cell or a subject. In some embodiments, the target gene contains a mutation that renders the cell or subject haploinsufficient for the target gene. Thus, in some embodiments, the methods and compositions of the present disclosure can be used to treat diseases and disorders associated with haploinsufficiency of the target gene product, such as Dravet syndrome, which is typically caused by haploinsufficiency of the voltage-gated sodium channel α subunit Nav1.1 due to a mutation in one copy of the SCN1A gene.
[0043] Transactivator fusion protein
[0044] Some aspects of the present disclosure relate to fusion proteins comprising a DNA binding domain (DBD) and a transactivator domain. As used herein, a fusion protein comprises two or more linked polypeptides encoded by two or more separate amino acid sequences. As used herein, a chimeric protein is a fusion protein in which two or more linked genes are from different species. Fusion proteins are typically produced recombinantly, wherein the gene encoding the fusion protein is located in a system that supports the expression of the two or more linked genes and the translation of the resulting mRNA into a recombinant protein. In some embodiments, the fusion protein is produced recombinantly in a prokaryotic or eukaryotic cell. Fusion proteins can be constructed in a variety of arrangements. For example, one protein (Protein A) is located upstream of a second protein (Protein B). In other fusion protein configurations, Protein B is located upstream of Protein A. In some embodiments, the nucleic acid sequence encoding the DNA binding domain is located upstream of the nucleic acid sequence encoding the transactivator domain and produces a fusion protein comprising a DBD linked to the transactivator. In some embodiments, the nucleic acid sequence encoding the transactivator domain is located upstream of the nucleic acid sequence encoding the DNA binding domain and produces a fusion protein comprising a transactivator domain linked to the DNA binding domain. In some embodiments, the fusion protein comprises a transactivator domain located upstream of the DNA binding domain. In some embodiments, the fusion protein comprises a DNA binding domain located upstream of the transactivator domain.
[0045] In some embodiments, the fusion proteins described in the present disclosure comprise a DNA binding domain. As used herein, "DNA binding domain (DBD)" refers to an independently folded protein that contains at least one structural motif that recognizes double-stranded or single-stranded DNA (dsDNA or ssDNA). Certain DBDs recognize specific sequences (recognition sequences or motifs), while other types of DBDs have general affinity for DNA. In some embodiments, the fusion proteins described in the present disclosure comprise a sequence-specific DBD. In some embodiments, the DBD recognizes (e.g., specifically binds) a nucleic acid sequence within or near the gene encoding the SCN1A protein (e.g., Nav1.1). Proteins containing DBDs are typically involved in cellular processes such as transcription, replication, repair, and DNA storage. The DBD in a transcription factor recognizes a specific DNA sequence in a promoter region or enhancer element to promote gene expression. Transcription factor DBDs are used as fusion proteins in genetic engineering to regulate the expression of target genes and can be mutated to alter DNA binding specificity or DNA binding affinity, thereby regulating the expression of desired target genes. Examples of DBDs include, but are not limited to, helix-turn-helix motifs, zinc finger motifs (including Cys2-His2 zinc fingers), transcription activator-like effectors (TALEs), winged helix motifs, HMG-boxes, dCas proteins (e.g., dCas9 or dCas12a), homeodomains, and OB-fold domains.
[0046] In some embodiments, the present disclosure relates to zinc finger DBD fusion proteins. As used herein, "zinc finger protein (ZFP)" refers to a protein that contains at least one structural motif characterized by the coordination of one or more zinc ions that stabilize protein folding. Zinc fingers are one of the most diverse structural motifs found in proteins, and up to 3% of human genes encode zinc fingers. Most ZFPs contain multiple zinc fingers that make tandem contacts with target molecules, including DNA, RNA, and the small protein ubiquitin. The "classical" zinc finger motif consists of 2 cysteine amino acids and 2 histidine amino acids (C2H2) and binds DNA in a sequence-specific manner. These ZFPs (including transcription factor IIIA (TFIIIA)) are typically involved in gene expression. Multiple zinc finger motifs in a DNA binding protein bind and wrap around the outside of the DNA double helix. Due to their relatively small size (e.g., approximately 25-40, typically 27-35 amino acids per finger), zinc finger domain fusion proteins are used to create DBDs with new DNA binding specificities. These DBDs can deliver other fusion domains (e.g., transcription activation or repression domains or epigenetic modification domains) to alter the transcriptional regulation of target genes. In some embodiments, the zinc finger protein comprises 2 to 8 fingers, where each finger contains 27 to 40 amino acids (e.g., 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 amino acids).
[0047] In some embodiments, the ZFP comprises 1, 2, 3, 4, 5, 6, 7, or 8 zinc fingers. Each zinc finger may comprise 25-40, 25-30, 30-35, 35-40, or 40-45 amino acids. In some embodiments, the zinc finger comprises 27-35 amino acids. In some embodiments, the zinc finger comprises 27, 28, 29, 30, 31, 32, 33, 34, or 35 amino acids. The zinc finger may specifically recognize or bind to a target sequence that is haploinsufficient in a subject, e.g., a target gene or a regulatory region of a target gene. In some embodiments, the zinc finger binds to a target sequence of the SCN1A gene (e.g., human SCN1A, such as that listed in SEQ ID NO:49). In some embodiments, the zinc finger that binds to the target sequence of the SCN1A gene comprises one or more amino acid sequences of SEQ ID NOs: 63-80 or a combination thereof. In some embodiments, the zinc finger specifically recognizes or binds to a target sequence comprising a trinucleotide sequence.
[0048] In some embodiments, the zinc finger comprises a recognition helix that recognizes or binds to a target sequence (e.g., a target sequence comprising a trinucleotide sequence). In some embodiments, the recognition helix binds to a trinucleotide. In some embodiments, the recognition helix comprises 4-10 amino acids. In some embodiments, the recognition helix comprises 4, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the recognition helix binds to the trinucleotide sequence of the SCN1A gene. In some embodiments, the recognition sequence that binds to the SCN1A gene comprises the amino acid sequence of any one of SEQ ID NOs: 17-22, 29-34, or 41-46. In some embodiments, the recognition sequence that binds to the SCN1A gene is encoded by any one of SEQ ID NOs: 11-16, 23-28, or 35-40. In some embodiments, the zinc finger binds to the same nucleotide sequence as the recognition helix comprising the amino acid sequence of any one of SEQ ID NOs: 17-22, 29-34, or 41-46.
[0049] In some embodiments, the zinc finger comprises a linker sequence at its C-terminus, which can be used to link or connect the zinc finger to another zinc finger. In some embodiments, the linker sequence can be, for example, a classical linker comprising the amino acid sequence of TGEKP (SEQ ID NO:120). In some embodiments, the linker sequence can be, for example, a non-classical linker comprising the amino acid sequence of TGSQKP (SEQ ID NO:121). In some embodiments, the linker sequence can be 2-10 amino acids, e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids.
[0050] In some embodiments, the ZFP that binds to a target gene (e.g., the SCN1A gene) comprises six zinc fingers, each zinc finger recognizing or binding to a different trinucleotide sequence of the target gene (e.g., the SCN1A gene). In some embodiments, the ZFP that binds to the SCN1A gene comprises the amino acid sequence of SEQ ID NO:57. In some embodiments, the ZFP that binds to the SCN1A gene comprises zinc fingers containing the amino acid sequences of SEQ ID NO:63, 64, 65, 66, 67, and / or 68. In some embodiments, the ZFP that binds to the SCN1A gene comprises recognition helices containing the amino acid sequences of SEQ ID NO:17, 18, 19, 20, 21, and / or 22. In some embodiments, the ZFP that binds to the SCN1A gene comprises the amino acid sequence of SEQ ID NO:59. In some embodiments, the ZFP that binds to the SCN1A gene comprises zinc fingers containing the amino acid sequences of SEQ ID NO:69, 70, 71, 72, 73, and / or 74. In some embodiments, the ZFP that binds to the SCN1A gene comprises recognition helices containing the amino acid sequences of SEQ ID NO:29, 30, 31, 32, 33, and / or 34. In some embodiments, the ZFP that binds to the SCN1A gene comprises the amino acid sequence of SEQ ID NO:61. In some embodiments, the ZFP that binds to the SCN1A gene comprises zinc fingers containing the amino acid sequences of SEQ ID NO:75, 76, 77, 78, 79, and / or 80. In some embodiments, the ZFP that binds to the SCN1A gene comprises recognition helices containing the amino acid sequences of SEQ ID NO:41, 42, 43, 44, 45, and / or 46. In some embodiments, the ZFP that binds to the SCN1A gene comprises at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, or at least 99% sequence identity with SEQ ID NO:57, 59, or 61 as shown below.
[0051] SEQ ID NO:57 (Amino acid sequence of ZFP1 protein)
[0052] RPFQCRICMRNFSQRGNLVRHIRTHTGEKPFACDICGKKFALSFNLTRHTKIHTGSQKPFQCRICMRNFSRSDNLTRHIRTHTGEKPFACDICGKKFADRSHLARHTKIHTGSQKPFQCRICMRNFSQKAHLTAHIRTHTGEKPFACDICGRKFARSDNLTRHTKIHLRQKD
[0053] SEQ ID NO:59 (Amino acid sequence of ZFP2 protein)
[0054] RPFQCRICMRNFSRSSNLTRHIRTHTGEKPFACDICGKKFADKRTLIRHTKIHTGSQKPFQCRICMRNFSQRGNLVRHIRTHTGEKPFACDICGKKFALSFNLTRHTKIHTGSQKPFQCRICMRNFSRSDNLTRHIRTHTGEKPFACDICGRKFADRSHLARHTKIHLRQKD
[0055] SEQ ID NO:61 (Amino acid sequence of ZFP3 protein)
[0056] RPFQCRICMRNFSDRSALARHIRTHTGEKPFACDICGKKFARSDNLTRHTKIHTGSQKPFQCRICMRNFSQSGDLTRHIRTHTGEKPFACDICGKKFAVRQTLKQHTKIHTGSQKPFQCRICMRNFSAAGNLTRHIRTHTGEKPFACDICGRKFARSDNLTRHTKIHLRQKD
[0057] In some embodiments, the DBD is a transcription activator-like effector protein (TALE). A TALE can specifically recognize or bind to a target sequence, e.g., a target gene or a regulatory region of a target gene. In some embodiments, the subject is haploinsufficient for the target gene. In some embodiments, the TALE binds to a target sequence of the SCN1A gene (e.g., human SCN1A as provided in SEQ ID NO: 49). TALE proteins are secreted by bacteria and bind to promoter sequences in the host plant to activate the expression of plant genes that contribute to bacterial infection. Typically, TALE proteins are engineered to bind to new DNA sequences because the target sequence is recognized by a central repeat domain consisting of a variable number of approximately 30 - 35 amino acid repeats, where each repeat recognizes a single base pair within the target sequence. An array of these repeats is generally required for DNA sequence recognition.
[0058] In some embodiments, the DBD is a homeodomain. The homeodomain can specifically recognize or bind to a target sequence, e.g., a target gene or a regulatory region of a target gene. In some embodiments, the subject is haploinsufficient for the target gene. In some embodiments, the homeodomain binds to a target sequence of the SCN1A gene (e.g., human SCN1A as provided in SEQ ID NO:49). The homeodomain is a protein containing three α -helices and an N-terminal arm that is responsible for recognizing the target sequence. Homeodomains typically recognize small DNA sequences (about 4 to 8 base pairs), however these domains can be tandemly fused with other DNA-binding domains (other homeodomains or zinc finger proteins) to recognize longer extended sequences (12 to 24 base pairs). Thus, the homeodomain can be a component of a DBD that recognizes unique sequences in the human genome.
[0059] In some embodiments, at least one DNA binding domain is a catalytically inactivated CRISPR-associated protein (Cas protein). A catalytically inactivated Cas protein (also referred to as dCas or “dead Cas protein”) is a Cas protein that has been modified or mutated such that its nuclease activity (e.g., endonuclease activity) is reduced or lacks all nuclease activity (e.g., endonuclease activity). In some embodiments, the catalytically inactivated Cas protein is a dCas9 or dCas12 protein. In some embodiments, the DBD is a dCas protein (also referred to as ‘dead Cas’), such as dCas9 or dCas12a. A dCas protein is a mutant variant of a CRISPR-associated protein (Cas, e.g., Cas9 or Cas12a) that has been mutated such that it is catalytically inactivated (i.e., unable to perform nucleotide cleavage). The dCas can specifically recognize or bind to a target sequence, e.g., a target gene or a regulatory region of a target gene. A complex comprising a dCas protein and a guide nucleic acid (e.g., gRNA) can target and / or bind to a specific nucleotide sequence or gene that is complementary to the guide nucleic acid. In some embodiments, the subject is haploinsufficient for the target gene. In some embodiments, the dCas binds to a target sequence of the SCN1A gene (e.g., human SCN1A as provided in SEQ ID NO:49). However, the dCas protein retains its ability to recognize and bind to the target DNA sequence when bound to a guide nucleic acid (e.g., guide RNA, gRNA, or sgRNA) that is complementary or partially complementary to the target DNA sequence. In some embodiments, the guide nucleic acid for targeting the dCas (e.g., dCas9) protein to SCN1A comprises a spacer sequence having any one of SEQ ID NO:85, 86, 89, 90, 93, or 94. In some embodiments, the guide nucleic acid for targeting the dCas (e.g., dCas9) protein to SCN1A comprises a spacer sequence having at least 15 (e.g., at least 16, 17, 18, 19, or 20) contiguous nucleotides of any one of SEQ ID NO:85, 86, 89, 90, 93, or 94. In some embodiments, the guide nucleic acid for targeting the dCas (e.g., dCas9) protein to SCN1A comprises any one of SEQ ID NO:83, 84, 87, 88, 91, or 92. In some embodiments, the guide nucleic acid for targeting the dCas (e.g., dCas9) protein to SCN1A comprises or consists of any one of SEQ ID NO:83-94. Thus, the dCas endonuclease can be a component of a DBD that recognizes a unique sequence in the human genome. In some embodiments, the fusion protein comprises a dCas9 protein and a transactivation domain (e.g., VPR domain).
[0060] In some aspects, the present disclosure relates to binding to a gene encoding a voltage-gated sodium channel (e.g., Nav The DNA binding domain of the gene of 1.1). In some embodiments, the gene encoding the voltage-gated sodium channel is the SCN1A gene and comprises the sequence listed in SEQ ID NO: 49. In some embodiments, the DNA binding domain binds to the untranslated region of the target gene, such as the 3'-untranslated region (3'UTR) or the 5'-untranslated region (5'UTR). In some embodiments, the untranslated region comprises regulatory sequences, such as enhancer, promoter, intron or repressor sequences. In some embodiments, the DNA binding domain is a zinc finger protein that comprises the sequence listed in SEQ ID NOs: 57-62. In some embodiments, the DNA binding domain binds to the nucleic acid sequence listed in any of SEQ ID NOs: 5-7.
[0061] The number of DNA binding domains encoded by the transgene can vary. In some embodiments, the transgene encodes one DNA binding domain. In some embodiments, the transgene encodes 2 DNA binding domains. In some embodiments, the transgene encodes 3 DNA binding domains. In some embodiments, the transgene encodes 4 DNA binding domains. In some embodiments, the transgene encodes 5 DNA binding domains. In some embodiments, the transgene encodes 6 DNA binding domains. In some embodiments, the transgene encodes 7 DNA binding domains. In some embodiments, the transgene encodes 8 DNA binding domains. In some embodiments, the transgene encodes 9 DNA binding domains. In some embodiments, the transgene encodes 10 DNA binding domains. In some embodiments, the transgene encodes more than 10 (e.g., 20, 30, 50, 100, etc.) DNA binding domains. The DNA binding domains can be the same DNA binding domain (e.g., multiple copies of the same DBD), different DNA binding domains (e.g., each DBD binds a unique sequence), or a combination thereof.
[0062] In some aspects, the present disclosure relates to fusion proteins comprising a transactivation domain. As used herein, a "transactivation domain" refers to a scaffold domain in a transcription factor that contains binding sites for other proteins (such as transcriptional co-regulators) that regulate gene expression. In some embodiments, the transactivation domain (also referred to as a transcriptional activation domain) acts in conjunction with a DBD to directly activate transcription from a promoter or enhancer by contacting a transcription factor or indirectly through a co-activator protein. Transactivation domains (TADs) are typically named based on their amino acid composition, where the amino acids are required for activity or are the most abundant in the TAD. TADs are used as fusion proteins in genetic engineering to regulate the expression of target genes and can be mutated to alter the level of transcriptional activation and thus the expression of the target gene. Examples of transactivation domains include, but are not limited to, GAL4, HAP1, VP16, P65, RTA, and GCN4.
[0063] In some embodiments, the transactivation domain comprises a VP64 domain. VP64 is an acidic TAD consisting of four tandem copies of the VP16 protein naturally expressed by herpes simplex virus. When fused to a DBD that binds at or near a gene promoter, VP64 acts as a strong transcriptional activator and can thus be used to regulate the expression of a target gene (e.g., SCN1A). The VP64 domain typically consists of a tetrameric repeat of the minimal activation domain of the herpes simplex protein VP16. In some embodiments, the VP64 domain comprises four repeats of amino acid residues 437 - 448 in VP16. In some embodiments, the VP16 protein is encoded by the human herpesvirus 2 UL48 gene, which contains the sequence listed in NCBI reference sequence accession number: NC_001798.2. In some embodiments, the VP16 gene comprises a nucleotide sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequence encoded by the nucleic acid sequence listed in NCBI reference sequence accession number: YP_009137200.1. In some embodiments, the VP16 protein comprises an amino acid sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequence listed in NCBI reference sequence accession number Q69113-1. In some embodiments, the VP16 gene comprises a nucleotide sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequence encoded by the nucleic acid sequence listed in SEQ ID NO: 51. In some embodiments, the VP16 protein comprises an amino acid sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequence listed in SEQ ID NO:52.
[0064] In some embodiments, the transactivation domain comprises a P65 activation domain. P65 is a subunit of the NF-κβ transcription factor, and its C-terminus contains two adjacent acidic TADs. When fused to a DBD that binds to or near a gene promoter, the p65 protein acts as a strong transcriptional activator and can thus be used to regulate the expression of target genes, for example as described by Urlinger et al. in “The p65 domain from NF-kappaB is an efficient human activator in the tetracycline-regulatable gene expression system,” Gene, 2000. In some embodiments, the p65 protein is encoded by the human RELA gene, which comprises the sequences listed in NCBI Reference Sequence accession numbers: NM_001145138.1, NM_001243984.1, NM_001243985.1, or NM_021975.3. In some embodiments, the RELA gene comprises a nucleotide sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequence encoded by the nucleic acid sequence listed in any of NCBI Reference Sequence accession numbers: NM_001145138.1, NM_001243984.1, NM_001243985.1, or NM_021975.3. In some embodiments, the p65 protein comprises an amino acid sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequences listed in NP_001138610.1, NP_001230913.1, NP_001230914.1, and NP_068110.3. In some embodiments, the RELA gene comprises a nucleotide sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequence encoded by the nucleic acid sequence listed in SEQ ID NO: 53. In some embodiments, the p65 protein comprises an amino acid sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequence listed in SEQ ID NO: 54.
[0065] In some embodiments, the transactivation domain comprises an RTA domain. RTA is a hydrophobic TAD derived from Epstein Barr virus, which is an efficient transactivation domain that binds to enhancer regions to promote the expression of several viral genes. When fused to a DBD that binds at or near a gene promoter, the RTA protein acts as a strong transcriptional activator and can thus be used to regulate the expression of target genes, such as described by Miyazawa et al., “IL-10 promoter transactivation by the viral K-RTA protein involves the host-cell transcription factors, specificity proteins 1 and 3,” Journal of Biological Chemistry, 2018. In some embodiments, the RTA protein is encoded by the Epstein Barr virus BRLF1 gene, which comprises the sequence listed in NCBI reference sequence accession number: YP_041674.1. In some embodiments, the BRLF1 gene comprises a nucleotide sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequence encoded by any of the nucleic acid sequences listed in NCBI reference Seq ID No: YP_041674.1. In some embodiments, the RTA protein comprises an amino acid sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequence listed in YP_041674.1. In some embodiments, the BRLF1 gene comprises a nucleotide sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequence encoded by the nucleic acid sequence listed in SEQ ID NO: 55. In some embodiments, the RTA protein comprises an amino acid sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the amino acid sequence listed in SEQ ID NO: 56.
[0066] The present disclosure is in part based on fusion proteins comprising a hybrid trans-activation domain. As used herein, a "hybrid trans-activation domain" refers to a fusion protein comprising more than one transcriptional activator protein or a portion thereof (e.g., 2, 3, 4, 5 or more transcriptional activator proteins or portions thereof). Hybrid trans-activation domains are used in genetic engineering to increase the expression of target genes. In some embodiments of the present disclosure, a ternary hybrid trans-activation domain comprising the nucleotide sequence of VP64-P65-RTA (VPR) (as described in Chavez et al., "Highly efficient Cas9-mediated transcriptional programming", Nat Methods, 2015, (SEQ ID NO: 47)) is used to increase the expression of a target gene (e.g., SCN1A).
[0067] In some embodiments, the fusion proteins described herein may comprise a DBD (e.g., a ZFP) and a transcriptional repressor protein. In some aspects, the present disclosure relates to fusion proteins comprising a transcriptional repressor domain. As used herein, a "transcriptional repressor" protein generally refers to a polypeptide that downregulates the expression of a target gene. Examples of transcriptional repressors include, but are not limited to, KRAB, SMRT / TRAC-2, and NCoR / RIP-13. In some embodiments, such transcriptional repressor fusion proteins can be used to reduce the expression level of a target gene (e.g., a gene that is overexpressed in a gain-of-function disease).
[0068] Isolated nucleic acid
[0069] An isolated nucleic acid sequence refers to a DNA or RNA sequence. In some embodiments, the proteins and nucleic acids of the present disclosure are isolated. As used herein, the term "isolated" means produced artificially. As used herein with respect to nucleic acids, the term "isolated" means: (i) amplified in vitro, such as by polymerase chain reaction (PCR); (ii) produced recombinantly by cloning; (iii) purified, such as by lysis and gel separation; or (iv) synthesized, such as by chemical synthesis. An isolated nucleic acid is a nucleic acid that is readily manipulable by recombinant DNA techniques well known in the art. Thus, a nucleotide sequence contained in a vector with known 5' and 3' restriction sites or published polymerase chain reaction (PCR) primer sequences is considered isolated, but a nucleic acid sequence that exists in its natural state in a natural host is not isolated. An isolated nucleic acid may be substantially purified, but it is not required to be. For example, an isolated nucleic acid in a cloning or expression vector is not pure because it may represent only a small percentage of the material in the cell in which it resides. However, as used herein, such a nucleic acid is isolated because it is readily manipulable by standard techniques known to those of ordinary skill in the art. As used herein with respect to a protein or peptide, the term "isolated" refers to a protein or peptide that has been separated from its natural environment or produced artificially (e.g., by chemical synthesis, by recombinant DNA techniques, etc.).
[0070] In some aspects, the present disclosure relates to isolated nucleic acids (e.g., expression constructs, such as rAAV vectors) configured to express one or more ZFP transactivation domain fusion proteins. In some embodiments, the fusion protein comprises 1 to 10 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) DBDs and / or 1 to 10 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) transactivation factor domains. In some embodiments, the fusion protein comprises more than 10 DBSs and / or more than 10 transactivation factor domains.
[0071] In some aspects of the present disclosure, the DNA binding domain is indirectly fused to the transcriptional regulator domain via a linker. As used herein, a "linker" is generally a stretch of polypeptide that structurally connects two different polypeptides within a single transgene. In some embodiments, the linker is flexible to allow movement of the different polypeptides. In some embodiments, the flexible linker comprises glycine residues. In some embodiments, the flexible linker comprises a mixture of glycine and serine residues. In some embodiments, the linker is cleavable to allow separation of the polypeptides. In some embodiments, the cleavable linker is cleaved by a protease. In some embodiments, the protease is trypsin or factor X.
[0072] In some embodiments, the linker comprises 5 to 30 amino acids (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids). In some embodiments, the linker comprises 3 to 30 amino acids (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids). In some embodiments, the linker comprises 3 to 20 amino acids (e.g., 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids).
[0073] The present disclosure is in part based on fusion proteins that are engineered to increase the expression of genes encoding voltage-gated sodium channel subunit proteins (also referred to as SCN proteins) (e.g., SCN1A). As used herein, an “SCN protein” refers to a sodium channel protein that mediates the voltage-dependent sodium permeability of excitable membranes, thereby allowing sodium ions to cross the membrane. Examples of SCN proteins in humans include, but are not limited to, SCN1A, SCN3A, SCN5A, SCN10A, and SCN11A. In some embodiments, the SCN protein is SCN1A (also referred to as Nav1.1), which encodes a type 1 α1 ion channel subunit. In some embodiments, the SCN protein is an SCN1B protein, which encodes a type 1 β1 ion channel subunit or an SCN1C protein. In some embodiments, the SCN protein is a combination of SCN1A, SCN1B, and / or SCN1C proteins. As disclosed herein, the SCN protein can be a part or fragment of an SCN protein. In some embodiments, the SCN protein as disclosed herein is a variant of an SCN protein, such as a point mutant or a truncated mutant.
[0074] In humans, SCN1A is encoded by the SCN1A gene (Gene ID: 6323, human), which is conserved in chimpanzee, rhesus macaque, dog, cow, mouse, rat, and chicken. The SCN1A gene in humans is mainly expressed in the brain, lung, and testis. In some embodiments, the SCN1A protein comprises five structural repeats (I, II, III, IV, Q).
[0075] In some embodiments, the SCN1A protein is encoded by the human SCN1A gene, which comprises the sequences listed in NCBI Reference Seq ID No: NM_001165963.2, NM_00165964.2, NM_001202435.2, NM_001353948.1, NM_001353949.1, NM_001353950.1, NM_00135395.1, NM_001353952.1, NM_001353954.1, NM_00353955.1, NM_001353957.1, NM_001353958.1, NM_001353960.1, NM_001353961.1 or NM_006920.5. In some embodiments, the SCN1A protein is encoded by the mouse SCN1A gene, which comprises the sequences listed in NCBI Reference Seq ID No: NM_001313997.1 or NM_018733.2. In some embodiments, the SCN1A protein comprises an amino acid sequence that is 99%, 95%, 90%, 80%, 70%, 60% or 50% identical to the amino acid sequence encoded by the nucleic acid sequence listed in NCBI Reference Seq ID No: NG_011906.1, NM_001313997.1 or NM_018733.2. In some embodiments, the SCN1A gene comprises an amino acid sequence that is 99%, 95%, 90%, 80%, 70%, 60% or 50% identical to the sequence listed in SEQ ID NO: 50. In some embodiments, the human SCN1A protein comprises the sequences listed in NCBI Reference Seq ID No: NP_001159435.1, NP_0011159436.1, NP_001189364.1, NP_001340877.1, NP_001340878.1, NP_001340879.1, NP_001340880.1, NP_001340881.1, NP_001340883.1, NP_001340884.1, NP_001340886.1, NP_001340887.1, NP_001340889.1, NP_001340890.1, NP_00851.3. In some embodiments, the SCN1A protein comprises an amino acid sequence that is 99%, 95%, 90%, 80%, 70%, 60% or 50% identical to the amino acid sequence encoded by the nucleic acid sequence listed in NCBI Reference Seq ID No: NG_011906.1, NM_001313997.1 or NM_018733.2.In some embodiments, the murine SCN1A protein comprises the sequence listed in NCBI Reference Seq ID No: NP_001300926.1 or NP_061203.2. In some embodiments, the human SCN1A protein comprises an amino acid sequence that is 99%, 95%, 90%, 80%, 70%, 60%, or 50% identical to the nucleic acid sequence listed in SEQ ID NO: 49.
[0076] The isolated nucleic acids of the present disclosure can be recombinant adeno-associated virus (AAV) vectors (rAAV vectors). In some embodiments, the isolated nucleic acid as described in the present disclosure comprises a region (e.g., a first region) containing a first adeno-associated virus (AAV) inverted terminal repeat (ITR) or a variant thereof. The isolated nucleic acid (e.g., the recombinant AAV vector) can be packaged into a capsid protein and administered to a subject and / or delivered to a selected target cell. A "recombinant AAV (rAAV) vector" generally consists of at least a transgene and its regulatory sequences, as well as 5' and 3' AAV inverted terminal repeats (ITRs). As described elsewhere in the present disclosure, the transgene can comprise a region encoding, for example, a protein and / or an expression control sequence (e.g., a poly-A tail).
[0077] Typically, the length of the ITR sequence is about 145 bp. Preferably, substantially the entire sequence encoding the ITR is used in the molecule, although some minor modifications to these sequences are allowed. The ability to modify these ITR sequences is within the skill in the art. (See, for example, texts such as Sambrook et al., "Molecular Cloning. A Laboratory Manual", 2nd ed., Cold Spring Harbor Laboratory, New York (1989); and K. Fisher et al., J Virol., 70:520 - 532 (1996)). An example of such a molecule employed in the present disclosure is a "cis-acting" plasmid containing a transgene, wherein the selected transgene sequence and associated regulatory elements are flanked by 5' and 3' AAV ITR sequences. The AAV ITR sequences can be obtained from any known AAV, including currently identified mammalian AAV types. In some embodiments, the isolated nucleic acid further comprises a region (e.g., a second region, a third region, a fourth region, etc.) containing a second AAV ITR.
[0078] In addition to the major elements of the recombinant AAV vector identified above, the vector also includes conventional control elements operably linked to the elements of the transgene in such a way as to permit its transcription, translation, and / or expression in cells transfected with the vector or infected with virus produced by the present disclosure. As used herein, "operably linked" sequences include expression control sequences adjacent to the gene of interest and expression control sequences that act in trans or at a distance to control the expression of the gene of interest. Expression control sequences include appropriate transcriptional start, stop, promoter, and enhancer sequences; efficient RNA processing signals, such as splicing and polyadenylation (polyA) signals; sequences that stabilize cytoplasmic mRNA; sequences that enhance translation efficiency (e.g., Kozak consensus sequences); sequences that enhance protein stability; and, when desired, sequences that enhance the secretion of the encoded product. Many expression control sequences, including native, constitutive, inducible, and / or tissue-specific promoters, are known in the art and can be utilized.
[0079] As used herein, nucleic acid sequences (e.g., coding sequences) and regulatory sequences are said to be operably linked when they are covalently linked in such a way as to place the expression or transcription of the nucleic acid sequence under the influence or control of the regulatory sequence. If it is desired to translate a nucleic acid sequence into a functional protein, two DNA sequences are said to be operably linked if induction of a promoter in the 5' regulatory sequence results in transcription of the coding sequence and if the nature of the linkage between the two DNA sequences does not (1) result in the introduction of a frameshift mutation, (2) interfere with the ability of the promoter region to direct transcription of the coding sequence, or (3) interfere with the ability of the corresponding RNA transcript to be translated into protein. Thus, a promoter region will be operably linked to a nucleic acid sequence if the promoter region can affect the transcription of that DNA sequence such that the resulting transcript can be translated into the desired protein or polypeptide. Similarly, two or more coding regions are operably linked when they are linked in such a way that transcription from a common promoter results in the expression of two or more proteins that have been translated in frame. In some embodiments, the operably linked coding sequences produce a fusion protein.
[0080] The region containing the transgene (e.g., containing a fusion protein, etc.) can be located at any suitable position of the isolated nucleic acid that will be capable of expressing the fusion protein.
[0081] It should be understood that in the case where the transgene encodes more than one polypeptide, each polypeptide can be located at any suitable position within the transgene. For example, the nucleic acid encoding the first polypeptide can be located within an intron of the transgene, and the nucleic acid sequence encoding the second polypeptide can be located in another untranslated region (e.g., between the last codon of the protein-coding sequence and the first base of the transgene poly-A signal).
[0082] "Promoter" means a DNA sequence recognized by the synthetic machinery of a cell that is required to initiate specific transcription of a gene or introduced synthetic machinery. The phrases "operably linked", "operably positioned", "controlled" or "under transcriptional control" mean that the promoter is in the correct position and orientation relative to the nucleic acid to control the initiation of RNA polymerase and the expression of the gene.
[0083] For nucleic acids encoding proteins, a polyadenylation sequence is typically inserted after the transgene sequence and before the 3' AAV ITR sequence. The rAAV constructs useful in the present disclosure may also contain an intron, which is desirably located between the promoter / enhancer sequence and the transgene. One possible intron sequence is derived from SV-40 and is referred to as the SV-40T intron sequence. Another vector element that can be used is an internal ribosome entry site (IRES). The IRES sequence is used to produce more than one polypeptide from a single gene transcript. The IRES sequence will be used to produce proteins containing more than one polypeptide chain. The selection of these and other common vector elements is routine, and many such sequences are available [see, for example, Sambrook et al. and the references cited therein at pages 3.18 - 3.26 and 16.17 - 16.27, as well as Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, New York, 1989]. In some embodiments, the foot-and-mouth disease virus 2A sequence is included in the polyprotein; this is a small peptide (about 18 amino acids in length) that has been shown to mediate cleavage of the polyprotein (Ryan, M D et al., EMBO, 1994; 4:928 - 933; Mattion, N M et al., J Virology, November 1996; pp. 8124 - 8127; Furler, S et al., Gene Therapy, 2001; 8:864 - 873; and Halpin, C et al., The Plant Journal, 1999; 4:453 - 459). The cleavage activity of the 2A sequence has previously been demonstrated in artificial systems including plasmids and gene therapy vectors (AAV and retroviruses) (Ryan, M D et al., EMBO, 1994; 4:928 - 933; Mattion, N M et al., J Virology, November 1996; pp. 8124 - 8127; Furler, S et al., Gene Therapy, 2001; 8:864 - 873; and Halpin, C et al., The Plant Journal, 1999; 4:453 - 459; de Felipe, P et al., Gene Therapy, 1999; 6:198 - 208; de Felipe, P et al., Human Gene Therapy, 2000; 11:1921 - 1931; and Klump, H et al., Gene Therapy, 2001; 8:811 - 817).
[0084] Examples of constitutive promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer) [see, e.g., Boshart et al., Cell, 41:521-530 (1985)], the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerate kinase (PGK) promoter, and the EF1α promoter [Invitrogen]. In some embodiments, the promoter is the P2 promoter. In some embodiments, the promoter is the chicken β-actin (CBA) promoter. In some embodiments, the promoter is two CBA promoters. In some embodiments, the promoter is two CBA promoters isolated by the CMV enhancer. In some embodiments, the promoter is the CAG promoter.
[0085] Inducible promoters permit the regulation of gene expression and can be regulated by an exogenously supplied compound, an environmental factor such as temperature, or the presence of a specific physiological state (e.g., acute phase, a specific differentiation state of a cell, or only in replicating cells). Inducible promoters and inducible systems are available from a variety of commercial sources, including but not limited to Invitrogen, Clontech, and Ariad. Many other systems have been described and can be readily selected by those skilled in the art. Examples of inducible promoters regulated by exogenously supplied promoters include the zinc-inducible sheep metallothionein (MT) promoter, the dexamethasone (Dex)-inducible mouse mammary tumor virus (MMTV) promoter, the T7 polymerase promoter system (WO98 / 10088); the ecdysone insect promoter (No et al., Proc. Natl. Acad. Sci. USA, 93:3346-3351 (1996)), the tetracycline repression system (Gossen et al., Proc. Natl. Acad. Sci. USA, 89:5547-5551 (1992)), the tetracycline-inducible system (Gossen et al., Science, 268:1766-1769 (1995), see also Harvey et al., Curr. Opin. Chem. Biol., 2:512-518 (1998)), the RU486-inducible system (Wang et al., Nat. Biotech., 15:239-243 (1997) and Wang et al., Gene Ther., 4:432-441 (1997)), and the rapamycin-inducible system (Magari et al., J. Clin. Invest., 100:2865-2872 (1997)). Other types of inducible promoters that may be useful in this context are those that are regulated by a specific physiological state (e.g., temperature, acute phase, a specific differentiation state of a cell, or only in replicating cells).
[0086] In another embodiment, the native promoter of the transgene will be used. The native promoter may be preferred when it is desired that the expression of the transgene mimic native expression. The native promoter can be used when the expression of the transgene must be temporal or developmental, or in a tissue-specific manner, or in response to a specific transcriptional stimulator. In another embodiment, other native expression control elements such as enhancer elements, polyadenylation sites, or Kozak consensus sequences can also be used to mimic native expression.
[0087] In some embodiments, the regulatory sequence confers the ability of tissue-specific gene expression. In some cases, tissue-specific regulatory sequences bind to tissue-specific transcription factors that induce transcription in a tissue-specific manner. Such tissue-specific regulatory sequences (e.g., promoters, enhancers, etc.) are well known in the art. Exemplary tissue-specific regulatory sequences include, but are not limited to, the following tissue-specific promoters: liver-specific thyroxine-binding globulin (TBG) promoter, insulin promoter, glucagon promoter, somatostatin promoter, pancreatic polypeptide (PPY) promoter, synaptophysin-1 (Syn) promoter, creatine kinase (MCK) promoter, mammalian desmin (DES) promoter, α-myosin heavy chain (α-MHC) promoter, or cardiac troponin T (cTnT) promoter. Other exemplary promoters include the β-actin promoter, hepatitis B virus core promoter, Sandig et al., Gene Ther., 3:1002-9 (1996); alpha-fetoprotein (AFP) promoter, Arbuthnot et al., Hum. Gene Ther., 7:1503-14 (1996)), osteocalcin promoter of bone (Stein et al., Mol. Biol. Rep., 24:185-96 (1997)); bone sialoprotein promoter (Chen et al., J. Bone Miner. Res., 11∶654-64 (1996)), CD2 promoter (Hansal et al., J. Immunol., 161∶1063-8 (1998); immunoglobulin heavy chain promoter; T cell receptor alpha chain promoter, neurons such as neuron-specific enolase (NSE) promoter (Andersen et al., Cell. Mol. Neurobiol., 13:503-15 (1993)), neurofilament light chain gene promoter (Piccioli et al., Proc. Natl. Acad. Sci. USA, 88∶5611-5 (1991)) and neuron-specific vgf gene promoter (Piccioli et al., Neuron, 15∶373-84 (1995)) and other promoters that will be obvious to those skilled in the art.
[0088] In some embodiments, the transgene encoding a fusion protein comprising a DBD and a transactivator is operably linked to a promoter. In some embodiments, the promoter is a constitutive promoter. In some embodiments, the promoter is an inducible promoter. In some embodiments, the promoter is a tissue-specific promoter. In some embodiments, the promoter is specific for neural tissue. In some embodiments, the promoter is the SST or NPY promoter.
[0089] Aspects of the present disclosure relate to an isolated nucleic acid comprising more than one promoter (e.g., 2, 3, 4, 5 or more promoters). For example, in the case of a construct having a transgene with a first region encoding a protein and a second region encoding a protein, it may be desirable to use a first promoter sequence (e.g., a first promoter sequence operably linked to the protein-coding region) to drive the expression of the first protein-coding region, and a second promoter sequence (e.g., a second promoter sequence operably linked to the second protein-coding region) to drive the expression of the second protein-coding region. Generally, the first promoter sequence and the second promoter sequence may be the same promoter sequence or different promoter sequences. In some embodiments, the first promoter sequence (e.g., the promoter driving the expression of the protein-coding region) is an RNA polymerase III (pol III) promoter sequence. Non-limiting examples of pol III promoter sequences include U6 and H1 promoter sequences. In some embodiments, the second promoter sequence (e.g., the promoter sequence driving the expression of the second protein) is an RNA polymerase II (pol II) promoter sequence. Non-limiting examples of pol II promoter sequences include T7, T3, SP6, RSV, and cytomegalovirus promoter sequences. In some embodiments, the pol III promoter sequence drives the expression of the first protein-coding region. In some embodiments, the pol II promoter sequence drives the expression of the second protein-coding region.
[0090] Recombinant adeno-associated virus (rAAV)
[0091] In some aspects, the present disclosure provides isolated adeno-associated virus (AAV). As used herein with respect to AAV, the term "isolated" refers to AAV that is produced or obtained artificially. Isolated AAV can be produced using recombinant methods. Such AAV is referred to herein as "recombinant AAV". Recombinant AAV (rAAV) preferably has tissue-specific targeting ability such that the nuclease and / or transgene of rAAV will be specifically delivered to one or more predetermined tissues. The AAV capsid is an important element in determining these tissue-specific targeting abilities. Thus, rAAV having a capsid suitable for the targeted tissue can be selected.
[0092] Methods for obtaining recombinant AAVs with desired capsid proteins are well known in the art. (See, e.g., US 2003 / 0138772), the content of which is incorporated herein by reference in its entirety. Generally, the methods involve culturing host cells containing a nucleic acid sequence encoding an AAV capsid protein; a functional rep gene; a recombinant AAV vector consisting of AAV inverted terminal repeats (ITRs) and a transgene; and sufficient helper functions to allow packaging of the recombinant AAV vector into the AAV capsid protein. In some embodiments, the capsid protein is a structural protein encoded by the cap gene of AAV. AAV contains three capsid proteins, viral proteins 1 to 3 (named VP1, VP2, and VP3), all of which are transcribed from a single cap gene by alternative splicing. In some embodiments, the molecular weights of VP1, VP2, and VP3 are approximately 87 kDa, approximately 72 kDa, and approximately 62 kDa, respectively. In some embodiments, upon translation, the capsid proteins form a spherical 60-mer protein shell around the viral genome. In some embodiments, the function of the capsid protein is to protect the viral genome, deliver the genome, and interact with the host. In some aspects, the capsid protein delivers the viral genome to the host in a tissue-specific manner.
[0093] In some embodiments, the AAV capsid protein has an AAV serotype selected from the group consisting of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAV9, AAV10, AAVrh10, and AAV.PHP.B. In some embodiments, the AAV capsid protein has a serotype derived from a non-human primate, such as the AAVrh8 serotype. In some embodiments, the AAV capsid protein has a serotype derived for broad and efficient CNS transduction, such as AAV.PHP.B. In some embodiments, the capsid protein has the AAV serotype 9.
[0094] Components that are cultured in a host cell to package an rAAV vector in an AAV capsid can be provided to the host cell in trans. Alternatively, any one or more of the required components (e.g., recombinant AAV vector, rep sequences, cap sequences, and / or helper functions) can be provided by a stable host cell that has been engineered using methods known to those of skill in the art to contain one or more of the required components. Most desirably, such a stable host cell will contain one or more of the required components under the control of an inducible promoter. However, one or more of the required components may be under the control of a constitutive promoter. Examples of suitable inducible and constitutive promoters are provided herein when discussing regulatory elements suitable for use with a transgene. In another alternative, a selected stable host cell can contain one or more selected components under the control of a constitutive promoter and one or more other selected components under the control of one or more inducible promoters. For example, a stable host cell can be generated that is derived from 293 cells (which contain E1 helper function under the control of a constitutive promoter), but which contains rep and / or cap proteins under the control of an inducible promoter. Those of skill in the art can also generate other stable host cells.
[0095] In some embodiments, the present disclosure relates to a host cell containing a nucleic acid comprising a coding sequence encoding a transgene (e.g., a DNA binding domain fused to a transcriptional regulator domain). In some embodiments, the host cell is a mammalian cell, yeast cell, bacterial cell, insect cell, plant cell, or fungal cell.
[0096] The recombinant AAV vector, rep sequences, cap sequences, and helper functions required to produce the rAAV of the present disclosure can be delivered to the packaging host cell using any suitable genetic element (vector). The selected genetic element can be delivered by any suitable method, including those described herein. The methods for constructing any embodiment of the present disclosure are known to those skilled in nucleic acid manipulation and include genetic engineering, recombination engineering, and synthetic techniques. See, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, N.Y. Similarly, methods for generating rAAV viral particles are well known, and the selection of a suitable method is not a limitation of the present disclosure. See, e.g., K. Fisher et al., J. Virol., 70:520-532 (1993) and U.S. Patent No. 5,478,745.
[0097] In some embodiments, recombinant AAV can be produced using a triple transfection method (described in detail in U.S. Patent No. 6,001,650). Generally, recombinant AAV is produced by transfecting a host cell with an AAV vector to be packaged into AAV particles, an AAV helper function vector, and a helper function vector (containing a transgene flanked by ITR elements). The AAV helper function vector encodes "AAV helper function" sequences (e.g., rep and cap), which act in trans for productive AAV replication and packaging. Preferably, the AAV helper function vector supports efficient AAV vector production without generating any detectable wild-type AAV virions (e.g., AAV virions containing functional rep and cap genes). Non-limiting examples of vectors suitable for use with the present disclosure include pHLP19 described in U.S. Patent No. 6,001,650 and pRep6cap6 vector described in U.S. Patent No. 6,156,303, the entire contents of both are incorporated herein by reference. The helper function vector encodes nucleotide sequences of non-AAV-derived viruses and / or cellular functions (e.g., "helper functions") on which AAV replication depends. Helper functions include those functions required for AAV replication, including but not limited to those involved in the activation of AAV gene transcription, stage-specific AAV mRNA splicing, AAV DNA replication, synthesis of cap expression products, and AAV capsid assembly. Virus-based helper functions can be derived from any known helper virus, such as adenovirus, herpesvirus (except herpes simplex virus type 1), and vaccinia virus.
[0098] In some aspects, the present disclosure provides transfected host cells. The term "transfection" is used to refer to the uptake of exogenous DNA by a cell, and when the exogenous DNA has been introduced into the cell membrane, the cell has been "transfected". Many transfection techniques are well known in the art. See, e.g., Graham et al. (1973) Virology, 52:456; Sambrook et al. (1989) Molecular Cloning, a laboratory manual, Cold Spring Harbor Laboratories, New York; Davis et al. (1986) Basic Methods in Molecular Biology, Elsevier; and Chu et al. (1981) Gene 13:197. Such techniques can be used to introduce one or more exogenous nucleic acids, such as nucleotide integration vectors and other nucleic acid molecules, into a suitable host cell.
[0099] "Host cell" means any cell that harbors or is capable of harboring a substance of interest. Typically, the host cell is a mammalian cell. In some embodiments, the host cell is a neuron, optionally a GABAergic neuron. As used herein, a "GABAergic neuron" is a nerve cell that produces gamma-aminobutyric acid (GABA). In mammals, GABA is a neurotransmitter that is widely distributed in the nervous system and binds to and inhibits the neurons to which it binds. Thus, GABA is associated with many disorders that affect the nervous system, including epilepsy, autism, and anxiety. Studies of SCN1A hemizygous and knockout mice have observed severe sodium current defects in GABAergic neurons in the brain. A host cell can serve as a recipient for an AAV helper construct, an AAV minigene plasmid, a helper function vector, or other transfer DNA associated with the production of recombinant AAV. The term includes progeny of the original cell that have been transfected. Thus, as used herein, a "host cell" can refer to a cell that has been transfected with an exogenous DNA sequence. It should be understood that due to natural, accidental, or intentional mutations, the progeny of a single parental cell may not be identical in morphology or in genomic or total DNA complementary sequences to the original parent.
[0100] As used herein, the term "cell line" refers to a population of cells capable of continuous or extended growth and division in vitro. Typically, a cell line is a clonal population derived from a single progenitor cell. It is further known in the art that karyotypic changes can occur spontaneously or be induced during the storage or transfer of such clonal populations. Thus, cells derived from the indicated cell line may not be identical to the ancestral cells or culture, and the indicated cell line includes such variants.
[0101] As used herein, the term "recombinant cell" refers to a cell that has been introduced with an exogenous DNA fragment, such as a DNA fragment that results in the transcription of a bioactive polypeptide or the production of a bioactive nucleic acid such as RNA.
[0102] As used herein, the term "vector" includes any genetic element, such as a plasmid, phage, transposon, cosmid, chromosome, artificial chromosome, virus, virion, etc., that is capable of replicating when associated with appropriate control elements and that can transfer gene sequences between cells. In some embodiments, the vector is a viral vector, such as an rAAV vector, a lentiviral vector, an adenoviral vector, a retroviral vector, etc. Thus, the term includes cloning and expression vectors, as well as viral vectors. In some embodiments, the useful vectors contemplated are those in which the nucleic acid fragment to be transcribed is under the transcriptional control of a promoter.
[0103] "Promoter" refers to a DNA sequence recognized by the synthetic machinery of a cell that needs to initiate specific transcription of a gene or introduced synthetic machinery. The phrases "operably linked," "operably positioned," "controlled," or "under transcriptional control" mean that the promoter is in the correct position and orientation relative to the nucleic acid to control the initiation of RNA polymerase and the expression of the gene. The term "expression vector or construct" means any type of genetic construct containing a nucleic acid, where part or all of the nucleic acid coding sequence is capable of being transcribed. In some embodiments, expression includes the transcription of the nucleic acid to produce, for example, a bioactive polypeptide product from the transcribed gene. The foregoing methods for packaging recombinant vectors in the desired AAV capsid to produce the rAAV of the present disclosure are not meant to be limiting, and other suitable methods will be apparent to those skilled in the art.
[0104] Methods for regulating the expression of a target gene
[0105] The present disclosure provides methods for regulating gene expression in a cell or a subject. The methods generally involve administering to the cell or subject an isolated nucleic acid or rAAV that contains a transgene encoding a fusion protein comprising a DNA binding domain (e.g., a ZFP domain) and a transactivation domain. In some embodiments, the fusion protein comprises a ZFP and a VP64 transactivation factor. In some embodiments, the fusion protein comprises a ZFP and a p65 transactivation factor. In some embodiments, the fusion protein comprises a ZFP and an RTA transactivation factor. In some embodiments, the fusion protein comprises a ZFP and a VPR transactivation factor. In some embodiments, the method involves administering to the cell or subject a dCas9 protein and at least one guide nucleic acid targeting SCN1A (e.g., a guide nucleic acid comprising any one of SEQ ID NOs: 83-94 or encoded by any one of SEQ ID NOs: 83-94).
[0106] In some embodiments, administering to the cell or subject an isolated nucleic acid or rAAV encoding a fusion protein (e.g., a fusion protein comprising a transactivation factor) results in increased expression of a target gene (e.g., SCN1A). Thus, in some embodiments, the compositions and methods described in the present disclosure can be used to treat conditions caused by haploinsufficiency of a target gene, such as Dravet syndrome caused by haploinsufficiency of the SCN1A gene.
[0107] As used herein, "haploinsufficiency" refers to a genetic disorder in which one copy of a gene (e.g., SCN1A) is inactivated, for example, by a gene mutation, or deleted, and the remaining functional copy of the gene is not sufficient to produce an amount of gene product sufficient to maintain normal function of the gene.
[0108] Dravet syndrome (also known as severe myoclonic epilepsy of infancy) is a rare lifelong epilepsy that typically appears in the first three years of life. Dravet syndrome is characterized by long and frequent seizures, behavioral and developmental delays, motor and balance problems, language and speech delays, and autonomic nervous system disorders. In some embodiments, the subject has haploinsufficiency associated with Dravet syndrome, such as a mutation in one copy of the SCN1A gene, resulting in a decrease in SCN1A protein in the cell or subject. Most Dravet syndrome patients carry SCN1A mutations that are translated into truncated proteins; other SCN1A mutations associated with Dravet syndrome include splice site and missense mutations, as well as mutations randomly distributed throughout the SCN1A gene. In some embodiments, the fusion proteins of the present disclosure comprise a ZFP domain that specifically targets (e.g., binds) the SCN1A gene and a transactivation domain. In some embodiments, the composition for targeting SCNA1 comprises (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds) the SCN1A gene.
[0109] In some embodiments, the subject has haploinsufficiency associated with MED13L haploinsufficiency syndrome, wherein the subject has only a single functional copy of the MED13L gene. Subjects with MED13L haploinsufficiency syndrome typically have a mutation in the second non-functional copy of the MED13L gene. MED13L haploinsufficiency syndrome is characterized by intellectual disability, speech problems, unique facial features, and developmental delays. In some embodiments, the fusion proteins of the present disclosure comprise a ZFP domain that specifically targets (e.g., binds) the MED13L gene and a transactivation domain. In some embodiments, the composition for targeting MED13L comprises (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds) the MED13L gene.
[0110] In some embodiments, the subject has haploinsufficiency associated with myelodysplastic syndrome. Subjects with myelodysplastic syndrome typically have a mutation in one copy of the isocitrate dehydrogenase 1 (IDH1), isocitrate dehydrogenase 2 (IDH2), and / or GATA2 gene. Myelodysplastic syndrome is a group of cancers in which immature blood cells in the bone marrow do not mature into healthy blood cells. Sometimes, this syndrome can lead to acute myeloid leukemia. In some embodiments, the fusion protein of the present disclosure comprises a ZFP domain that specifically targets (e.g., binds) the IDH1 gene and a transactivation domain. In some embodiments, the composition for targeting IDH1 comprises (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds) the IDH1 gene. In some embodiments, the fusion protein of the present disclosure comprises a ZFP domain that specifically targets (e.g., binds) the IDH2 gene and a transactivation domain. In some embodiments, the composition for targeting IDH2 comprises (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds) the IDH2 gene. In some embodiments, the fusion protein of the present disclosure comprises a ZFP domain that specifically targets (e.g., binds) the GATA2 gene and a transactivation domain. In some embodiments, the composition for targeting GATA2 comprises (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds) the GATA2 gene.
[0111] In some embodiments, the subject has haploinsufficiency associated with DiGeorge syndrome. Subjects with DiGeorge syndrome typically have a deletion of 30 to 40 genes at a location called 22q11.2 in the middle of chromosome 22. In particular, the disease may be characterized by haploinsufficiency of the TBX gene. DiGeorge syndrome is characterized by congenital heart problems, specific facial features, frequent infections, developmental delays, learning problems, and cleft palate. In some embodiments, the fusion protein of the present disclosure comprises a ZFP domain that specifically targets (e.g., binds) the TBX gene and a transactivation domain. In some embodiments, the composition for targeting TBX comprises (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds) the TBX gene.
[0112] In some embodiments, the subject has haploinsufficiency associated with CHARGE syndrome. In most cases, subjects with CHARGE syndrome have haploinsufficiency of the CHD7 gene. CHARGE syndrome is characterized by eye defects, heart defects, choanal atresia, growth and / or developmental delay, genital and / or urinary tract abnormalities, and ear abnormalities and deafness. In some embodiments, the fusion proteins of the present disclosure comprise a ZFP domain that specifically targets (e.g., binds to) the CHD7 gene and a transactivation domain. In some embodiments, the compositions for targeting CHD7 comprise (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds to) the CHD7 gene.
[0113] In some embodiments, the subject has haploinsufficiency associated with Ehlers–Danlos syndrome. Subjects with Ehlers–Danlos syndrome may have haploinsufficiency of the COL1A1, COL1A2, COL3A1, COL5A1, COL5A2, TNXB, ADAMTS2, PLOD1, B4GALT7, DSE, and / or D4ST1 / CHST14 genes. Ehlers–Danlos syndrome is characterized by overly elastic skin and can result in aortic dissection, scoliosis, and early-onset osteoarthritis. In some embodiments, the fusion proteins of the present disclosure comprise a ZFP domain that specifically targets (e.g., binds to) any one of the COL1A1, COL1A2, COL3A1, COL5A1, COL5A2, TNXB, ADAMTS2, PLOD1, B4GALT7, DSE, or D4ST1 / CHST14 genes and a transactivation domain. In some embodiments, the compositions for targeting any one of COL1A1, COL1A2, COL3A1, COL5A1, COL5A2, TNXB, ADAMTS2, PLOD1, B4GALT7, DSE, or D4ST1 / CHST14 comprise (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds to) any one of the COL1A1, COL1A2, COL3A1, COL5A1, COL5A2, TNXB, ADAMTS2, PLOD1, B4GALT7, DSE, or D4ST1 / CHST14 genes.
[0114] In some embodiments, the subject has haploinsufficiency associated with frontotemporal dementia. Subjects with FTD have haploinsufficiency of the MAPT gene encoding the Tau protein and / or the GRN gene. FTD is characterized by memory loss, lack of social awareness, poor impulse control, and speech difficulties. In some embodiments, the fusion proteins of the present disclosure comprise a ZFP domain that specifically targets (e.g., binds to) the MAPT gene and a transactivation domain. In some embodiments, the composition for targeting MAPT comprises (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds to) the MAPT gene. In some embodiments, the fusion proteins of the present disclosure comprise a ZFP domain that specifically targets (e.g., binds to) the GRN gene and a transactivation domain. In some embodiments, the composition for targeting GRN comprises (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds to) the GRN gene.
[0115] In some embodiments, the subject has haploinsufficiency associated with Holt–Oram syndrome. Subjects with Holt–Oram syndrome have haploinsufficiency of the TBX5 gene. Holt–Oram syndrome is characterized by cardiac complications, including congenital heart defects and cardiac conduction disease. In some embodiments, the fusion proteins of the present disclosure comprise a ZFP domain that specifically targets (e.g., binds to) the TBX5 gene and a transactivation domain. In some embodiments, the composition for targeting TBX5 comprises (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds to) the TBX5 gene.
[0116] In some embodiments, the subject has haploinsufficiency associated with Marfan syndrome. Subjects with Marfan syndrome typically have haploinsufficiency of the FBN1 gene encoding the fibrillin-1 protein. Marfan syndrome is characterized by disproportionate limb length, early-onset arthritis, cardiac complications, and / or autonomic nervous system dysfunction. In some embodiments, the fusion proteins of the present disclosure comprise a ZFP domain that specifically targets (e.g., binds to) the FBN1 gene and a transactivation domain. In some embodiments, the composition for targeting FBN1 comprises (i) a fusion protein comprising a dCas protein and a transactivation domain, and (ii) a guide nucleic acid (e.g., gRNA) that specifically targets (e.g., binds to) the FBN1 gene.
[0117] The present disclosure is in part based on methods of administering a fusion protein as described herein to a subject. In some embodiments, the fusion protein comprises a DBD and a transcriptional activator. In some embodiments, the DBD is a ZNF, a TALE, a dCas protein (e.g., dCas9 or dCas12a), or a homeodomain that binds to the SCN1A gene. In some embodiments, the transcriptional activator is VP64, p65, RTA, or a ternary transcriptional activator comprising VP64-p65-RTA (VPR). In some embodiments, the fusion protein is flanked by AAV inverted terminal repeat (ITR) sequences. In some embodiments, the fusion protein is operably linked to a promoter. In some embodiments, the subject has or is suspected of having an SCN1A mutation that results in SCN1A haploinsufficiency. In some embodiments, the subject has or is suspected of having Dravet syndrome.
[0118] In some aspects, the present disclosure provides methods of modulating (e.g., increasing, decreasing, etc.) the expression of a target gene in a cell. In some embodiments, the present disclosure provides methods of increasing the expression of a target gene (e.g., SCN1A) in a cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is in a subject (e.g., in vivo). In some embodiments, the subject is a mammalian subject, such as a human. In some embodiments, the cell is a nervous system cell (a central nervous system cell or a peripheral nervous system cell), such as a neuron (e.g., a GABAergic neuron, a unipolar neuron, a bipolar neuron, a basket cell, a Betz cell, a Lugaro cell, a spiny neuron, a Purkinje cell, a pyramidal cell, a Renshaw cell, a granule cell, a motor neuron, a fusiform cell, etc.) or a glial cell (e.g., an astrocyte, an oligodendrocyte, an ependymal cell, a radial glial cell, a Schwann cell, a satellite cell, etc.).
[0119] In a "normal" cell or subject, expression of a target gene (e.g., SCN1A) is sufficient such that the cell or subject is not haploinsufficient with respect to the target gene (e.g., SCN1A). In some embodiments, "improved" or "increased" expression or activity of a transgene is measured relative to the expression or activity of the transgene in a cell or subject that has not been administered one or more isolated nucleic acids, rAAV, or compositions as described herein. In some embodiments, "improved" or "increased" expression or activity of a transgene is measured relative to the expression or activity of the transgene in a subject after administration (e.g., measuring gene expression before and after administration) of one or more isolated nucleic acids, rAAV, or compositions as described herein. For example, in some embodiments, "improved" or "increased" expression of SCN1A in a cell or subject is measured relative to a cell or subject that has not been administered a transgene encoding a fusion ZFP transactivator. In some embodiments, the methods described in the present disclosure result in a 2-fold to 100-fold (e.g., 2-fold, 5-fold, 10-fold, 50-fold, 100-fold, etc.) increase in SCN1A expression and / or activity in a subject relative to SCN1A expression and / or activity in a subject that has not been administered one or more compositions described in the present disclosure.
[0120] As used herein, the terms "treatment / treating" and "therapy" refer to therapeutic treatment and prophylactic / preventative operations. The terms also include ameliorating existing symptoms, preventing additional symptoms, ameliorating or preventing the root cause of symptoms, preventing or reversing the cause of symptoms, e.g., symptoms associated with a haploinsufficient gene (e.g., the haploinsufficient SCN1A gene). Thus, the terms denote that a beneficial result has been conferred on a subject having a disorder (e.g., a disease or condition associated with a haploinsufficient gene, e.g., Dravet syndrome) or having the potential to develop such a disorder. In addition, the term "treatment" also includes the application or administration of an agent (e.g., a therapeutic agent or therapeutic composition, e.g., an isolated nucleic acid or rAAV that targets or binds to a target gene or a regulatory region of a target gene) to a subject or an isolated tissue or cell line from a subject that may have a disease, disease symptom, or disease predisposition, with the aim of treating, curing, alleviating, reducing, modifying, remedying, mitigating, improving, or affecting the disease, disease symptom, or disease predisposition.
[0121] A therapeutic agent or composition may include a compound in a pharmaceutically acceptable form that prevents and / or reduces the symptoms of a particular disease (e.g., a disease or condition associated with a haploinsufficient gene, such as Dravet syndrome). For example, the therapeutic composition may be a pharmaceutical composition that prevents and / or reduces the symptoms of a disease or condition associated with a haploinsufficient gene (e.g., Dravet syndrome). It is contemplated that the therapeutic compositions of the present invention will be provided in any suitable form. The form of the therapeutic composition will depend on a number of factors, including the mode of administration as described herein. The therapeutic composition may contain diluents, adjuvants, excipients, and other ingredients as described herein.
[0122] Mode of administration
[0123] The isolated nucleic acids, rAAVs, and compositions of the present disclosure can be delivered to a subject in a composition form by any suitable method known in the art. For example, rAAV, preferably suspended in a physiologically compatible carrier (e.g., in a composition), can be administered to a subject, i.e., a host animal, such as a human, mouse, rat, cat, dog, sheep, rabbit, horse, cow, goat, pig, guinea pig, hamster, chicken, turkey, or non-human primate (e.g., macaque). In some embodiments, the host animal does not include a human.
[0124] Delivery of rAAV to a mammalian subject can be by, for example, intramuscular injection or by administration into the bloodstream of the mammalian subject. Administration into the bloodstream can be by injection into a vein, artery, or any other blood vessel catheter. In some embodiments, rAAV is administered into the bloodstream by isolated limb perfusion, a technique well known in the surgical arts that essentially enables one of ordinary skill in the art to isolate a limb from the systemic circulation prior to administration of the rAAV viral particles. Variations of the isolated limb perfusion technique described in U.S. Patent No. 6,177,403 can also be used by one of ordinary skill in the art to administer viral particles into the vasculature of an isolated limb to potentially enhance transduction into muscle cells or tissues. In addition, in certain instances, it may be desirable to deliver viral particles to the CNS of a subject. "CNS" means all cells and tissues of the vertebrate brain and spinal cord. Thus, the term includes, but is not limited to, neuronal cells, glial cells, astrocytes, cerebrospinal fluid (CSF), interstitial spaces, bone, cartilage, and the like. Recombinant AAV can be directly delivered to the CNS or brain by injection, using neurosurgical techniques known in the art, with a needle, catheter, or related device, into, for example, ventricular regions as well as the striatum (e.g., the caudate or putamen of the striatum), thalamus, spinal cord, and neuromuscular junctions or cerebellar lobules, such as by stereotactic injection (see, e.g., Stein et al., J Virol 73:3424-3429, 1999; Davidson et al., PNAS 97:3428-3432, 2000; Davidson et al., Nat. Genet. 3:219-223, 1993; and Alisky and Davidson, Hum. Gene Ther. 11:2315-2329, 2000). In some embodiments, rAAV as described in the present disclosure is administered by intravenous injection. In some embodiments, rAAV is administered by intracerebral injection. In some embodiments, rAAV is administered by intrathecal injection. In some embodiments, rAAV is administered by intrastriatal injection. In some embodiments, rAAV is delivered by intracranial injection. In some embodiments, rAAV is delivered by cisterna magna injection. In some embodiments, rAAV is delivered by lateral cerebral ventricle injection.
[0125] Aspects of the present disclosure relate to a composition comprising a recombinant AAV, the recombinant AAV comprising a capsid protein and a nucleic acid encoding a transgene, wherein the transgene comprises a nucleic acid sequence encoding one or more proteins. In some embodiments, the nucleic acid further comprises AAV ITRs. In some embodiments, the composition further comprises a pharmaceutically acceptable carrier.
[0126] The compositions of the present disclosure may comprise rAAV alone or in combination with one or more other viruses (e.g., a second rAAV encoding one or more different transgenes). In some embodiments, the composition comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more different rAAVs, each having one or more different transgenes.
[0127] Given the indication targeted by the rAAV, one of ordinary skill in the art can readily select a suitable carrier. For example, a suitable carrier includes saline, which can be formulated with a variety of buffer solutions (e.g., phosphate buffered saline). Other exemplary carriers include sterile saline, lactose, sucrose, calcium phosphate, gelatin, dextran, agar, pectin, peanut oil, sesame oil, and water. The choice of carrier is not a limitation of the present disclosure.
[0128] Optionally, in addition to the rAAV and one or more carriers, the compositions of the present disclosure may also contain other conventional pharmaceutical ingredients, such as preservatives or chemical stabilizers. Suitable exemplary preservatives include chlorobutanol, potassium sorbate, sorbic acid, sulfur dioxide, propyl gallate, parabens, ethyl vanillin, glycerol, phenol, p-chlorophenol, and poloxamers (nonionic surfactants) such as F-68. Suitable chemical stabilizers include gelatin and albumin.
[0129] The rAAV is administered in an amount sufficient to transfect the cells of the desired tissue and provide a sufficient level of gene transfer and expression without producing excessive side effects. Conventional and pharmaceutically acceptable routes of administration include, but are not limited to, direct delivery to the selected organ (e.g., delivery to the liver via the portal vein), oral, inhalation (including intranasal and intratracheal delivery), intraocular, intravenous, intramuscular, subcutaneous, intradermal, intratumoral, and other parenteral routes of administration. If desired, routes of administration may be combined.
[0130] The dose of rAAV viral particles required to achieve a particular "therapeutic effect" (e.g., in dose units of genome copies per kilogram body weight (GC / kg)) will vary depending on a variety of factors, including but not limited to: the route of administration of the rAAV viral particles, the level of gene or RNA expression required to achieve the therapeutic effect, the particular disease or disorder being treated, and the stability of the gene or RNA product. One of ordinary skill in the art can readily determine the range of rAAV viral particle doses for treating a patient with a particular disease or disorder based on the foregoing factors as well as other factors well known in the art.
[0131] An effective amount of rAAV is an amount sufficient to target infection of an animal and the desired tissue. In some embodiments, an effective amount of rAAV is administered to a subject during the pre-symptomatic stage of a lysosomal storage disease. In some embodiments, the pre-symptomatic stage of a lysosomal storage disease occurs between birth (e.g., perinatal) and 4 weeks of age.
[0132] In some embodiments, the rAAV composition is formulated to reduce aggregation of AAV particles in the composition, particularly in the presence of high rAAV concentrations (e.g., about 10 13 GC / mL or higher). Methods for reducing rAAV aggregation are well known in the art and include, for example, adding surfactants, pH adjustment, salt concentration adjustment, etc. (See, e.g., Wright FR et al., Molecular Therapy (2005) 12, 171–178, the content of which is incorporated herein by reference.)
[0133] Those skilled in the art are familiar with the formulation of pharmaceutically acceptable excipients and carrier solutions, as well as the development of appropriate dosages and treatment regimens for using the specific compositions described herein in various treatment regimens.
[0134] Generally, these formulations may contain at least about 0.1% of the active compound or more, although the percentage of one or more active ingredients may of course vary and may conveniently be between about 1% or 2% and about 70% or 80% or more of the total weight or volume of the formulation. Naturally, the amount of the active compound in each therapeutically useful composition can be prepared in such a way that a suitable dosage will be obtained in any given unit dose of the compound. Those skilled in the art of preparing such pharmaceutical formulations will envision factors such as solubility, bioavailability, biological half-life, route of administration, product shelf-life, and other pharmacological considerations, and thus, a variety of dosages and treatment regimens may be desirable.
[0135] In certain cases, it is desirable to deliver the rAAV-based therapeutic construct in a suitably formulated pharmaceutical composition by subcutaneous, intra-pancreatic, intranasal, parenteral, intravenous, intramuscular, intrathecal or oral, intraperitoneal or by inhalation. In some embodiments, the modes of administration described in U.S. Patent Nos. 5,543,158; 5,641,515 and 5,399,363 (each specifically incorporated herein by reference in its entirety) may be used to deliver rAAV. In some embodiments, the preferred mode of administration is by portal vein injection.
[0136] Pharmaceutical dosage forms suitable for injection include sterile aqueous solutions or dispersions and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersions. The dispersions can also be prepared in glycerol, liquid polyethylene glycols, and their mixtures and in oils. Under ordinary storage and use conditions, these preparations contain preservatives to prevent the growth of microorganisms. In many cases, the form is sterile and is a fluid that is easy to inject. Under the manufacturing and storage conditions, the composition must be stable and must be protected against the contaminating action of microorganisms such as bacteria and fungi during storage. The carrier can be a solvent or a dispersion medium that contains, for example, water, ethanol, polyols (e.g., glycerol, propylene glycol, and liquid polyethylene glycols, etc.), suitable mixtures thereof, and / or vegetable oils. For example, appropriate fluidity can be maintained by the use of coatings such as lecithin, by maintaining the required particle size (in the case of dispersions), and by the use of surfactants. Prevention of microbial action can be achieved by various antibacterial and antifungal agents (e.g., parabens, chlorobutanol, phenol, sorbic acid, thimerosal, etc.). In many cases, it will be preferable to include isotonic agents, such as sugars or sodium chloride. Prolonged absorption of injectable compositions can be achieved by the use of agents that delay absorption (e.g., aluminum monostearate and gelatin) in the composition.
[0137] For example, for the administration of an injectable aqueous solution, if necessary, the solution can be appropriately buffered, and the liquid diluent is first made isotonic with sufficient saline or glucose. These particular aqueous solutions are particularly suitable for intravenous, intramuscular, subcutaneous, and intraperitoneal administration. In this regard, sterile aqueous media that can be used are known to those skilled in the art. For example, one dose can be dissolved in 1 mL of isotonic NaCl solution and then added to 1000 mL of subcutaneous infusion or injected at the proposed infusion site (see, for example, "Remington's Pharmaceutical Sciences", 15th edition, pages 1035 - 1038 and 1570 - 1580). Depending on the condition of the host, the dosage will necessarily vary to some extent. In any case, the person responsible for administration will determine the appropriate dosage for each individual host.
[0138] Sterile injectable solutions are prepared by incorporating the active rAAV in the required amount, if necessary, together with various other ingredients enumerated herein, into a suitable solvent, followed by filtration sterilization. Generally, dispersions are prepared by incorporating the various sterilized active ingredients into a sterile vehicle that contains a basic dispersion medium and the required other ingredients from those enumerated above. In the case of sterile powders for the preparation of sterile injectable solutions, the preferred methods of preparation are vacuum drying techniques and freeze-drying techniques, whereby a powder of the active ingredient plus any additional required ingredients from its previously sterile filtered solution is produced.
[0139] The rAAV compositions disclosed herein can also be formulated in neutral or salt forms. Pharmaceutically acceptable salts include acid addition salts (formed with the free amino groups of the protein) formed with inorganic acids such as, for example, hydrochloric acid or phosphoric acid and organic acids such as acetic acid, oxalic acid, tartaric acid, mandelic acid, etc. Salts formed with free carboxyl groups can also be derived from inorganic bases such as, for example, sodium hydroxide, potassium hydroxide, ammonium hydroxide, calcium hydroxide or ferric hydroxide and organic bases such as isopropylamine, trimethylamine, histidine, procaine, etc. When formulated, the solution will be administered in a manner compatible with the dosage formulation and in a therapeutically effective amount. The formulations are readily administered in a variety of dosage forms such as, for example, injectable solutions, drug release capsules, etc.
[0140] As used herein, "formulation" includes any and all solvents, dispersion media, vehicles, coatings, diluents, antibacterial and antifungal agents, isotonic and absorption delaying agents, buffers, carrier solutions, suspensions, colloids, etc. The use of such media and agents for pharmaceutical active substances is well known in the art. Supplementary active ingredients can also be incorporated into the compositions. The phrase "pharmaceutically acceptable" refers to molecular entities and compositions that do not produce an allergic or similar untoward reaction when administered to a host.
[0141] Delivery vehicles such as liposomes, nanocapsules, microparticles, microspheres, lipid particles, vesicles, etc. can be used to introduce the compositions of the present disclosure into suitable host cells. In particular, the transgenes delivered by the rAAV vectors can be formulated for delivery encapsulated in lipid particles, liposomes, vesicles, nanospheres or nanoparticles, etc.
[0142] Such formulations can preferably be pharmaceutically acceptable formulations for introducing the nucleic acids or rAAV constructs disclosed herein. The formation and use of liposomes are generally known to those skilled in the art. Recently, the development of liposomes has improved serum stability and circulation half-life (U.S. Patent No. 5,741,516). In addition, various methods of liposomes and liposome-like formulations as potential drug carriers have been described (U.S. Patent Nos. 5,567,434; 5,552,157; 5,565,213; 5,738,868 and 5,795,587).
[0143] Liposomes have been successfully used in many cell types that are typically resistant to transfection by other procedures. In addition, liposomes are not subject to the DNA length limitations typical of virus-based delivery systems. Liposomes have been effectively used to introduce genes, drugs, radiotherapeutic agents, viruses, transcription factors and allosteric effectors into a variety of cultured cell lines and animals. In addition, several successful clinical trials have been completed to examine the effectiveness of liposome-mediated drug delivery.
[0144] Liposomes are formed from phospholipids that are dispersed in an aqueous medium and spontaneously form multilamellar concentric bilayer vesicles (also known as multilamellar vesicles (MLV)). MLVs typically have a diameter of 25 nm to 4 μm. Sonication of MLVs results in the formation of small unilamellar vesicles (SUV) with an aqueous solution in the core having a diameter in the range of to range.
[0145] Alternatively, a nanocapsule formulation of rAAV can be used. Nanocapsules can generally capture substances in a stable and reproducible manner. To avoid side effects due to intracellular polymer overload, polymers that can be degraded in vivo should be used to design such ultrafine particles (sized approximately 0.1 μm). Biodegradable polyalkylcyanoacrylate nanoparticles that meet these requirements are envisioned.
[0146] In addition to the above delivery methods, the following techniques are also envisioned as alternative methods for delivering rAAV compositions to a host. Sonophoresis (i.e., ultrasound) has been used and described in U.S. Patent No. 5,656,016 as a device for enhancing the rate and efficacy of drug penetration into and through the circulatory system. Other alternative drug delivery options envisioned are intraosseous injection (U.S. Patent No. 5,779,708), microchip devices (U.S. Patent No. 5,797,898), ophthalmic preparations (Bourlais et al., 1998), transdermal matrices (U.S. Patent Nos. 5,770,219 and 5,783,208), and feedback-controlled delivery (U.S. Patent No. 5,697,899).
[0147] Examples
[0148] Example 1. Design of zinc finger proteins that upregulate SCN1A gene expression
[0149] The homologous region between the human (HEK293T cells) and mouse (HEPG2 cells) SCN1A promoter sequences was identified by aligning the sequences around two prominent transcription start sites for each species identified in the RIKEN CAGE-seq dataset ( Figure 1 ). A highly conserved sequence between human (HEK) and mouse (HEPG2) exists in the proximal promoter region of SCN1A ( Figure 2 ). Three ZFPs consisting of six fingers were designed to bind to overlapping 15-22 nucleotide homologous regions in the SCN1A proximal promoter region by assembling one-finger and two-finger modules with predefined DNA-binding specificities ( Figure 3 ). Three ZFPs (ZFP1-ZFP3), each consisting of six fingers, were designed to bind to Figure 3The overlapping highly conserved sequences identified in. Each finger is designed to bind to a three-base region (triplet) in the highly conserved region of the SCN1A proximal promoter.
[0150] ZFP-1 recognizes a separate three-base region (DNA triplets shown in red, separated by "·") in the proximal promoter region of the SCN1A gene (SEQ ID NO: 2), as Figure 4A shown. Each recognition helix (seven amino acids) of fingers 1 to 6 of ZFP-1 binds to three nucleotide sequences, as Figure 4B shown. The amino acid sequences of the six fingers of ZFP-1 (SEQ ID NO: 17-22) are shown in Figure 4C ; the linkers between the fingers are highlighted to designate classical (TGEKP) and non-classical (TGSQKP) linker sequences. The nucleotide sequences of the six fingers of ZFP-1 (SEQ ID NO: 11-16) are shown in Figure 4D .
[0151] Table 1. ZFP-1 recognition helices targeting SCN1A
[0152]
[0153] ZFP-2 recognizes a separate three-base region (DNA triplets shown in red, separated by "·") in the proximal promoter region of the SCN1A gene (SEQ ID NO: 3), as Figure 5A shown. Each recognition helix (seven amino acids) of fingers 1 to 6 of ZFP-2 binds to three nucleotide sequences, as Figure 5B shown. The amino acid sequences of the six fingers of ZFP-2 (SEQ ID NO: 29-34) are shown in Figure 5C ; the linkers between the fingers are highlighted to designate classical (TGEKP) and non-classical (TGSQKP) linker sequences. The nucleotide sequences of the six fingers of ZFP-1 (SEQ ID NO: 23-28) are shown in Figure 5D .
[0154] Table 2. ZFP-2 recognition helices targeting SCN1A
[0155]
[0156] ZFP-3 recognizes a separate three-base region (DNA triplets shown in red, separated by "·") in the proximal promoter region of the SCN1A gene (SEQ ID NO: 4), as Figure 6A shown. Each recognition helix (seven amino acids) of fingers 1 to 6 of ZFP-3 binds to three nucleotide sequences, as Figure 6Bshown. The amino acid sequences of the six fingers of ZFP-3 (SEQ ID NO: 41-46) are shown in Figure 6C ; the linkers between the fingers are highlighted to designate classical (TGEKP) and non-classical (TGSQKP) linker sequences. The nucleotide sequences of the six fingers of ZFP-1 (SEQ ID NO: 35-40) are shown in Figure 6D .
[0157] Table 3. ZFP-3 recognition helices targeting SCN1A
[0158]
[0159]
[0160] Additional ZFPs designed to target conserved sequences in the proximal promoter region of the SCN1A gene will each contain five or six finger domains and will bind to regions of 15-22 nucleotides that are highly conserved between human and mouse SCN1A.
[0161] Table 4. Zinc finger proteins targeting SCN1A
[0162]
[0163]
[0164] Example 2. ZFPs increase the expression of the SCN1A gene in human cells
[0165] To examine the ability of ZFP1-ZFP3 to upregulate SCN1A transcription, the ZFP1-ZFP3 DNA-binding domains were fused to the hybrid VP64, p53, and RTA (VPR) triple-strength transcriptional activator domain to form chimeric transactivators. The VPR fusion activator domain is used to recruit transcriptional regulatory complexes and increase chromatin accessibility, and contributes to achieving high levels of gene expression. Thus, the ZFP domain targets the VPR activator to highly conserved sequences in the proximal promoter region to increase SCN1A gene expression.
[0166] The expression plasmids encoding the VPR-ZFP1, VPR-ZFP2, and / or VPR-ZFP3 fusion proteins were transfected into HEK293 cells by transient transfection, and SCN1A gene expression was measured by qRT-PCR (using TBP expression as a normalization reference). The VPR-ZFP fusions comprise ZFP1, ZFP2, and / or ZFP3 fused to VPR. Transfection of three constructs for multiplex regulation containing the DNA-binding domains of ZFP1, ZFP2, and ZFP3, each fused to VPR, resulted in a 45-fold increase in SCN1A gene expression relative to untransfected cells, indicating that the VPR-ZFP chimeric transactivators are capable of increasing SCN1A gene expression by binding to the promoter-proximal region of the gene( Figure 7 ).
[0167] The VPR-[ZFP1-ZFP3] fusion protein and VPR-ZFP fusion proteins in which the ZFP DNA-binding domains are currently being engineered are being transfected into HeLa and HEPG2 cells, both of which have low levels of SCN1A expression. The VPR-ZFP fusion proteins contain either a single ZFP DNA-binding domain or a combination of multiple ZFP DNA-binding domains fused to the VPR transactivator domain. SCN1A gene expression is measured by qRT-PCR to determine whether these VPR-ZFP fusions are capable of increasing gene expression. The ability of the most promising VPR-ZFP fusion candidates to increase SCN1A expression is tested in primary mouse cortical neurons after delivery of the fusion protein by adeno-associated virus (AAV).
[0168] The specificity of the ZFP domains is being further optimized using a bacterial one-hybrid selection system (see, e.g., Meng et al., “Targeted gene inactivation in zebrafish using engineered zinc-fingernucleases,” Nat Biotechnol, 2008) to identify the ideal ZFPs from a random library in which the residues important for DNA binding are diversified. The newly selected ZFPs will be fused to the VPR transactivator domain, either individually or in combination of multiple ZFPs, and transfected into HEK293, HeLa, and HEPG2 cells as well as primary mouse cortical neurons to identify the candidate ZFP domains that most potently increase SCN1A gene expression after qRT-PCR analysis.
[0169] Example 3. Generation of ZFPs with Different Potencies SCN1A Transactivator Series
[0170] In Example 2, the most effective ZFPs that upregulate SCN1A gene expression were fused to a series of human transactivation domains with an expected potency gradient (e.g., Rta, p65, Hsf1, etc.) to identify assemblies that achieve a 2-fold upregulation of SCN1A gene expression within the range of AAV multiplicity of infection (MOI). AAV vectors expressing ZFP SCN1A fusion transactivators were used to infect primary murine cortical neurons from normal and SCN1A + / - mice. The expression levels of Na V v1.1 protein were evaluated using western blotting and qPCR. Primary neurons treated with TGF-α for 8 hours were used as positive controls because this treatment increased Na V v1.1 protein expression by approximately 6- to 8-fold (Chen et al., 2015, Neuroinflammation 12:126). Changes in the expression levels of other Na V α subunit genes were also evaluated to demonstrate the specificity of ZFP SCN1A transactivation. Immunofluorescence was used to determine whether Na SCN1A (HA tag) and markers specific for GABAergic neurons (e.g., parvalbumin + or somatostatin + ) or general neuronal markers (e.g., NeuN, TUBIII, and / or Map2) by double immunofluorescence staining with antibodies. Whether Na V v1.1 expression remained restricted to GABAergic interneurons. The specificity of ZFP SCN1A transactivation of the SCN1A gene was also evaluated by ChIP-Seq and RNA-Seq to map the genomic binding sites and transcriptome proliferation generated after gene transfer.
[0171] Example 4. Mapping of the histone organization and epigenome of the SCN1A promoter in GABAergic inhibitors that direct the design of promoter activity-dependent SCN1A-ZFP transactivators
[0172] The ability of a ZFP to bind to genomic targets depends on the accessibility of the target sequence (e.g., the presence of a nucleosome-free region). This requirement for DNA accessibility has been used to design ZFP transactivators that function only in subsets of cell types based on the presence of DNA target sequence accessibility. Additional restriction of cell type activity is achieved by using tissue-specific promoters for ZFP transactivator expression. Small promoters from the somatostatin and neuropeptide Y genes of the pufferfish (Takifugu rubripes) have been shown to drive highly specific transgene expression in cortical and hippocampal inhibitory interneurons in the context of AAV vectors and lentiviruses. In some embodiments, the combination of AAV-based transcriptional restriction of an SCN1A-specific ZFP sensitive to DNA accessibility results in highly specific upregulation of Na V 1.1 protein expression in inhibitory interneurons throughout the brain. This dual regulatory approach will minimize the potential side effects of ectopic expression of the Na V 1.1 protein in cells that do not normally express it.
[0173] The nucleosome structure and epigenetic landscape of the SCN1A promoter were analyzed in mouse and human GABAergic inhibitory and glutamatergic excitatory neurons. This information was used to design GABAergic inhibitory neuron-restricted ZFP transactivators by targeting sequences accessible only around the SCN1A locus in this cell type.
[0174] GABAergic inhibitory neurons of transgenic mice expressing TdTomato under the GAD67 promoter and GFP-positive glutamatergic excitatory neurons generated by crossing Emx1-IRES-Cre with ROSA26 / stop / EGFP mice were isolated using fluorescence-activated cell sorting (FACS). Human GABAergic and excitatory neurons were generated from induced pluripotent stem (iPS) cells, and markers specific to these cell types and electrophysiological activity were confirmed using immunostaining and RT-PCR. The accessible genomic regions around the SCN1A promoter were characterized in mouse and human neuronal populations using assay for transposase-accessible chromatin (ATAC-Seq).
[0175] Design ZFPs that recognize sequences accessible only in GABAergic neurons based on differential chromatin accessibility of genomic regions around the SCN1A promoter in inhibitory and excitatory neurons SCN1A transactivators. A series of candidate ZFP-VPR transactivator fusions are being generated to target different accessible regions of SCN1A, where binding of the transactivator is expected to effectively upregulate Na in the inhibitory region V 1.1 expression, as well as revealing Na in excitatory neurons VAny unwanted induced expression of 1.1.
[0176] Expression studies were performed in cultured human iPS-derived neurons and mouse SCN1A + / - primary neurons that mimic Dravet syndrome to determine whether a ZFP SCN1A transactivator designed to recognize DNA sequences accessible only in inhibitory neurons provides the necessary specificity when expressed from an AAV vector under the pan-neuronal human synapsin 1 or inhibitory interneuron-specific promoter. Na V 1.1 expression levels were determined by qRT-PCR, western blotting, and double immunofluorescence assays, which have inhibitory GABAergic (e.g., GABA + , GAD65 / 67 + , somatostatin, and / or parvalbumin) and excitatory glutamatergic (e.g., Cux1+, FoxG1, +, GABA A receptor, GABA - ) neuronal type-specific markers. The cell type specificity of the ZFP SCN1A transactivator was designed to target different sequences in the mouse and human SCN1A promoters because chromatin structure and DNA sequence in the syntenic region differ between species. Controls in these experiments included neuronal cultures infected with a similar AAV vector encoding GFP, a ZFP without a transactivation domain, or a transactivator without a ZFP DNA-binding domain.
[0177] MicroRNA (miRNA) binding sites were incorporated into the 3' untranslated region (3'UTR) of the ZFP SCN1A transactivator, which is restricted to the cell type (e.g., glutamatergic excitatory neurons) in which unwanted expression occurs. This method has previously been used to restrict the expression of AAV-delivered transgenes (Xie et al., “MicroRNA-regulated, systemically delivered rAAV9: a step closer to CNS-restricted transgene expression,” Mol. Ther. 2011). Differences in miRNA expression profiles between GABAergic inhibitory neurons and other cell types are being determined by small RNA sequencing.
[0178] Example 5. Evaluation of the potential of AAV-ZFP SCN1A gene therapy to correct sodium current defects in patient-derived iPS-generated GABAergic interneurons
[0179] Development of one or more ZFPs for Dravet syndromeSCN1A A key step in the transactivator is to demonstrate that these artificial transactivators have the desired function in human neurons. To this end, iPS cells are being obtained from Dravet patients (n = 4 - 6) and non-Dravet patients (n = 4). The non-Dravet genetic background is expressed in these cells without artificial manipulation of gene expression, and thus iPS cells have become the state-of-the-art cell line for biomedical research. The CRISPR-Cas9 genome editing technology is being used to create isogenic cell lines by repairing the genetic mutations in SCN1A to wild-type sequences or by introducing Dravet-related mutations into the normal alleles within control cell lines. Thus, the isogenic lines eliminate the natural variability that arises from comparing cell lines from different human subjects and are thus valuable for confirming and enhancing disease-specific phenotypes. The established inhibitory neuron differentiation protocol and validation pipeline are being used to differentiate the iPS cell lines into forebrain GABAergic inhibitory interneurons.
[0180] As determined by whole-cell patch-clamp electrophysiology measurements, inhibitory neurons from Dravet patients exhibit reduced sodium currents and impaired action potential firing. Similar measurements are being performed to confirm that the Dravet-derived neurons described herein recapitulate these disease-related phenotypes. The sodium current defect occurs in inhibitory neurons but not in excitatory neurons in Dravet patients (Sun et al.), and thus only inhibitory neurons are used in this disclosure. The mutation-induced sodium channel defect in inhibitory neurons derived from Dravet patients can be rescued by the ectopic expression of wild-type SCN1A (Reference 20). Thus, the methods described in this disclosure are applicable to testing the efficacy of ZFP SCN1A transactivators in restoring wild-type sodium channel function and physiology in the context of Dravet syndrome.
[0181] GABAergic inhibitory neuron cultures are infected with AAV vectors encoding ZFP SCN1A transactivators under a general neuronal or inhibitory neuron-specific promoter. The change in Na V 1.1 expression levels is being evaluated by Western blotting. The restoration of functional sodium currents in inhibitory neurons is being evaluated by whole-cell patch-clamp of untransfected cells compared to transfected cells. The binding of ZFP SCN1A transactivators to the genome within all patient-derived inhibitory neurons is being analyzed by ChIP-seq and correlated with any identified transcriptomic changes detected by RNA-seq. The controls in these experiments are neuronal cultures infected with similar AAV vectors encoding GFP, ZFP without the VPR transactivation domain, and the VPR transactivation domain without the ZFP DNA-binding domain.
[0182] Example 6. Evaluation of AAV-ZFP SCN1A Therapeutic potential of the intervention in SCN1A mice at different ages and delivery routes
[0183] The broad tropism of AAV is a key feature for gene therapy applications that widely express genes, but it can become a major challenge when the transgene of interest is expressed in a cell type-specific manner. This problem in major tissues of the body such as the liver, muscle, and heart has been largely solved by using tissue-specific promoters such as thyroxine-binding protein (TBP), creatine kinase, and troponin T, respectively. Additional levels of control can be superimposed on tissue-specific promoters to achieve a higher degree of off-targeting from specific tissues by incorporating multiple copies of binding sites for microRNAs highly abundant in those tissues (such as miR-122 in the liver and miR-1 in skeletal muscle). In the case of transducing a wide range of cell types, the recently described AAV-PHP.B serotype is very effective for CNS gene transfer after systemic delivery. In addition, its tropism for peripheral tissues is largely as broad as that of AAV9. The goal of the gene therapy approach for Dravet syndrome is to fully restore Na V 1.1 expression while preventing harmful effects from ectopic expression in other neurons and elsewhere. AAV and lentiviral vectors (<2.8 kb) encoding GFP under small promoters derived from the pufferfish (Takifugu rubripes) somatostatin (fSST) and neuropeptide Y (fNPY) genes have been shown to drive inhibitory neuron-specific expression in the mouse brain after intracranial injection. AAV-PHP.B vectors carrying these promoters driving GFP expression are being compared with control vectors in which transgene expression is driven by the ubiquitous strong CAG promoter and the minimal relatively weak mouse MeCP2 promoter. After systemic administration to 6-week-old (tail vein) and postnatal day 1 (postorbital) mice, neonatal CSF delivery, and finally unilateral injection targeting the dentate gyrus (DG) to the CNS, the specificity of AAV-PHP.B-GFP vectors with fSST and fNYP promoters for GABAergic inhibitory interneurons is being investigated (Table 5). The efficiency of CNS gene transfer varies greatly depending on the delivery route, and because Scn1a mice of different ages are being treated, extensive analyses are being performed to establish a baseline for the neuronal transduction efficacy and promoter specificity of each delivery route for GABAergic inhibitory interneurons throughout the CNS. AAV vectors driven by short fSST and fNYP promoters expressing GFP have previously been shown to be highly specific for inhibitory interneurons in the hippocampus after direct injection. The AAV-PHP.B vectors of the present disclosure are being validated in the same manner as subsequent studies in which the restoration of Scn1a + / - mice is being evaluated+ / - Therapeutic impact of Na V 1.1 expression (specifically located in the dentate gyrus and the inner wall of the granular cell layer) (the rationale is clearly expressed below). Experiments are being conducted in 129SvJ / C57BL / 6 mice generated at UMMS by mating 129SvJ with C57BL / 6 mice obtained from Jackson Laboratories (Bar Harbor, ME). Mice are euthanized one month after injection, and the brain and spinal cord are collected for histological analysis of transduction efficiency and specificity using double immunofluorescence with antibodies against cell-specific markers and GFP. Gene transfer efficiency and specificity of GABAergic inhibitory interneurons in the whole brain and spinal cord are being evaluated by double immunofluorescence staining with antibodies against glutamate decarboxylase (GAD; a marker for GABAergic neurons) and GFP. Additionally, the preferential specificity of the promoter and / or AAV-PHP.B for subsets of inhibitory interneurons expressing somatostatin (SST), parvalbumin (PV), calretinin (CR), vasoactive intestinal peptide (VIP), or neuropeptide Y (NPY) is evaluated using antibodies specific for those proteins and GFP. Liver, heart, and skeletal muscle are collected from mice treated by systemic and ICV administration to histologically evaluate GFP expression, and Western blotting is being used to determine the possibility of ectopic expression in peripheral tissues.
[0184] Table 5. Experimental groups
[0185]
[0186] * Each group consists of an equal number of mice from both sexes.
[0187] # Each litter is injected with each vector
[0188] Abbreviations: ICV – intracerebroventricular injection; IC – intracranial injection; PND1 – postnatal day 1
[0189] Six-week-old Scn1a + / - mice are administered bilateral injections of an AAV-PHP.B vector encoding different ZFP Scn1a transactivator proteins (a construct with a ZFP Scn1a activation domain but no DNA-binding domain to control the effect of the transactivator alone) or an equal volume of phosphate-buffered saline (PBS) into the dentate gyrus (n = 3 males + 3 females / group). The single-stranded AAV vectors used in these experiments also carry ZFP Scn1aThe IRES-GFP cassette downstream of the cDNA to facilitate the identification of transduced cells. At least two ZFPs were tested Scn1a transactivators, which may have a broader activation in a variety of neurons, and the above two most promising GABAergic inhibitory neuron-restricted ZFPs SCN1A transactivators. One month after injection, the brains were harvested and the hippocampi from one cerebral hemisphere were dissected to evaluate the ZFPs by Western blotting using β-actin or tubulin as loading controls Scn1a , Na V 1.1, Na V 1.3, GAD65, GAD67 protein expression levels. The other cerebral hemisphere was examined by histological studies using serial brain sections (10 μm) to analyze the percentage of transduced inhibitory interneurons in the granule cell layer of the dentate gyrus and the inner lamina by double immunofluorescence staining with antibodies against GAD and GFP or GAD and epitope tags (HA or myc tags) included in all ZFP Scn1a proteins. In addition, the percentage of GAD-positive neurons expressing Na V 1.1 and Na V 1.3 was determined to demonstrate the restoration of the normal pattern of sodium channel expression. In addition to immunofluorescence detection of Na V 1.1 and Na V 1.3 protein expression, RNAscope probes against Na V 1.1, Na V 1.3, ZFP Scn1a and GAD were used to evaluate changes in mRNA levels in GABAergic interneurons. RNAScope is a highly sensitive in situ hybridization technique for analyzing mRNA levels in neurons within the brain. The combination of these two methods for evaluating changes in Na Scn1a levels induced by ZFP V expression provides a comprehensive understanding of how changes in interneurons are achieved by the gene therapy methods of the present disclosure.
[0190] The therapeutic efficacy of AAV-PHP.B-ZFP + / - gene therapy was analyzed in Scn1a Scn1a mice of both sexes starting via the tail vein at postnatal day 1 or 6 weeks of age. Controls included mice treated with an AAV vector encoding a ZFP-like protein without the ZFP DNA-binding domain; and age-matched untreated Scn1a + / -Mice and wild-type littermates (n = 15 males and 15 females per group). A subset of the mice in each group (n = 3 males and 3 females) was euthanized at 12 weeks of age to evaluate the efficiency of gene transfer to GABAergic interneurons using Western blotting and immunofluorescence with antibodies against GAD (and other neuron type-specific markers, such as GAD65, GAD67) and ZFP, and to restore Na V 1.1 expression in those cells throughout the brain and spinal cord. In addition, ectopic expression of ZFP and Na V 1.1 expression in peripheral tissues were evaluated. Another subset of animals in each group (n = 24) was being used to study the effects on survival (up to 1 year), motor performance, and behavior, and was tested every two months starting at 2 - 12 months of age. Since Scn1a + / - mice exhibit impaired forelimb and hindlimb coordination caused by PND21, motor function and coordination were evaluated using accelerating rotarod and beam crossing tests. Additionally, behavioral tests in which Scn1a + / - mice exhibit impaired performance, including: open field, elevated plus maze, nesting, marble burying, and Barnes maze, were used to test spatial learning and memory that appears to be severely impaired in Scn1a + / - mice. The spontaneous seizure characteristics of patients with Dravet syndrome are also evident in Scn1a + / - mice, and the frequency increases with age and body temperature. In addition, premature sudden death in Scn1a + / - mice occurs immediately after tonic-clonic seizures. Therefore, continuous video monitoring for 24 hours was used at 2 months, 6 months, and 12 months of age to evaluate seizure frequency and duration. If significant changes were detected in the primary outcomes measured in the above tests, social interaction studies using chamber preference readings that respond to novel objects, odors, and mice were considered. Brains, spinal cords, and peripheral organs were collected and evaluated at humane endpoint experiments for the above molecular and histological analyses.
[0191] Example 7. ZFP and dCas9 systems increase the expression of the SCN1A gene in human cells
[0192] To examine the ability of ZFP1 - ZFP3 to upregulate SCN1A transcription, the ZFP1 - ZFP3 DNA-binding domains were fused to the hybrid VP64, p53, and RTA (VPR) triple strong transcriptional activator domain to form chimeric transactivators. The VPR fusion activator domain is used to recruit transcriptional regulatory complexes and increase chromatin accessibility, and contributes to achieving high levels of gene expression. Thus, the ZFP domain targets the VPR activator to highly conserved sequences in the proximal promoter region to increase SCN1A gene expression.
[0193] In addition, to examine the ability of the dCas9 system targeting SCN1A to upregulate SCN1A transcription, three guide RNAs targeting SCN1A were complexed with the dCas9 protein.
[0194] HEK293T cells were transiently transfected with one of the following experimental conditions – (1) VPR-ZFP1 construct; (2) VPR-ZFP2 construct; (3) VPR-ZFP3 construct; (4) all three of the VPR-ZFP1, VPR-ZFP2, and VPR-ZFP3 constructs; (5) dCas9-VPR construct and SCN1A guide RNA 1; (6) dCas9-VPR construct and SCN1A guide RNA 2; (7) dCas9-VPR construct and SCN1A guide RNA 3; (8) dCas9-VPR construct and all three of SCN1A guide RNA 1, SCN1A guide RNA 2, and SCN1A guide RNA 3; and (9) dCas9-VPR construct without any guide RNA (control). SCN1A gene expression was measured by qRT-PCR. Fold activation of SCN1A was normalized to the control experiment (dCas9-VPR construct without any guide RNA).
[0195] Relative to the control experiment, all tested experimental conditions produced an increase in SCN1A gene activation ( Figure 8 ). These data indicate that the zinc finger proteins described in this example and throughout this disclosure are capable of targeting SCN1A to affect gene expression. These data further demonstrate that the guide RNA sequences (SEQ ID NO: 83-94) of this example are capable of targeting dCas9 to SCN1A to affect gene expression.
[0196] Table 6. Guide Nucleic Acids Targeting SCN1A (Spacer Sequences in Bold)
[0197]
[0198] Sequence Listing <110> University of Massachusetts <120> DNA-Binding Domain Transactivators and Their Uses <130> U0120.70106WO00 <140> Not Assigned <141> At the Same Time <150> US 62 / 810,005 <151> 2019-02-25 <160> 121 <170> PatentIn version 3.5 <210> 1 <211> 672 <212> DNA <213> Homo sapiens <400> 1 aatttccatg gactcttttt ccaaaggaat aactggaatg aataaactta aaatcaagat 60 gaaacaatta gatggcttac ctgattaaaa ggaaaattat ccatctgcag tgaggaacag 120 catcacccaa agacgagatg ataacaatgt gccttcagtt gcaattgttc agttccttct 180 tgcaaaaggt gtcaaagtat ttacaagggc tgcagtctca ctggggcaga acacacagac 240 acacaaacac acacaaacgc acacatacac acatgcacca gagacctctg cagtatcctc 300 tcggcttcat cctcgcctca ctctatggta cctaatacaa atcagcaaat agcttgtttc 360 aaaaaaaaaa aaaagtcaag acagcacctt acattacatc gccatctagt ggctaaatat 420 taaacacttt ctcacaatcc agatttatga tttcttcctc aacctctttt ctctcagctt 480 ttttcctttc ttctctgtaa tctcccagta ttgcttctcc ttgcttctct ttcattccct 540 attgctatat aatatcatga acctaatgac tcaaagagga aaaggtttga aagtaaatat 600 agctattttc aagtagtact tgaaaaactt agcattattt tagtttgaaa ctgttacttt 660 attcctaata tg 672 <210> 2 <211> 669 <212> DNA <213> Mus musculus <400> 2 tatttccgtg ggctcttctc cccaaggatt taccaggtaa gaattcacca ccaaagaaga 60 tcacaatgag ataatcagat ggcttacctg ataaaaagga aaattatcca tctgcagtca 120 ggagcaacat ctccccacga cgagtccgca ccttccgttg caacgattca gattccttct 180 tgcaaaaggt gaccaagtgc ttcacaaggg ctgcagcctc ataggggaga acacacgtac 240 acaaacacac gcacacacac acacacatgc accagagacc tctgcagtat cctctggctt 300 catcctcgcc tcactctatg gtacctaata caaatcagca aatagcttgt tttaaaaaaa 360 agaaagaaaa aaagcggaga cagcacctaa cgttacagtg ccatctagtg gctacatcgt 420 aaataggttc tcacagcctg gatttctgtg ttctttctca accgcttcct tctggttcct 480 ttttcttttt tcctctttat tttggtttta ttacttcctc agatgccttt ttttcattcc 540 cctttgctct gcctacatgg aactattgac ttaaagatta aaacaatcag aactggagag 600 cgttgctttt aagttaaaaa aaaaaaggtt gctaattttg tttgtaaatg ttactttatt 660 ttctctatt 669 <210> 3 <211> 130 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 3 tttttttttt tttttttgaa acaagctatt tgctgatttg tattaggtac catagagtga 60 ggcgaggatg aagccgagag gatactgcag aggtctctgg tgcatgtgtg tatgtgtgcg 120 tttgtgtgtg 130 <210> 4 <211> 41 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 4 gagtgaggcg aggatgaagc cgagaggata ctgcagaggt c 41 <210> 5 <211> 18 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 5 gagtgaggcg aggatgaa 18 <210> 6 <211> 18 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 6 ggcgaggatg aagccgag 18 <210> 7 <211> 18 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 7 gaggatactg cagaggtc 18 <210> 8 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 8 Glu Gly Glu Asp Glu 1 5 <210> 9 <211> 6 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 9 Gly Glu Asp Glu Ala Glu 1 5 <210> 10 <211> 6 <212> PRT <213> Artificial sequence <220> <223> Synthesis <400> 10 Glu Asp Thr Ala Glu Val 1 5 <210> 11 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 11 cagcggggaa acctggtgag g 21 <210> 12 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 12 ctgagcttca atctaaccag a 21 <210> 13 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 13 cggagtgaca acttaacgcg g 21 <210> 14 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 14 gaccggtctc accttgcccg a 21 <210> 15 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 15 cagaaggccc atttgactgc c 21 <210> 16 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 16 cggtcggaca acctcacacg c 21 <210> 17 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 17 Gln Arg Gly Asn Leu Val Arg 1 5 <210> 18 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 18 Leu Ser Phe Asn Leu Thr Arg 1 5 <210> 19 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 19 Arg Ser Asp Asn Leu Thr Arg 1 5 <210> 20 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 20 Asp Arg Ser His Leu Ala Arg 1 5 <210> 21 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 21 Gln Lys Ala His Leu Thr Ala 1 5 <210> 22 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 22 Arg Ser Asp Asn Leu Thr Arg 1 5 <210> 23 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 23 cgaagttcca acctgacacg g 21 <210> 24 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 24 gacaagcgga ccttaatccg c 21 <210> 25 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 25 cagcggggaa atctagtgcg a 21 <210> 26 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 26 ctgagcttca acttgactcg t 21 <210> 27 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 27 cggagtgaca atcttacgag a 21 <210> 28 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 28 gaccggagcc acttagccag g 21 <210> 29 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 29 Arg Ser Ser Asn Leu Thr Arg 1 5 <210> 30 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 30 Asp Lys Arg Thr Leu Ile Arg 1 5 <210> 31 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 31 Gln Arg Gly Asn Leu Val Arg 1 5 <210> 32 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 32 Leu Ser Phe Asn Leu Thr Arg 1 5 <210> 33 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 33 Arg Ser Asp Asn Leu Thr Arg 1 5 <210> 34 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 34 Asp Arg Ser His Leu Ala Arg 1 5 <210> 35 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 35 gaccggagcg cgctggcacg g 21 <210> 36 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 36 cgaagtgaca acttaacgcg c 21 <210> 37 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 37 cagtcagggg acctcactcg t 21 <210> 38 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 38 gtacgacaga cgcttaaaca a 21 <210> 39 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 39 gccgctggta acttgacacg a 21 <210> 40 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 40 agatctgata atctaacgcg t 21 <210> 41 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthesis <400> 41 Asp Arg Ser Ala Leu Ala Arg 1 5 <210> 42 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthesis <400> 42 Arg Ser Asp Asn Leu Thr Arg 1 5 <210> 43 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 43 Gln Ser Gly Asp Leu Thr Arg 1 5 <210> 44 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 44 Val Arg Gln Thr Leu Lys Gln 1 5 <210> 45 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 45 Ala Ala Gly Asn Leu Thr Arg 1 5 <210> 46 <211> 7 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 46 Arg Ser Asp Asn Leu Thr Arg 1 5 <210> 47 <211> 1569 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 47 gaggccagcg gttccggacg ggctgacgca ttggacgatt ttgatctgga tatgctggga 60 agtgacgccc tcgatgattt tgaccttgac atgcttggtt cggatgccct tgatgacttt 120 gacctcgaca tgctcggcag tgacgccctt gatgatttcg acctggacat gctgattaac 180 tctagaagtt ccggatctag ccagtacctg cccgacaccg acgaccggca ccggatcgag 240 gaaaagcgga agcggaccta cgagacattc aagagcatca tgaagaagtc ccccttcagc 300 ggccccaccg accctagacc tccacctaga agaatcgccg tgcccagcag atccagcgcc 360 agcgtgccaa aacctgcccc ccagccttac cccttcacca gcagcctgag caccatcaac 420 tacgacgagt tccctaccat ggtgttcccc agcggccaga tctctcaggc ctctgctctg 480 gctccagccc ctcctcaggt gctgcctcag gctcctgctc ctgcaccagc tccagccatg 540 gtgtctgcac tggctcaggc accagcaccc gtgcctgtgc tggctcctgg acctccacag 600 gctgtggctc caccagcccc taaacctaca caggccggcg agggcacact gtctgaagct 660 ctgctgcagc tgcagttcga cgacgaggat ctgggagccc tgctgggaaa cagcaccgat 720 cctgccgtgt tcaccgacct ggccagcgtg gacaacagcg agttccagca gctgctgaac 780 cagggcatcc ctgtggcccc tcacaccacc gagcccatgc tgatggaata ccccgaggcc 840 atcacccggc tcgtgacagg cgctcagagg cctcctgatc cagctcctgc ccctctggga 900 gcaccaggcc tgcctaatgg actgctgtct ggcgacgagg acttcagctc tatcgccgat 960 atggatttct cagccttgct gggctctggc agcggcagcc gggattccag ggaagggatg 1020 tttttgccga agcctgaggc cggctccgct attagtgacg tgtttgaggg ccgcgaggtg 1080 tgccagccaa aacgaatccg gccatttcat cctccaggaa gtccatgggc caaccgccca 1140 ctccccgcca gcctcgcacc aacaccaacc ggtccagtac atgagccagt cgggtcactg 1200 accccggcac cagtccctca gccactggat ccagcgcccg cagtgactcc cgaggccagt 1260 cacctgttgg aggatcccga tgaagagacg agccaggctg tcaaagccct tcgggagatg 1320 gccgatactg tgattcccca gaaggaagag gctgcaatct gtggccaaat ggacctttcc 1380 catccgcccc caaggggcca tctggatgag ctgacaacca cacttgagtc catgaccgag 1440 gatctgaacc tggactcacc cctgaccccg gaattgaacg agattctgga taccttcctg 1500 aacgacgagt gcctcttgca tgccatgcat atcagcacag gactgtccat cttcgacaca 1560 tctctgttt 1569 <210> 48 <211> 523 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 48 Glu Ala Ser Gly Ser Gly Arg Ala Asp Ala Leu Asp Asp Phe Asp Leu 1 5 10 15 Asp Met Leu Gly Ser Asp Ala Leu Asp Asp Phe Asp Leu Asp Met Leu 20 25 30 Gly Ser Asp Ala Leu Asp Asp Phe Asp Leu Asp Met Leu Gly Ser Asp 35 40 45 Ala Leu Asp Asp Phe Asp Leu Asp Met Leu Ile Asn Ser Arg Ser Ser 50 55 60 Gly Ser Ser Gln Tyr Leu Pro Asp Thr Asp Asp Arg His Arg Ile Glu 65 70 75 80 Glu Lys Arg Lys Arg Thr Tyr Glu Thr Phe Lys Ser Ile Met Lys Lys 85 90 95 Ser Pro Phe Ser Gly Pro Thr Asp Pro Arg Pro Pro Pro Arg Arg Ile 100 105 110 Ala Val Pro Ser Arg Ser Ser Ala Ser Val Pro Lys Pro Ala Pro Gln 115 120 125 Pro Tyr Pro Phe Thr Ser Ser Leu Ser Thr Ile Asn Tyr Asp Glu Phe 130 135 140 Pro Thr Met Val Phe Pro Ser Gly Gln Ile Ser Gln Ala Ser Ala Leu 145 150 155 160 Ala Pro Ala Pro Pro Gln Val Leu Pro Gln Ala Pro Ala Pro Ala Pro 165 170 175 Ala Pro Ala Met Val Ser Ala Leu Ala Gln Ala Pro Ala Pro Val Pro 180 185 190 Val Leu Ala Pro Gly Pro Pro Gln Ala Val Ala Pro Pro Ala Pro Lys 195 200 205 Pro Thr Gln Ala Gly Glu Gly Thr Leu Ser Glu Ala Leu Leu Gln Leu 210 215 220 Gln Phe Asp Asp Glu Asp Leu Gly Ala Leu Leu Gly Asn Ser Thr Asp 225 230 235 240 Pro Ala Val Phe Thr Asp Leu Ala Ser Val Asp Asn Ser Glu Phe Gln 245 250 255 Gln Leu Leu Asn Gln Gly Ile Pro Val Ala Pro His Thr Thr Glu Pro 260 265 270 Met Leu Met Glu Tyr Pro Glu Ala Ile Thr Arg Leu Val Thr Gly Ala 275 280 285 Gln Arg Pro Pro Asp Pro Ala Pro Ala Pro Leu Gly Ala Pro Gly Leu 290 295 300 Pro Asn Gly Leu Leu Ser Gly Asp Glu Asp Phe Ser Ser Ile Ala Asp 305 310 315 320 Met Asp Phe Ser Ala Leu Leu Gly Ser Gly Ser Gly Ser Arg Asp Ser 325 330 335 Arg Glu Gly Met Phe Leu Pro Lys Pro Glu Ala Gly Ser Ala Ile Ser 340 345 350 Asp Val Phe Glu Gly Arg Glu Val Cys Gln Pro Lys Arg Ile Arg Pro 355 360 365 Phe His Pro Pro Gly Ser Pro Trp Ala Asn Arg Pro Leu Pro Ala Ser 370 375 380 Leu Ala Pro Thr Pro Thr Gly Pro Val His Glu Pro Val Gly Ser Leu 385 390 395 400 Thr Pro Ala Pro Val Pro Gln Pro Leu Asp Pro Ala Pro Ala Val Thr 405 410 415 Pro Glu Ala Ser His Leu Leu Glu Asp Pro Asp Glu Glu Thr Ser Gln 420 425 430 Ala Val Lys Ala Leu Arg Glu Met Ala Asp Thr Val Ile Pro Gln Lys 435 440 445 Glu Glu Ala Ala Ile Cys Gly Gln Met Asp Leu Ser His Pro Pro Pro 450 455 460 Arg Gly His Leu Asp Glu Leu Thr Thr Thr Leu Glu Ser Met Thr Glu 465 470 475 480 Asp Leu Asn Leu Asp Ser Pro Leu Thr Pro Glu Leu Asn Glu Ile Leu 485 490 495 Asp Thr Phe Leu Asn Asp Glu Cys Leu Leu His Ala Met His Ile Ser 500 505 510 Thr Gly Leu Ser Ile Phe Asp Thr Ser Leu Phe 515 520 <210> 49 <211> 6027 <212> DNA <213> Homo sapiens <400> 49 atggaacaga ccgtgctggt gccgccgggc ccggatagct ttaacttttt tacccgcgaa 60 agcctggcgg cgattgaacg ccgcattgcg gaagaaaaag cgaaaaaccc gaaaccggat 120 aaaaaagatg atgatgaaaa cggcccgaaa ccgaacagcg atctggaagc gggcaaaaac 180 ctgccgttta tttatggcga tattccgccg gaaatggtga gcgaaccgct ggaagatctg 240 gatccgtatt atattaacaa aaaaaccttt attgtgctga acaaaggcaa agcgattttt 300 cgctttagcg cgaccagcgc gctgtatatt ctgaccccgt ttaacccgct gcgcaaaatt 360 gcgattaaaa ttctggtgca tagcctgttt agcatgctga ttatgtgcac cattctgacc 420 aactgcgtgt ttatgaccat gagcaacccg ccggattgga ccaaaaacgt ggaatatacc 480 tttaccggca tttatacctt tgaaagcctg attaaaatta ttgcgcgcgg cttttgcctg 540 gaagatttta cctttctgcg cgatccgtgg aactggctgg attttaccgt gattaccttt 600 gcgtatgtga ccgaatttgt ggatctgggc aacgtgagcg cgctgcgcac ctttcgcgtg 660 ctgcgcgcgc tgaaaaccat tagcgtgatt ccgggcctga aaaccattgt gggcgcgctg 720 attcagagcg tgaaaaaact gagcgatgtg atgattctga ccgtgttttg cctgagcgtg 780 tttgcgctga ttggcctgca gctgtttatg ggcaacctgc gcaacaaatg cattcagtgg 840 ccgccgacca acgcgagcct ggaagaacat agcattgaaa aaaacattac cgtgaactat 900 aacggcaccc tgattaacga aaccgtgttt gaatttgatt ggaaaagcta tattcaggat 960 agccgctatc attattttct ggaaggcttt ctggatgcgc tgctgtgcgg caacagcagc 1020 gatgcgggcc agtgcccgga aggctatatg tgcgtgaaag cgggccgcaa cccgaactat 1080 ggctatacca gctttgatac ctttagctgg gcgtttctga gcctgtttcg cctgatgacc 1140 caggattttt gggaaaacct gtatcagctg accctgcgcg cggcgggcaa aacctatatg 1200 attttttttg tgctggtgat ttttctgggc agcttttatc tgattaacct gattctggcg 1260 gtggtggcga tggcgtatga agaacagaac caggcgaccc tggaagaagc ggaacagaaa 1320 gaagcggaat ttcagcagat gattgaacag ctgaaaaaac agcaggaagc ggcgcagcag 1380 gcggcgaccg cgaccgcgag cgaacatagc cgcgaaccga gcgcggcggg ccgcctgagc 1440 gatagcagca gcgaagcgag caaactgagc agcaaaagcg cgaaagaacg ccgcaaccgc 1500 cgcaaaaaac gcaaacagaa agaacagagc ggcggcgaag aaaaagatga agatgaattt 1560 cgcaaaaaac gcaaacagaa agaacagagc ggcggcgaag aaaaagatga agatgaattt 1560 cagaaaagcg aaagcgaaga tagcattcgc cgcaaaggct ttcgctttag cattgaaggc 1620 cagaaaagcg aaagcgaaga tagcattcgc cgcaaaggct ttcgctttag cattgaaggc 1620 aaccgcctga cctatgaaaa acgctatagc agcccgcatc agagcctgct gagcattcgc 1680 aaccgcctga cctatgaaaa acgctatagc agcccgcatc agagcctgct gagcattcgc 1680 ggcagcctgt ttagcccgcg ccgcaacagc cgcaccagcc tgtttagctt tcgcggccgc 1740 ggcagcctgt ttagcccgcg ccgcaacagc cgcaccagcc tgtttagctt tcgcggccgc 1740 gcgaaagatg tgggcagcga aaacgatttt gcggatgatg aacatagcac ctttgaagat 1800 gcgaaagatg tgggcagcga aaacgatttt gcggatgatg aacatagcac ctttgaagat 1800 aacgaaagcc gccgcgatag cctgtttgtg ccgcgccgcc atggcgaacg ccgcaacagc 1860 aacgaaagcc gccgcgatag cctgtttgtg ccgcgccgcc atggcgaacg ccgcaacagc 1860 aacctgagcc agaccagccg cagcagccgc atgctggcgg tgtttccggc gaacggcaaa 1920 aacctgagcc agaccagccg cagcagccgc atgctggcgg tgtttccggc gaacggcaaa 1920 atgcatagca ccgtggattg caacggcgtg gtgagcctgg tgggcggccc gagcgtgccg 1980 atgcatagca ccgtggattg caacggcgtg gtgagcctgg tgggcggccc gagcgtgccg 1980 accagcccgg tgggccagct gctgccggaa gtgattattg ataaaccggc gaccgatgat 2040 accagcccgg tgggccagct gctgccggaa gtgattattg ataaaccggc gaccgatgat 2040 aacggcacca ccaccgaaac cgaaatgcgc aaacgccgca gcagcagctt tcatgtgagc 2100 aacggcacca ccaccgaaac cgaaatgcgc aaacgccgca gcagcagctt tcatgtgagc 2100 atggattttc tggaagatcc gagccagcgc cagcgcgcga tgagcattgc gagcattctg 2160 atggattttc tggaagatcc gagccagcgc cagcgcgcga tgagcattgc gagcattctg 2160 accaacaccg tggaagaact ggaagaaagc cgccagaaat gcccgccgtg ctggtataaa 2220 accaacaccg tggaagaact ggaagaaagc cgccagaaat gcccgccgtg ctggtataaa 2220 tttagcaaca tttttctgat ttgggattgc agcccgtatt ggctgaaagt gaaacatgtg 2280 tttagcaaca tttttctgat ttgggattgc agcccgtatt ggctgaaagt gaaacatgtg 2280 gtgaacctgg tggtgatgga tccgtttgtg gatctggcga ttaccatttg cattgtgctg 2340 gtgaacctgg tggtgatgga tccgtttgtg gatctggcga ttaccatttg cattgtgctg 2340 aacaccctgt ttatggcgat ggaacattat ccgatgaccg atcattttaa caacgtgctg 2400 aacaccctgt ttatggcgat ggaacattat ccgatgaccg atcattttaa caacgtgctg 2400 accgtgggca acctggtgtt taccggcatt tttaccgcgg aaatgtttct gaaaattatt 2460 accgtgggca acctggtgtt taccggcatt tttaccgcgg aaatgtttct gaaaattatt 2460 gcgatggatc cgtattatta ttttcaggaa ggctggaaca tttttgatgg ctttattgtg 2520 gcgatggatc cgtattatta ttttcaggaa ggctggaaca tttttgatgg ctttattgtg 2520 accctgagcc tggtggaact gggcctggcg aacgtggaag gcctgagcgt gctgcgcagc 2580 accctgagcc tggtggaact gggcctggcg aacgtggaag gcctgagcgt gctgcgcagc 2580 tttcgcctgc tgcgcgtgtt taaactggcg aaaagctggc cgaccctgaa catgctgatt 2640 tttcgcctgc tgcgcgtgtt taaactggcg aaaagctggc cgaccctgaa catgctgatt 2640 aaaattattg gcaacagcgt gggcgcgctg ggcaacctga ccctggtgct ggcgattatt 2700 aaaattattg gcaacagcgt gggcgcgctg ggcaacctga ccctggtgct ggcgattatt 2700 gtgtttattt ttgcggtggt gggcatgcag ctgtttggca aaagctataa agattgcgtg 2760 gtgtttattt ttgcggtggt gggcatgcag ctgtttggca aaagctataa agattgcgtg 2760 tgcaaaattg cgagcgattg ccagctgccg cgctggcata tgaacgattt ttttcatagc 2820 tgcaaaattg cgagcgattg ccagctgccg cgctggcata tgaacgattt ttttcatagc 2820 tttctgattg tgtttcgcgt gctgtgcggc gaatggattg aaaccatgtg ggattgcatg 2880 tttctgattg tgtttcgcgt gctgtgcggc gaatggattg aaaccatgtg ggattgcatg 2880 gaagtggcgg gccaggcgat gtgcctgacc gtgtttatga tggtgatggt gattggcaac 2940 gaagtggcgg gccaggcgat gtgcctgacc gtgtttatga tggtgatggt gattggcaac 2940 ctggtggtgc tgaacctgtt tctggcgctg ctgctgagca gctttagcgc ggataacctg 3000 gcggcgaccg atgatgataa cgaaatgaac aacctgcaga ttgcggtgga tcgcatgcat 3060 aaaggcgtgg cgtatgtgaa acgcaaaatt tatgaattta ttcagcagag ctttattcgc 3120 aaacagaaaa ttctggatga aattaaaccg ctggatgatc tgaacaacaa aaaagatagc 3180 tgcatgagca accataccgc ggaaattggc aaagatctgg attatctgaa agatgtgaac 3240 ggcaccacca gcggcattgg caccggcagc agcgtggaaa aatatattat tgatgaaagc 3300 gattatatga gctttattaa caacccgagc ctgaccgtga ccgtgccgat tgcggtgggc 3360 gaaagcgatt ttgaaaacct gaacaccgaa gattttagca gcgaaagcga tctggaagaa 3420 agcaaagaaa aactgaacga aagcagcagc agcagcgaag gcagcaccgt ggatattggc 3480 gcgccggtgg aagaacagcc ggtggtggaa ccggaagaaa ccctggaacc ggaagcgtgc 3540 tttaccgaag gctgcgtgca gcgctttaaa tgctgccaga ttaacgtgga agaaggccgc 3600 ggcaaacagt ggtggaacct gcgccgcacc tgctttcgca ttgtggaaca taactggttt 3660 gaaaccttta ttgtgtttat gattctgctg agcagcggcg cgctggcgtt tgaagatatt 3720 tatattgatc agcgcaaaac cattaaaacc atgctggaat atgcggataa agtgtttacc 3780 tatattttta ttctggaaat gctgctgaaa tgggtggcgt atggctatca gacctatttt 3840 accaacgcgt ggtgctggct ggattttctg attgtggatg tgagcctggt gagcctgacc 3900 gcgaacgcgc tgggctatag cgaactgggc gcgattaaaa gcctgcgcac cctgcgcgcg 3960 ctgcgcccgc tgcgcgcgct gagccgcttt gaaggcatgc gcgtggtggt gaacgcgctg 4020 ctgggcgcga ttccgagcat tatgaacgtg ctgctggtgt gcctgatttt ttggctgatt 4080 tttagcatta tgggcgtgaa cctgtttgcg ggcaaatttt atcattgcat taacaccacc 4140 accggcgatc gctttgatat tgaagatgtg aacaaccata ccgattgcct gaaactgatt 4200 gaacgcaacg aaaccgcgcg ctggaaaaac gtgaaagtga actttgataa cgtgggcttt 4260 ggctatctga gcctgctgca ggtggcgacc tttaaaggct ggatggatat tatgtatgcg 4320 gcggtggata gccgcaacgt ggaactgcag ccgaaatatg aagaaagcct gtatatgtat 4380 ctgtattttg tgatttttat tatttttggc agctttttta ccctgaacct gtttattggc 4440 gtgattattg ataactttaa ccagcagaaa aaaaaatttg gcggccagga tatttttatg 4500 accgaagaac agaaaaaata ttataacgcg atgaaaaaac tgggcagcaa aaaaccgcag 4560 aaaccgattc cgcgcccggg caacaaattt cagggcatgg tgtttgattt tgtgacccgc 4620 caggtgtttg atattagcat tatgattctg atttgcctga acatggtgac catgatggtg 4680 gaaaccgatg atcagagcga atatgtgacc accattctga gccgcattaa cctggtgttt 4740 attgtgctgt ttaccggcga atgcgtgctg aaactgatta gcctgcgcca ttattatttt 4800 accattggct ggaacatttt tgattttgtg gtggtgattc tgagcattgt gggcatgttt 4860 ctggcggaac tgattgaaaa atattttgtg agcccgaccc tgtttcgcgt gattcgcctg 4920 gcgcgcattg gccgcattct gcgcctgatt aaaggcgcga aaggcattcg caccctgctg 4980 tttgcgctga tgatgagcct gccggcgctg tttaacattg gcctgctgct gtttctggtg 5040 atgtttattt atgcgatttt tggcatgagc aactttgcgt atgtgaaacg cgaagtgggc 5100 attgatgata tgtttaactt tgaaaccttt ggcaacagca tgatttgcct gtttcagatt 5160 accaccagcg cgggctggga tggcctgctg gcgccgattc tgaacagcaa accgccggat 5220 tgcgatccga acaaagtgaa cccgggcagc agcgtgaaag gcgattgcgg caacccgagc 5280 gtgggcattt ttttttttgt gagctatatt attattagct ttctggtggt ggtgaacatg 5340 tatattgcgg tgattctgga aaactttagc gtggcgaccg aagaaagcgc ggaaccgctg 5400 agcgaagatg attttgaaat gttttatgaa gtgtgggaaa aatttgatcc ggatgcgacc 5460 cagtttatgg aatttgaaaa actgagccag tttgcggcgg cgctggaacc gccgctgaac 5520 ctgccgcagc cgaacaaact gcagctgatt gcgatggatc tgccgatggt gagcggcgat 5580 cgcattcatt gcctggatat tctgtttgcg tttaccaaac gcgtgctggg cgaaagcggc 5640 gaaatggatg cgctgcgcat tcagatggaa gaacgcttta tggcgagcaa cccgagcaaa 5700 gtgagctatc agccgattac caccaccctg aaacgcaaac aggaagaagt gagcgcggtg 5760 attattcagc gcgcgtatcg ccgccatctg ctgaaacgca ccgtgaaaca ggcgagcttt 5820 acctataaca aaaacaaaat taaaggcggc gcgaacctgc tgattaaaga agatatgatt 5880 attgatcgca ttaacgaaaa cagcattacc gaaaaaaccg atctgaccat gagcaccgcg 5940 gcgtgcccgc cgagctatga tcgcgtgacc aaaccgattg tggaaaaaca tgaacaggaa 6000 ggcaaagatg aaaaagcgaa aggcaaa 6027 <210> 50 <211> 2009 <212> PRT <213> Homo sapiens <400> 50 Met Glu Gln Thr Val Leu Val Pro Pro Gly Pro Asp Ser Phe Asn Phe 1 5 10 15 Phe Thr Arg Glu Ser Leu Ala Ala Ile Glu Arg Arg Ile Ala Glu Glu 20 25 30 Lys Ala Lys Asn Pro Lys Pro Asp Lys Lys Asp Asp Asp Glu Asn Gly 35 40 45 Pro Lys Pro Asn Ser Asp Leu Glu Ala Gly Lys Asn Leu Pro Phe Ile 50 55 60 Tyr Gly Asp Ile Pro Pro Glu Met Val Ser Glu Pro Leu Glu Asp Leu 65 70 75 80 Asp Pro Tyr Tyr Ile Asn Lys Lys Thr Phe Ile Val Leu Asn Lys Gly 85 90 95 Lys Ala Ile Phe Arg Phe Ser Ala Thr Ser Ala Leu Tyr Ile Leu Thr 100 105 110 Pro Phe Asn Pro Leu Arg Lys Ile Ala Ile Lys Ile Leu Val His Ser 115 120 125 Leu Phe Ser Met Leu Ile Met Cys Thr Ile Leu Thr Asn Cys Val Phe 130 135 140 Met Thr Met Ser Asn Pro Pro Asp Trp Thr Lys Asn Val Glu Tyr Thr 145 150 155 160 Phe Thr Gly Ile Tyr Thr Phe Glu Ser Leu Ile Lys Ile Ile Ala Arg 165 170 175 Gly Phe Cys Leu Glu Asp Phe Thr Phe Leu Arg Asp Pro Trp Asn Trp 180 185 190 Leu Asp Phe Thr Val Ile Thr Phe Ala Tyr Val Thr Glu Phe Val Asp 195 200 205 Leu Gly Asn Val Ser Ala Leu Arg Thr Phe Arg Val Leu Arg Ala Leu 210 215 220 Lys Thr Ile Ser Val Ile Pro Gly Leu Lys Thr Ile Val Gly Ala Leu 225 230 235 240 Ile Gln Ser Val Lys Lys Leu Ser Asp Val Met Ile Leu Thr Val Phe 245 250 255 Cys Leu Ser Val Phe Ala Leu Ile Gly Leu Gln Leu Phe Met Gly Asn 260 265 270 Leu Arg Asn Lys Cys Ile Gln Trp Pro Pro Thr Asn Ala Ser Leu Glu 275 280 285 Glu His Ser Ile Glu Lys Asn Ile Thr Val Asn Tyr Asn Gly Thr Leu 290 295 300 Ile Asn Glu Thr Val Phe Glu Phe Asp Trp Lys Ser Tyr Ile Gln Asp 305 310 315 320 Ser Arg Tyr His Tyr Phe Leu Glu Gly Phe Leu Asp Ala Leu Leu Cys 325 330 335 Gly Asn Ser Ser Asp Ala Gly Gln Cys Pro Glu Gly Tyr Met Cys Val 340 345 350 Lys Ala Gly Arg Asn Pro Asn Tyr Gly Tyr Thr Ser Phe Asp Thr Phe 355 360 365 Ser Trp Ala Phe Leu Ser Leu Phe Arg Leu Met Thr Gln Asp Phe Trp 370 375 380 Glu Asn Leu Tyr Gln Leu Thr Leu Arg Ala Ala Gly Lys Thr Tyr Met 385 390 395 400 Ile Phe Phe Val Leu Val Ile Phe Leu Gly Ser Phe Tyr Leu Ile Asn 405 410 415 Leu Ile Leu Ala Val Val Ala Met Ala Tyr Glu Glu Gln Asn Gln Ala 420 425 430 Thr Leu Glu Glu Ala Glu Gln Lys Glu Ala Glu Phe Gln Gln Met Ile 435 440 445 Glu Gln Leu Lys Lys Gln Gln Glu Ala Ala Gln Gln Ala Ala Thr Ala 450 455 460 Thr Ala Ser Glu His Ser Arg Glu Pro Ser Ala Ala Gly Arg Leu Ser 465 470 475 480 Asp Ser Ser Ser Glu Ala Ser Lys Leu Ser Ser Lys Ser Ala Lys Glu 485 490 495 Arg Arg Asn Arg Arg Lys Lys Arg Lys Gln Lys Glu Gln Ser Gly Gly 500 505 510 Glu Glu Lys Asp Glu Asp Glu Phe Gln Lys Ser Glu Ser Glu Asp Ser 515 520 525 Ile Arg Arg Lys Gly Phe Arg Phe Ser Ile Glu Gly Asn Arg Leu Thr 530 535 540 Tyr Glu Lys Arg Tyr Ser Ser Pro His Gln Ser Leu Leu Ser Ile Arg 545 550 555 560 Gly Ser Leu Phe Ser Pro Arg Arg Asn Ser Arg Thr Ser Leu Phe Ser 565 570 575 Phe Arg Gly Arg Ala Lys Asp Val Gly Ser Glu Asn Asp Phe Ala Asp 580 585 590 Asp Glu His Ser Thr Phe Glu Asp Asn Glu Ser Arg Arg Asp Ser Leu 595 600 605 Phe Val Pro Arg Arg His Gly Glu Arg Arg Asn Ser Asn Leu Ser Gln 610 615 620 Thr Ser Arg Ser Ser Arg Met Leu Ala Val Phe Pro Ala Asn Gly Lys 625 630 635 640 Met His Ser Thr Val Asp Cys Asn Gly Val Val Ser Leu Val Gly Gly 645 650 655 Pro Ser Val Pro Thr Ser Pro Val Gly Gln Leu Leu Pro Glu Val Ile 660 665 670 Ile Asp Lys Pro Ala Thr Asp Asp Asn Gly Thr Thr Thr Glu Thr Glu 675 680 685 Met Arg Lys Arg Arg Ser Ser Ser Phe His Val Ser Met Asp Phe Leu 690 695 700 Glu Asp Pro Ser Gln Arg Gln Arg Ala Met Ser Ile Ala Ser Ile Leu 705 710 715 720 Thr Asn Thr Val Glu Glu Leu Glu Glu Ser Arg Gln Lys Cys Pro Pro 725 730 735 Cys Trp Tyr Lys Phe Ser Asn Ile Phe Leu Ile Trp Asp Cys Ser Pro 740 745 750 Tyr Trp Leu Lys Val Lys His Val Val Asn Leu Val Val Met Asp Pro 755 760 765 Phe Val Asp Leu Ala Ile Thr Ile Cys Ile Val Leu Asn Thr Leu Phe 770 775 780 Met Ala Met Glu His Tyr Pro Met Thr Asp His Phe Asn Asn Val Leu 785 790 795 800 Thr Val Gly Asn Leu Val Phe Thr Gly Ile Phe Thr Ala Glu Met Phe 805 810 815 Leu Lys Ile Ile Ala Met Asp Pro Tyr Tyr Tyr Phe Gln Glu Gly Trp 820 825 830 Asn Ile Phe Asp Gly Phe Ile Val Thr Leu Ser Leu Val Glu Leu Gly 835 840 845 Leu Ala Asn Val Glu Gly Leu Ser Val Leu Arg Ser Phe Arg Leu Leu 850 855 860 Arg Val Phe Lys Leu Ala Lys Ser Trp Pro Thr Leu Asn Met Leu Ile 865 870 875 880 Lys Ile Ile Gly Asn Ser Val Gly Ala Leu Gly Asn Leu Thr Leu Val 885 890 895 Leu Ala Ile Ile Val Phe Ile Phe Ala Val Val Gly Met Gln Leu Phe 900 905 910 Gly Lys Ser Tyr Lys Asp Cys Val Cys Lys Ile Ala Ser Asp Cys Gln 915 920 925 Leu Pro Arg Trp His Met Asn Asp Phe Phe His Ser Phe Leu Ile Val 930 935 940 Phe Arg Val Leu Cys Gly Glu Trp Ile Glu Thr Met Trp Asp Cys Met 945 950 955 960 Glu Val Ala Gly Gln Ala Met Cys Leu Thr Val Phe Met Met Val Met 965 970 975 Val Ile Gly Asn Leu Val Val Leu Asn Leu Phe Leu Ala Leu Leu Leu 980 985 990 Ser Ser Phe Ser Ala Asp Asn Leu Ala Ala Thr Asp Asp Asp Asn Glu 995 1000 1005 Met Asn Asn Leu Gln Ile Ala Val Asp Arg Met His Lys Gly Val 1010 1015 1020 Ala Tyr Val Lys Arg Lys Ile Tyr Glu Phe Ile Gln Gln Ser Phe 1025 1030 1035 Ile Arg Lys Gln Lys Ile Leu Asp Glu Ile Lys Pro Leu Asp Asp 1040 1045 1050 Leu Asn Asn Lys Lys Asp Ser Cys Met Ser Asn His Thr Ala Glu 1055 1060 1065 Ile Gly Lys Asp Leu Asp Tyr Leu Lys Asp Val Asn Gly Thr Thr 1070 1075 1080 Ser Gly Ile Gly Thr Gly Ser Ser Val Glu Lys Tyr Ile Ile Asp 1085 1090 1095 Glu Ser Asp Tyr Met Ser Phe Ile Asn Asn Pro Ser Leu Thr Val 1100 1105 1110 Thr Val Pro Ile Ala Val Gly Glu Ser Asp Phe Glu Asn Leu Asn 1115 1120 1125 Thr Glu Asp Phe Ser Ser Glu Ser Asp Leu Glu Glu Ser Lys Glu 1130 1135 1140 Lys Leu Asn Glu Ser Ser Ser Ser Ser Glu Gly Ser Thr Val Asp 1145 1150 1155 Ile Gly Ala Pro Val Glu Glu Gln Pro Val Val Glu Pro Glu Glu 1160 1165 1170 Thr Leu Glu Pro Glu Ala Cys Phe Thr Glu Gly Cys Val Gln Arg 1175 1180 1185 Phe Lys Cys Cys Gln Ile Asn Val Glu Glu Gly Arg Gly Lys Gln 1190 1195 1200 Trp Trp Asn Leu Arg Arg Thr Cys Phe Arg Ile Val Glu His Asn 1205 1210 1215 Trp Phe Glu Thr Phe Ile Val Phe Met Ile Leu Leu Ser Ser Gly 1220 1225 1230 Ala Leu Ala Phe Glu Asp Ile Tyr Ile Asp Gln Arg Lys Thr Ile 1235 1240 1245 Lys Thr Met Leu Glu Tyr Ala Asp Lys Val Phe Thr Tyr Ile Phe 1250 1255 1260 Ile Leu Glu Met Leu Leu Lys Trp Val Ala Tyr Gly Tyr Gln Thr 1265 1270 1275 Tyr Phe Thr Asn Ala Trp Cys Trp Leu Asp Phe Leu Ile Val Asp 1280 1285 1290 Val Ser Leu Val Ser Leu Thr Ala Asn Ala Leu Gly Tyr Ser Glu 1295 1300 1305 Leu Gly Ala Ile Lys Ser Leu Arg Thr Leu Arg Ala Leu Arg Pro 1310 1315 1320 Leu Arg Ala Leu Ser Arg Phe Glu Gly Met Arg Val Val Val Asn 1325 1330 1335 Ala Leu Leu Gly Ala Ile Pro Ser Ile Met Asn Val Leu Leu Val 1340 1345 1350 Cys Leu Ile Phe Trp Leu Ile Phe Ser Ile Met Gly Val Asn Leu 1355 1360 1365 Phe Ala Gly Lys Phe Tyr His Cys Ile Asn Thr Thr Thr Gly Asp 1370 1375 1380 Arg Phe Asp Ile Glu Asp Val Asn Asn His Thr Asp Cys Leu Lys 1385 1390 1395 Leu Ile Glu Arg Asn Glu Thr Ala Arg Trp Lys Asn Val Lys Val 1400 1405 1410 Asn Phe Asp Asn Val Gly Phe Gly Tyr Leu Ser Leu Leu Gln Val 1415 1420 1425 Ala Thr Phe Lys Gly Trp Met Asp Ile Met Tyr Ala Ala Val Asp 1430 1435 1440 Ser Arg Asn Val Glu Leu Gln Pro Lys Tyr Glu Glu Ser Leu Tyr 1445 1450 1455 Met Tyr Leu Tyr Phe Val Ile Phe Ile Ile Phe Gly Ser Phe Phe 1460 1465 1470 Thr Leu Asn Leu Phe Ile Gly Val Ile Ile Asp Asn Phe Asn Gln 1475 1480 1485 Gln Lys Lys Lys Phe Gly Gly Gln Asp Ile Phe Met Thr Glu Glu 1490 1495 1500 Gln Lys Lys Tyr Tyr Asn Ala Met Lys Lys Leu Gly Ser Lys Lys 1505 1510 1515 Pro Gln Lys Pro Ile Pro Arg Pro Gly Asn Lys Phe Gln Gly Met 1520 1525 1530 Val Phe Asp Phe Val Thr Arg Gln Val Phe Asp Ile Ser Ile Met 1535 1540 1545 Ile Leu Ile Cys Leu Asn Met Val Thr Met Met Val Glu Thr Asp 1550 1555 1560 Asp Gln Ser Glu Tyr Val Thr Thr Ile Leu Ser Arg Ile Asn Leu 1565 1570 1575 Val Phe Ile Val Leu Phe Thr Gly Glu Cys Val Leu Lys Leu Ile 1580 1585 1590 Ser Leu Arg His Tyr Tyr Phe Thr Ile Gly Trp Asn Ile Phe Asp 1595 1600 1605 Phe Val Val Val Ile Leu Ser Ile Val Gly Met Phe Leu Ala Glu 1610 1615 1620 Leu Ile Glu Lys Tyr Phe Val Ser Pro Thr Leu Phe Arg Val Ile 1625 1630 1635 Arg Leu Ala Arg Ile Gly Arg Ile Leu Arg Leu Ile Lys Gly Ala 1640 1645 1650 Lys Gly Ile Arg Thr Leu Leu Phe Ala Leu Met Met Ser Leu Pro 1655 1660 1665 Ala Leu Phe Asn Ile Gly Leu Leu Leu Phe Leu Val Met Phe Ile 1670 1675 1680 Tyr Ala Ile Phe Gly Met Ser Asn Phe Ala Tyr Val Lys Arg Glu 1685 1690 1695 Val Gly Ile Asp Asp Met Phe Asn Phe Glu Thr Phe Gly Asn Ser 1700 1705 1710 Met Ile Cys Leu Phe Gln Ile Thr Thr Ser Ala Gly Trp Asp Gly 1715 1720 1725 Leu Leu Ala Pro Ile Leu Asn Ser Lys Pro Pro Asp Cys Asp Pro 1730 1735 1740 Asn Lys Val Asn Pro Gly Ser Ser Val Lys Gly Asp Cys Gly Asn 1745 1750 1755 Pro Ser Val Gly Ile Phe Phe Phe Val Ser Tyr Ile Ile Ile Ser 1760 1765 1770 Phe Leu Val Val Val Asn Met Tyr Ile Ala Val Ile Leu Glu Asn 1775 1780 1785 Phe Ser Val Ala Thr Glu Glu Ser Ala Glu Pro Leu Ser Glu Asp 1790 1795 1800 Asp Phe Glu Met Phe Tyr Glu Val Trp Glu Lys Phe Asp Pro Asp 1805 1810 1815 Ala Thr Gln Phe Met Glu Phe Glu Lys Leu Ser Gln Phe Ala Ala 1820 1825 1830 Ala Leu Glu Pro Pro Leu Asn Leu Pro Gln Pro Asn Lys Leu Gln 1835 1840 1845 Leu Ile Ala Met Asp Leu Pro Met Val Ser Gly Asp Arg Ile His 1850 1855 1860 Cys Leu Asp Ile Leu Phe Ala Phe Thr Lys Arg Val Leu Gly Glu 1865 1870 1875 Ser Gly Glu Met Asp Ala Leu Arg Ile Gln Met Glu Glu Arg Phe 1880 1885 1890 Met Ala Ser Asn Pro Ser Lys Val Ser Tyr Gln Pro Ile Thr Thr 1895 1900 1905 Thr Leu Lys Arg Lys Gln Glu Glu Val Ser Ala Val Ile Ile Gln 1910 1915 1920 Arg Ala Tyr Arg Arg His Leu Leu Lys Arg Thr Val Lys Gln Ala 1925 1930 1935 Ser Phe Thr Tyr Asn Lys Asn Lys Ile Lys Gly Gly Ala Asn Leu 1940 1945 1950 Leu Ile Lys Glu Asp Met Ile Ile Asp Arg Ile Asn Glu Asn Ser 1955 1960 1965 Ile Thr Glu Lys Thr Asp Leu Thr Met Ser Thr Ala Ala Cys Pro 1970 1975 1980 Pro Ser Tyr Asp Arg Val Thr Lys Pro Ile Val Glu Lys His Glu 1985 1990 1995 Gln Glu Gly Lys Asp Glu Lys Ala Lys Gly Lys 2000 2005 <210> 51 <211> 1470 <212> DNA <213> Homo sapiens <400> 51 atggatctgc tggtggatga actgtttgcg gatatgaacg cggatggcgc gagcccgccg 60 ccgccgcgcc cggcgggcgg cccgaaaaac accccggcgg cgccgccgct gtatgcgacc 120 ggccgcctga gccaggcgca gctgatgccg agcccgccga tgccggtgcc gccggcggcg 180 ctgtttaacc gcctgctgga tgatctgggc tttagcgcgg gcccggcgct gtgcaccatg 240 ctggatacct ggaacgaaga tctgtttagc gcgctgccga ccaacgcgga tctgtatcgc 300 gaatgcaaat ttctgagcac cctgccgagc gatgtggtgg aatggggcga tgcgtatgtg 360 ccggaacgca cccagattga tattcgcgcg catggcgatg tggcgtttcc gaccctgccg 420 gcgacccgcg atggcctggg cctgtattat gaagcgctga gccgcttttt tcatgcggaa 480 ctgcgcgcgc gcgaagaaag ctatcgcacc gtgctggcga acttttgcag cgcgctgtat 540 cgctatctgc gcgcgagcgt gcgccagctg catcgccagg cgcatatgcg cggccgcgat 600 cgcgatctgg gcgaaatgct gcgcgcgacc attgcggatc gctattatcg cgaaaccgcg 660 cgcctggcgc gcgtgctgtt tctgcatctg tatctgtttc tgacccgcga aattctgtgg 720 gcggcgtatg cggaacagat gatgcgcccg gatctgtttg attgcctgtg ctgcgatctg 780 gaaagctggc gccagctggc gggcctgttt cagccgttta tgtttgtgaa cggcgcgctg 840 accgtgcgcg gcgtgccgat tgaagcgcgc cgcctgcgcg aactgaacca tattcgcgaa 900 catctgaacc tgccgctggt gcgcagcgcg gcgaccgaag aaccgggcgc gccgctgacc 960 accccgccga ccctgcatgg caaccaggcg cgcgcgagcg gctattttat ggtgctgatt 1020 cgcgcgaaac tggatagcta tagcagcttt accaccagcc cgagcgaagc ggtgatgcgc 1080 gaacatgcgt atagccgcgc gcgcaccaaa aacaactatg gcagcaccat tgaaggcctg 1140 ctggatctgc cggatgatga tgcgccggaa gaagcgggcc tggcggcgcc gcgcctgagc 1200 tttctgccgg cgggccatac ccgccgcctg agcaccgcgc cgccgaccga tgtgagcctg 1260 ggcgatgaac tgcatctgga tggcgaagat gtggcgatgg cgcatgcgga tgcgctggat 1320 gattttgatc tggatatgct gggcgatggc gatagcccgg gcccgggctt taccccgcat 1380 gatagcgcgc cgtatggcgc gctggatatg gcggattttg aatttgaaca gatgtttacc 1440 gatgcgctgg gcattgatga atatggcggc 1470 <210> 52 <211> 490 <212> PRT <213> Homo sapiens <400> 52 Met Asp Leu Leu Val Asp Glu Leu Phe Ala Asp Met Asn Ala Asp Gly 1 5 10 15 Ala Ser Pro Pro Pro Pro Arg Pro Ala Gly Gly Pro Lys Asn Thr Pro 20 25 30 Ala Ala Pro Pro Leu Tyr Ala Thr Gly Arg Leu Ser Gln Ala Gln Leu 35 40 45 Met Pro Ser Pro Pro Met Pro Val Pro Pro Ala Ala Leu Phe Asn Arg 50 55 60 Leu Leu Asp Asp Leu Gly Phe Ser Ala Gly Pro Ala Leu Cys Thr Met 65 70 75 80 Leu Asp Thr Trp Asn Glu Asp Leu Phe Ser Ala Leu Pro Thr Asn Ala 85 90 95 Asp Leu Tyr Arg Glu Cys Lys Phe Leu Ser Thr Leu Pro Ser Asp Val 100 105 110 Val Glu Trp Gly Asp Ala Tyr Val Pro Glu Arg Thr Gln Ile Asp Ile 115 120 125 Arg Ala His Gly Asp Val Ala Phe Pro Thr Leu Pro Ala Thr Arg Asp 130 135 140 Gly Leu Gly Leu Tyr Tyr Glu Ala Leu Ser Arg Phe Phe His Ala Glu 145 150 155 160 Leu Arg Ala Arg Glu Glu Ser Tyr Arg Thr Val Leu Ala Asn Phe Cys 165 170 175 Ser Ala Leu Tyr Arg Tyr Leu Arg Ala Ser Val Arg Gln Leu His Arg 180 185 190 Gln Ala His Met Arg Gly Arg Asp Arg Asp Leu Gly Glu Met Leu Arg 195 200 205 Ala Thr Ile Ala Asp Arg Tyr Tyr Arg Glu Thr Ala Arg Leu Ala Arg 210 215 220 Val Leu Phe Leu His Leu Tyr Leu Phe Leu Thr Arg Glu Ile Leu Trp 225 230 235 240 Ala Ala Tyr Ala Glu Gln Met Met Arg Pro Asp Leu Phe Asp Cys Leu 245 250 255 Cys Cys Asp Leu Glu Ser Trp Arg Gln Leu Ala Gly Leu Phe Gln Pro 260 265 270 Phe Met Phe Val Asn Gly Ala Leu Thr Val Arg Gly Val Pro Ile Glu 275 280 285 Ala Arg Arg Leu Arg Glu Leu Asn His Ile Arg Glu His Leu Asn Leu 290 295 300 Pro Leu Val Arg Ser Ala Ala Thr Glu Glu Pro Gly Ala Pro Leu Thr 305 310 315 320 Thr Pro Pro Thr Leu His Gly Asn Gln Ala Arg Ala Ser Gly Tyr Phe 325 330 335 Met Val Leu Ile Arg Ala Lys Leu Asp Ser Tyr Ser Ser Phe Thr Thr 340 345 350 Ser Pro Ser Glu Ala Val Met Arg Glu His Ala Tyr Ser Arg Ala Arg 355 360 365 Thr Lys Asn Asn Tyr Gly Ser Thr Ile Glu Gly Leu Leu Asp Leu Pro 370 375 380 Asp Asp Asp Ala Pro Glu Glu Ala Gly Leu Ala Ala Pro Arg Leu Ser 385 390 395 400 Phe Leu Pro Ala Gly His Thr Arg Arg Leu Ser Thr Ala Pro Pro Thr 405 410 415 Asp Val Ser Leu Gly Asp Glu Leu His Leu Asp Gly Glu Asp Val Ala 420 425 430 Met Ala His Ala Asp Ala Leu Asp Asp Phe Asp Leu Asp Met Leu Gly 435 440 445 Asp Gly Asp Ser Pro Gly Pro Gly Phe Thr Pro His Asp Ser Ala Pro 450 455 460 Tyr Gly Ala Leu Asp Met Ala Asp Phe Glu Phe Glu Gln Met Phe Thr 465 470 475 480 Asp Ala Leu Gly Ile Asp Glu Tyr Gly Gly 485 490 <210> 53 <211> 2570 <212> DNA <213> Homo sapiens <400> 53 agcgcgcagg cgcggccgga ttccgggcag tgacgcgacg gcgggccgcg cggcgcattt 60 ccgcctctgg cgaatggctc gtctgtagtg cacgccgcgg gcccagctgc gaccccggcc 120 ccgcccccgg gaccccggcc atggacgaac tgttccccct catcttcccg gcagagccag 180 cccaggcctc tggcccctat gtggagatca ttgagcagcc caagcagcgg ggcatgcgct 240 tccgctacaa gtgcgagggg cgctccgcgg gcagcatccc aggcgagagg agcacagata 300 ccaccaagac ccaccccacc atcaagatca atggctacac aggaccaggg acagtgcgca 360 tctccctggt caccaaggac cctcctcacc ggcctcaccc ccacgagctt gtaggaaagg 420 actgccggga tggcttctat gaggctgagc tctgcccgga ccgctgcatc cacagtttcc 480 agaacctggg aatccagtgt gtgaagaagc gggacctgga gcaggctatc agtcagcgca 540 tccagaccaa caacaacccc ttccaagaag agcagcgtgg ggactacgac ctgaatgctg 600 tgcggctctg cttccaggtg acagtgcggg acccatcagg caggcccctc cgcctgccgc 660 ctgtcctttc tcatcccatc tttgacaatc gtgcccccaa cactgccgag ctcaagatct 720 gccgagtgaa ccgaaactct ggcagctgcc tcggtgggga tgagatcttc ctactgtgtg 780 acaaggtgca gaaagaggac attgaggtgt atttcacggg accaggctgg gaggcccgag 840 gctccttttc gcaagctgat gtgcaccgac aagtggccat tgtgttccgg acccctccct 900 acgcagaccc cagcctgcag gctcctgtgc gtgtctccat gcagctgcgg cggccttccg 960 accgggagct cagtgagccc atggaattcc agtacctgcc agatacagac gatcgtcacc 1020 ggattgagga gaaacgtaaa aggacatatg agaccttcaa gagcatcatg aagaagagtc 1080 ctttcagcgg acccaccgac ccccggcctc cacctcgacg cattgctgtg ccttcccgca 1140 gctcagcttc tgtccccaag ccagcacccc agccctatcc ctttacgtca tccctgagca 1200 ccatcaacta tgatgagttt cccaccatgg tgtttccttc tgggcagatc agccaggcct 1260 cggccttggc cccggcccct ccccaagtcc tgccccaggc tccagcccct gcccctgctc 1320 cagccatggt atcagctctg gcccaggccc cagcccctgt cccagtccta gccccaggcc 1380 ctcctcaggc tgtggcccca cctgccccca agcccaccca ggctggggaa ggaacgctgt 1440 ctcctcaggc tgtggcccca cctgccccca agcccaccca ggctggggaa ggaacgctgt 1440 cagaggccct gctgcagctg cagtttgatg atgaagacct gggggccttg cttggcaaca 1500 cagaggccct gctgcagctg cagtttgatg atgaagacct gggggccttg cttggcaaca 1500 gcacagaccc agctgtgttc acagacctgg catccgtcga caactccgag tttcagcagc 1560 gcacagaccc agctgtgttc acagacctgg catccgtcga caactccgag tttcagcagc 1560 tgctgaacca gggcatacct gtggcccccc acacaactga gcccatgctg atggagtacc 1620 tgctgaacca gggcatacct gtggcccccc acacaactga gcccatgctg atggagtacc 1620 ctgaggctat aactcgccta gtgacagggg cccagaggcc ccccgaccca gctcctgctc 1680 ctgaggctat aactcgccta gtgacagggg cccagaggcc ccccgaccca gctcctgctc 1680 cactgggggc cccggggctc cccaatggcc tcctttcagg agatgaagac ttctcctcca 1740 cactgggggc cccggggctc cccaatggcc tcctttcagg agatgaagac ttctcctcca 1740 ttgcggacat ggacttctca gccctgctga gtcagatcag ctcctaaggg ggtgacgcct 1800 ttgcggacat ggacttctca gccctgctga gtcagatcag ctcctaaggg ggtgacgcct 1800 gccctcccca gagcactggg ttgcagggga ttgaagccct ccaaaagcac ttacggattc 1860 gccctcccca gagcactggg ttgcagggga ttgaagccct ccaaaagcac ttacggattc 1860 tggtggggtg tgttccaact gcccccaact ttgtggatgt cttccttgga ggggggagcc 1920 tggtggggtg tgttccaact gcccccaact ttgtggatgt cttccttgga ggggggagcc 1920 atattttatt cttttattgt cagtatctgt atctctctct ctttttggag gtgcttaagc 1980 atattttatt cttttattgt cagtatctgt atctctctct ctttttggag gtgcttaagc 1980 agaagcatta acttctctgg aaagggggga gctggggaaa ctcaaacttt tcccctgtcc 2040 agaagcatta acttctctgg aaagggggga gctggggaaa ctcaaacttt tcccctgtcc 2040 tgatggtcag ctcccttctc tgtagggaac tctggggtcc cccatcccca tcctccagct 2100 tgatggtcag ctcccttctc tgtagggaac tctggggtcc cccatcccca tcctccagct 2100 tctggtactc tcctagagac agaagcaggc tggaggtaag gcctttgagc ccacaaagcc 2160 ttatcaagtg tcttccatca tggattcatt acagcttaat caaaataacg ccccagatac 2220 cagcccctgt atggcactgg cattgtccct gtgcctaaca ccagcgtttg aggggctggc 2280 cttcctgccc tacagaggtc tctgccggct ctttccttgc tcaaccatgg ctgaaggaaa 2340 ccagtgcaac agcactggct ctctccagga tccagaaggg gtttggtctg ggacttcctt 2400 gctctccctc ttctcaagtg ccttaatagt agggtaagtt gttaagagtg ggggagagca 2460 ggctggcagc tctccagtca ggaggcatag tttttactga acaatcaaag cacttggact 2520 cttgctcttt ctactctgaa ctaataaatc tgttgccaag ctggctagaa 2570 <210> 54 <211> 548 <212> PRT <213> Homo sapiens <400> 54 Met Asp Glu Leu Phe Pro Leu Ile Phe Pro Ala Glu Pro Ala Gln Ala 1 5 10 15 Ser Gly Pro Tyr Val Glu Ile Ile Glu Gln Pro Lys Gln Arg Gly Met 20 25 30 Arg Phe Arg Tyr Lys Cys Glu Gly Arg Ser Ala Gly Ser Ile Pro Gly 35 40 45 Glu Arg Ser Thr Asp Thr Thr Lys Thr His Pro Thr Ile Lys Ile Asn 50 55 60 Gly Tyr Thr Gly Pro Gly Thr Val Arg Ile Ser Leu Val Thr Lys Asp 65 70 75 80 Pro Pro His Arg Pro His Pro His Glu Leu Val Gly Lys Asp Cys Arg 85 90 95 Asp Gly Phe Tyr Glu Ala Glu Leu Cys Pro Asp Arg Cys Ile His Ser 100 105 110 Phe Gln Asn Leu Gly Ile Gln Cys Val Lys Lys Arg Asp Leu Glu Gln 115 120 125 Ala Ile Ser Gln Arg Ile Gln Thr Asn Asn Asn Pro Phe Gln Glu Glu 130 135 140 Gln Arg Gly Asp Tyr Asp Leu Asn Ala Val Arg Leu Cys Phe Gln Val 145 150 155 160 Thr Val Arg Asp Pro Ser Gly Arg Pro Leu Arg Leu Pro Pro Val Leu 165 170 175 Ser His Pro Ile Phe Asp Asn Arg Ala Pro Asn Thr Ala Glu Leu Lys 180 185 190 Ile Cys Arg Val Asn Arg Asn Ser Gly Ser Cys Leu Gly Gly Asp Glu 195 200 205 Ile Phe Leu Leu Cys Asp Lys Val Gln Lys Glu Asp Ile Glu Val Tyr 210 215 220 Phe Thr Gly Pro Gly Trp Glu Ala Arg Gly Ser Phe Ser Gln Ala Asp 225 230 235 240 Val His Arg Gln Val Ala Ile Val Phe Arg Thr Pro Pro Tyr Ala Asp 245 250 255 Pro Ser Leu Gln Ala Pro Val Arg Val Ser Met Gln Leu Arg Arg Pro 260 265 270 Ser Asp Arg Glu Leu Ser Glu Pro Met Glu Phe Gln Tyr Leu Pro Asp 275 280 285 Thr Asp Asp Arg His Arg Ile Glu Glu Lys Arg Lys Arg Thr Tyr Glu 290 295 300 Thr Phe Lys Ser Ile Met Lys Lys Ser Pro Phe Ser Gly Pro Thr Asp 305 310 315 320 Pro Arg Pro Pro Pro Arg Arg Ile Ala Val Pro Ser Arg Ser Ser Ala 325 330 335 Ser Val Pro Lys Pro Ala Pro Gln Pro Tyr Pro Phe Thr Ser Ser Leu 340 345 350 Ser Thr Ile Asn Tyr Asp Glu Phe Pro Thr Met Val Phe Pro Ser Gly 355 360 365 Gln Ile Ser Gln Ala Ser Ala Leu Ala Pro Ala Pro Pro Gln Val Leu 370 375 380 Pro Gln Ala Pro Ala Pro Ala Pro Ala Pro Ala Met Val Ser Ala Leu 385 390 395 400 Ala Gln Ala Pro Ala Pro Val Pro Val Leu Ala Pro Gly Pro Pro Gln 405 410 415 Ala Val Ala Pro Pro Ala Pro Lys Pro Thr Gln Ala Gly Glu Gly Thr 420 425 430 Leu Ser Glu Ala Leu Leu Gln Leu Gln Phe Asp Asp Glu Asp Leu Gly 435 440 445 Ala Leu Leu Gly Asn Ser Thr Asp Pro Ala Val Phe Thr Asp Leu Ala 450 455 460 Ser Val Asp Asn Ser Glu Phe Gln Gln Leu Leu Asn Gln Gly Ile Pro 465 470 475 480 Val Ala Pro His Thr Thr Glu Pro Met Leu Met Glu Tyr Pro Glu Ala 485 490 495 Ile Thr Arg Leu Val Thr Gly Ala Gln Arg Pro Pro Asp Pro Ala Pro 500 505 510 Ala Pro Leu Gly Ala Pro Gly Leu Pro Asn Gly Leu Leu Ser Gly Asp 515 520 525 Glu Asp Phe Ser Ser Ile Ala Asp Met Asp Phe Ser Ala Leu Leu Ser 530 535 540 Gln Ile Ser Ser 545 <210> 55 <211> 1815 <212> DNA <213> Homo sapiens <400> 55 atgcgcccga aaaaagatgg cctggaagat tttctgcgcc tgaccccgga aattaaaaaa 60 cagctgggca gcctggtgag cgattattgc aacgtgctga acaaagaatt taccgcgggc 120 agcgtggaaa ttaccctgcg cagctataaa atttgcaaag cgtttattaa cgaagcgaaa 180 gcgcatggcc gcgaatgggg cggcctgatg gcgaccctga acatttgcaa cttttgggcg 240 attctgcgca acaaccgcgt gcgccgccgc gcggaaaacg cgggcaacga tgcgtgcagc 300 attgcgtgcc cgattgtgat gcgctatgtg ctggatcatc tgattgtggt gaccgatcgc 360 ttttttattc aggcgccgag caaccgcgtg atgattccgg cgaccattgg caccgcgatg 420 tataaactgc tgaaacatag ccgcgtgcgc gcgtatacct atagcaaagt gctgggcgtg 480 gatcgcgcgg cgattatggc gagcggcaaa caggtggtgg aacatctgaa ccgcatggaa 540 aaagaaggcc tgctgagcag caaatttaaa gcgttttgca aatgggtgtt tacctatccg 600 gtgctggaag aaatgtttca gaccatggtg agcagcaaaa ccggccatct gaccgatgat 660 gtgaaagatg tgcgcgcgct gattaaaacc ctgccgcgcg cgagctatag cagccatgcg 720 ggccagcgca gctatgtgag cggcgtgctg ccggcgtgcc tgctgagcac caaaagcaaa 780 gcggtggaaa ccccgattct ggtgagcggc gcggatcgca tggatgaaga actgatgggc 840 aacgatggcg gcgcgagcca taccgaagcg cgctatagcg aaagcggcca gtttcatgcg 900 tttaccgatg aactggaaag cctgccgagc ccgaccatgc cgctgaaacc gggcgcgcag 960 agcgcggatt gcggcgatag cagcagcagc agcagcgata gcggcaacag cgataccgaa 1020 cagagcgaac gcgaagaagc gcgcgcggaa gcgccgcgcc tgcgcgcgcc gaaaagccgc 1080 cgcaccagcc gcccgaaccg cggccagacc ccgtgcccga gcaacgcggc ggaaccggaa 1140 cagccgtgga ttgcggcggt gcatcaggaa agcgatgaac gcccgatttt tccgcatccg 1200 agcaaaccga cctttctgcc gccggtgaaa cgcaaaaaag gcctgcgcga tagccgcgaa 1260 ggcatgtttc tgccgaaacc ggaagcgggc agcgcgatta gcgatgtgtt tgaaggccgc 1320 gaagtgtgcc agccgaaacg cattcgcccg tttcatccgc cgggcagccc gtgggcgaac 1380 cgcccgctgc cggcgagcct ggcgccgacc ccgaccggcc cggtgcatga accggtgggc 1440 agcctgaccc cggcgccggt gccgcagccg ctggatccgg cgccggcggt gaccccggaa 1500 gcgagccatc tgctggaaga tccggatgaa gaaaccagcc aggcggtgaa agcgctgcgc 1560 gaaatggcgg ataccgtgat tccgcagaaa gaagaagcgg cgatttgcgg ccagatggat 1620 ctgagccatc cgccgccgcg cggccatctg gatgaactga ccaccaccct ggaaagcatg 1680 accgaagatc tgaacctgga tagcccgctg accccggaac tgaacgaaat tctggatacc 1740 tttctgaacg atgaatgcct gctgcatgcg atgcatatta gcaccggcct gagcattttt 1800 gataccagcc tgttt 1815 <210> 56 <211> 605 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 56 Met Arg Pro Lys Lys Asp Gly Leu Glu Asp Phe Leu Arg Leu Thr Pro 1 5 10 15 Glu Ile Lys Lys Gln Leu Gly Ser Leu Val Ser Asp Tyr Cys Asn Val 20 25 30 Leu Asn Lys Glu Phe Thr Ala Gly Ser Val Glu Ile Thr Leu Arg Ser 35 40 45 Tyr Lys Ile Cys Lys Ala Phe Ile Asn Glu Ala Lys Ala His Gly Arg 50 55 60 Glu Trp Gly Gly Leu Met Ala Thr Leu Asn Ile Cys Asn Phe Trp Ala 65 70 75 80 Ile Leu Arg Asn Asn Arg Val Arg Arg Arg Ala Glu Asn Ala Gly Asn 85 90 95 Asp Ala Cys Ser Ile Ala Cys Pro Ile Val Met Arg Tyr Val Leu Asp 100 105 110 His Leu Ile Val Val Thr Asp Arg Phe Phe Ile Gln Ala Pro Ser Asn 115 120 125 Arg Val Met Ile Pro Ala Thr Ile Gly Thr Ala Met Tyr Lys Leu Leu 130 135 140 Lys His Ser Arg Val Arg Ala Tyr Thr Tyr Ser Lys Val Leu Gly Val 145 150 155 160 Asp Arg Ala Ala Ile Met Ala Ser Gly Lys Gln Val Val Glu His Leu 165 170 175 Asn Arg Met Glu Lys Glu Gly Leu Leu Ser Ser Lys Phe Lys Ala Phe 180 185 190 Cys Lys Trp Val Phe Thr Tyr Pro Val Leu Glu Glu Met Phe Gln Thr 195 200 205 Met Val Ser Ser Lys Thr Gly His Leu Thr Asp Asp Val Lys Asp Val 210 215 220 Arg Ala Leu Ile Lys Thr Leu Pro Arg Ala Ser Tyr Ser Ser His Ala 225 230 235 240 Gly Gln Arg Ser Tyr Val Ser Gly Val Leu Pro Ala Cys Leu Leu Ser 245 250 255 Thr Lys Ser Lys Ala Val Glu Thr Pro Ile Leu Val Ser Gly Ala Asp 260 265 270 Arg Met Asp Glu Glu Leu Met Gly Asn Asp Gly Gly Ala Ser His Thr 275 280 285 Glu Ala Arg Tyr Ser Glu Ser Gly Gln Phe His Ala Phe Thr Asp Glu 290 295 300 Leu Glu Ser Leu Pro Ser Pro Thr Met Pro Leu Lys Pro Gly Ala Gln 305 310 315 320 Ser Ala Asp Cys Gly Asp Ser Ser Ser Ser Ser Ser Asp Ser Gly Asn 325 330 335 Ser Asp Thr Glu Gln Ser Glu Arg Glu Glu Ala Arg Ala Glu Ala Pro 340 345 350 Arg Leu Arg Ala Pro Lys Ser Arg Arg Thr Ser Arg Pro Asn Arg Gly 355 360 365 Gln Thr Pro Cys Pro Ser Asn Ala Ala Glu Pro Glu Gln Pro Trp Ile 370 375 380 Ala Ala Val His Gln Glu Ser Asp Glu Arg Pro Ile Phe Pro His Pro 385 390 395 400 Ser Lys Pro Thr Phe Leu Pro Pro Val Lys Arg Lys Lys Gly Leu Arg 405 410 415 Asp Ser Arg Glu Gly Met Phe Leu Pro Lys Pro Glu Ala Gly Ser Ala 420 425 430 Ile Ser Asp Val Phe Glu Gly Arg Glu Val Cys Gln Pro Lys Arg Ile 435 440 445 Arg Pro Phe His Pro Pro Gly Ser Pro Trp Ala Asn Arg Pro Leu Pro 450 455 460 Ala Ser Leu Ala Pro Thr Pro Thr Gly Pro Val His Glu Pro Val Gly 465 470 475 480 Ser Leu Thr Pro Ala Pro Val Pro Gln Pro Leu Asp Pro Ala Pro Ala 485 490 495 Val Thr Pro Glu Ala Ser His Leu Leu Glu Asp Pro Asp Glu Glu Thr 500 505 510 Ser Gln Ala Val Lys Ala Leu Arg Glu Met Ala Asp Thr Val Ile Pro 515 520 525 Gln Lys Glu Glu Ala Ala Ile Cys Gly Gln Met Asp Leu Ser His Pro 530 535 540 Pro Pro Arg Gly His Leu Asp Glu Leu Thr Thr Thr Leu Glu Ser Met 545 550 555 560 Thr Glu Asp Leu Asn Leu Asp Ser Pro Leu Thr Pro Glu Leu Asn Glu 565 570 575 Ile Leu Asp Thr Phe Leu Asn Asp Glu Cys Leu Leu His Ala Met His 580 585 590 Ile Ser Thr Gly Leu Ser Ile Phe Asp Thr Ser Leu Phe 595 600 605 <210> 57 <211> 172 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 57 Arg Pro Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Gln Arg Gly 1 5 10 15 Asn Leu Val Arg His Ile Arg Thr His Thr Gly Glu Lys Pro Phe Ala 20 25 30 Cys Asp Ile Cys Gly Lys Lys Phe Ala Leu Ser Phe Asn Leu Thr Arg 35 40 45 His Thr Lys Ile His Thr Gly Ser Gln Lys Pro Phe Gln Cys Arg Ile 50 55 60 Cys Met Arg Asn Phe Ser Arg Ser Asp Asn Leu Thr Arg His Ile Arg 65 70 75 80 Thr His Thr Gly Glu Lys Pro Phe Ala Cys Asp Ile Cys Gly Lys Lys 85 90 95 Phe Ala Asp Arg Ser His Leu Ala Arg His Thr Lys Ile His Thr Gly 100 105 110 Ser Gln Lys Pro Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Gln 115 120 125 Lys Ala His Leu Thr Ala His Ile Arg Thr His Thr Gly Glu Lys Pro 130 135 140 Phe Ala Cys Asp Ile Cys Gly Arg Lys Phe Ala Arg Ser Asp Asn Leu 145 150 155 160 Thr Arg His Thr Lys Ile His Leu Arg Gln Lys Asp 165 170 <210> 58 <211> 516 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 58 cgaccattcc agtgtcgaat ctgcatgcgc aacttcagcc agcggggaaa cctggtgagg 60 catatccgca cccacacggg agagaagcct tttgcctgcg atatttgtgg aaagaagttt 120 gctctgagct tcaatctaac cagacacacc aagattcata ctgggtccca gaaaccgttc 180 cagtgtagga tatgcatgag gaatttctct cggagtgaca acttaacgcg gcatataagg 240 acgcacacag gtgaaaaacc atttgcatgc gacatctgtg gcaaaaagtt tgcggaccgg 300 tctcaccttg cccgacacac aaaaatccat accggcagtc aaaagccctt tcaatgtcgc 360 atttgcatgc gaaacttctc acagaaggcc catttgactg cccatattcg tactcatact 420 ggcgagaaac ctttcgcttg cgatatatgt ggtcgtaagt ttgcacggtc ggacaacctc 480 acacgccaca ctaagataca cctgcggcag aaggac 516 <210> 59 <211> 172 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 59 Arg Pro Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Arg Ser Ser 1 5 10 15 Asn Leu Thr Arg His Ile Arg Thr His Thr Gly Glu Lys Pro Phe Ala 20 25 30 Cys Asp Ile Cys Gly Lys Lys Phe Ala Asp Lys Arg Thr Leu Ile Arg 35 40 45 His Thr Lys Ile His Thr Gly Ser Gln Lys Pro Phe Gln Cys Arg Ile 50 55 60 Cys Met Arg Asn Phe Ser Gln Arg Gly Asn Leu Val Arg His Ile Arg 65 70 75 80 Thr His Thr Gly Glu Lys Pro Phe Ala Cys Asp Ile Cys Gly Lys Lys 85 90 95 Phe Ala Leu Ser Phe Asn Leu Thr Arg His Thr Lys Ile His Thr Gly 100 105 110 Ser Gln Lys Pro Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Arg 115 120 125 Ser Asp Asn Leu Thr Arg His Ile Arg Thr His Thr Gly Glu Lys Pro 130 135 140 Phe Ala Cys Asp Ile Cys Gly Arg Lys Phe Ala Asp Arg Ser His Leu 145 150 155 160 Ala Arg His Thr Lys Ile His Leu Arg Gln Lys Asp 165 170 <210> 60 <211> 516 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 60 cgaccattcc agtgtcgaat ctgcatgcgc aacttcagcc gaagttccaa cctgacacgg 60 catatccgca cccacacggg agagaagcct tttgcctgcg atatttgtgg aaagaagttt 120 gctgacaagc ggaccttaat ccgccacacc aagattcata ctgggtccca gaaaccgttc 180 cagtgtagga tatgcatgag gaatttctct cagcggggaa atctagtgcg acatataagg 240 acgcacacag gtgaaaaacc atttgcatgc gacatctgtg gcaaaaagtt tgcgctgagc 300 ttcaacttga ctcgtcacac aaaaatccat accggcagtc aaaagccctt tcaatgtcgc 360 atttgcatgc gaaacttctc acggagtgac aatcttacga gacatattcg tactcatact 420 ggcgagaaac ctttcgcttg cgatatatgt ggtcgtaagt ttgcagaccg gagccactta 480 gccaggcaca ctaagataca cctgcggcag aaggac 516 <210> 61 <211> 172 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 61 Arg Pro Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Asp Arg Ser 1 5 10 15 Ala Leu Ala Arg His Ile Arg Thr His Thr Gly Glu Lys Pro Phe Ala 20 25 30 Cys Asp Ile Cys Gly Lys Lys Phe Ala Arg Ser Asp Asn Leu Thr Arg 35 40 45 His Thr Lys Ile His Thr Gly Ser Gln Lys Pro Phe Gln Cys Arg Ile 50 55 60 Cys Met Arg Asn Phe Ser Gln Ser Gly Asp Leu Thr Arg His Ile Arg 65 70 75 80 Thr His Thr Gly Glu Lys Pro Phe Ala Cys Asp Ile Cys Gly Lys Lys 85 90 95 Phe Ala Val Arg Gln Thr Leu Lys Gln His Thr Lys Ile His Thr Gly 100 105 110 Ser Gln Lys Pro Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Ala 115 120 125 Ala Gly Asn Leu Thr Arg His Ile Arg Thr His Thr Gly Glu Lys Pro 130 135 140 Phe Ala Cys Asp Ile Cys Gly Arg Lys Phe Ala Arg Ser Asp Asn Leu 145 150 155 160 Thr Arg His Thr Lys Ile His Leu Arg Gln Lys Asp 165 170 <210> 62 <211> 516 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 62 cgaccattcc agtgtcgaat ctgcatgcgc aacttcagcg accggagcgc gctggcacgg 60 catatccgca cccacacggg agagaagcct tttgcctgcg atatttgtgg aaagaagttt 120 gctcgaagtg acaacttaac gcgccacacc aagattcata ctgggtccca gaaaccgttc 180 cagtgtagga tatgcatgag gaatttctct cagtcagggg acctcactcg tcatataagg 240 acgcacacag gtgaaaaacc atttgcatgc gacatctgtg gcaaaaagtt tgcggtacga 300 cagacgctta aacaacacac aaaaatccat accggcagtc aaaagccctt tcaatgtcgc 360 atttgcatgc gaaacttctc agccgctggt aacttgacac gacatattcg tactcatact 420 ggcgagaaac ctttcgcttg cgatatatgt ggtcgtaagt ttgcaagatc tgataatcta 480 acgcgtcaca ctaagataca cctgcggcag aaggac 516 <210> 63 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 63 Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Gln Arg Gly Asn Leu 1 5 10 15 Val Arg His Ile Arg Thr His Thr Gly Glu Lys Pro 20 25 <210> 64 <211> 29 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 64 Phe Ala Cys Asp Ile Cys Gly Lys Lys Phe Ala Leu Ser Phe Asn Leu 1 5 10 15 Thr Arg His Thr Lys Ile His Thr Gly Ser Gln Lys Pro 20 25 <210> 65 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 65 Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Arg Ser Asp Asn Leu 1 5 10 15 Thr Arg His Ile Arg Thr His Thr Gly Glu Lys Pro 20 25 <210> 66 <211> 29 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 66 Phe Ala Cys Asp Ile Cys Gly Lys Lys Phe Ala Asp Arg Ser His Leu 1 5 10 15 Ala Arg His Thr Lys Ile His Thr Gly Ser Gln Lys Pro 20 25 <210> 67 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 67 Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Gln Lys Ala His Leu 1 5 10 15 Thr Ala His Ile Arg Thr His Thr Gly Glu Lys Pro 20 25 <210> 68 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 68 Phe Ala Cys Asp Ile Cys Gly Arg Lys Phe Ala Arg Ser Asp Asn Leu 1 5 10 15 Thr Arg His Thr Lys Ile His Leu Arg Gln Lys Asp 20 25 <210> 69 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 69 Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Arg Ser Ser Asn Leu 1 5 10 15 Thr Arg His Ile Arg Thr His Thr Gly Glu Lys Pro 20 25 <210> 70 <211> 29 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 70 Phe Ala Cys Asp Ile Cys Gly Lys Lys Phe Ala Asp Lys Arg Thr Leu 1 5 10 15 Ile Arg His Thr Lys Ile His Thr Gly Ser Gln Lys Pro 20 25 <210> 71 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 71 Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Gln Arg Gly Asn Leu 1 5 10 15 Val Arg His Ile Arg Thr His Thr Gly Glu Lys Pro 20 25 <210> 72 <211> 29 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 72 Phe Ala Cys Asp Ile Cys Gly Lys Lys Phe Ala Leu Ser Phe Asn Leu 1 5 10 15 Thr Arg His Thr Lys Ile His Thr Gly Ser Gln Lys Pro 20 25 <210> 73 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthesis <400> 73 Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Arg Ser Asp Asn Leu 1 5 10 15 Thr Arg His Ile Arg Thr His Thr Gly Glu Lys Pro 20 25 <210> 74 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthesis <400> 74 Phe Ala Cys Asp Ile Cys Gly Arg Lys Phe Ala Asp Arg Ser His Leu 1 5 10 15 Ala Arg His Thr Lys Ile His Leu Arg Gln Lys Asp 20 25 <210> 75 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthesis <400> 75 Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Asp Arg Ser Ala Leu 1 5 10 15 Ala Arg His Ile Arg Thr His Thr Gly Glu Lys Pro 20 25 <210> 76 <211> 29 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 76 Phe Ala Cys Asp Ile Cys Gly Lys Lys Phe Ala Arg Ser Asp Asn Leu 1 5 10 15 Thr Arg His Thr Lys Ile His Thr Gly Ser Gln Lys Pro 20 25 <210> 77 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 77 Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Gln Ser Gly Asp Leu 1 5 10 15 Thr Arg His Ile Arg Thr His Thr Gly Glu Lys Pro 20 25 <210> 78 <211> 29 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 78 Phe Ala Cys Asp Ile Cys Gly Lys Lys Phe Ala Val Arg Gln Thr Leu 1 5 10 15 Lys Gln His Thr Lys Ile His Thr Gly Ser Gln Lys Pro 20 25 <210> 79 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 79 Phe Gln Cys Arg Ile Cys Met Arg Asn Phe Ser Ala Ala Gly Asn Leu 1 5 10 15 Thr Arg His Ile Arg Thr His Thr Gly Glu Lys Pro 20 25 <210> 80 <211> 28 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 80 Phe Ala Cys Asp Ile Cys Gly Arg Lys Phe Ala Arg Ser Asp Asn Leu 1 5 10 15 Thr Arg His Thr Lys Ile His Leu Arg Gln Lys Asp 20 25 <210> 81 <211> 4104 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 81 atggacaaga agtactccat tgggctcgct atcggtacca acagcgtcgg ctgggccgtc 60 attacggacg agtacaaggt gccgagcaaa aaattcaaag ttctgggcaa taccgatcgc 120 cacagcataa agaagaacct cattggagcc ctcctgttcg actccgggga gacggccgaa 180 gccacgcggc tcaaaagaac agcacggcgc agatataccc gcagaaagaa tcggatctgc 240 tacctgcagg agatctttag taatgagatg gctaaggtgg atgactcttt cttccatagg 300 ctggaggagt cctttttggt ggaggaggat aaaaagcacg agcgccaccc aatctttggc 360 aatatcgtgg acgaggtggc gtaccatgaa aagtacccaa ccatatatca tctgaggaag 420 aagctggtag acagtactga taaggctgac ttgcggttga tctatctcgc gctggcgcac 480 atgatcaaat ttcggggaca cttcctcatc gagggggacc tgaacccaga caacagcgat 540 gtcgacaaac tctttatcca actggttcag acttacaatc agcttttcga ggagaacccg 600 atcaacgcat ccggcgttga cgccaaagca atcctgagcg ctaggctgtc caaatcccgg 660 cggctcgaaa acctcatcgc acagctccct ggggagaaga agaacggcct gtttggtaat 720 cttatcgccc tgtcactcgg gctgaccccc aactttaaat ctaacttcga cctggccgaa 780 gatgccaagc tgcaactgag caaagacacc tacgatgatg atctcgacaa tctgctggcc 840 cagatcggcg accagtacgc agaccttttt ttggcggcaa agaacctgtc agacgccatt 900 ctgctgagtg atattctgcg agtgaacacg gagatcacca aagctccgct gagcgctagt 960 atgatcaagc gctatgatga gcaccaccaa gacttgactt tgctgaaggc ccttgtcaga 1020 cagcaactgc ctgagaagta caaggaaatt ttcttcgatc agtctaaaaa tggctacgcc 1080 ggatacattg acggcggagc aagccaggag gaattttaca aatttattaa gcccatcttg 1140 gaaaaaatgg acggcaccga ggagctgctg gtaaagctga acagagaaga tctgttgcgc 1200 aaacagcgca ctttcgacaa tggaagcatc ccccaccaga ttcacctggg cgaactgcac 1260 gctatcctca ggcggcaaga ggatttctac ccctttttga aagataacag ggaaaagatt 1320 gagaaaatcc tcacatttcg gataccctac tatgtaggcc ccctcgctcg gggaaattcc 1380 agattcgcgt ggatgactcg caaatcagaa gagaccatca ctccctggaa cttcgaggaa 1440 gtcgtggata agggggcctc tgcccagtcc ttcatcgaaa ggatgactaa ctttgataaa 1500 aatctgccta acgaaaaggt gcttcctaaa cactctctgc tgtacgagta cttcacagtt 1560 tataacgagc tcaccaaggt caaatacgtc acagaaggga tgagaaagcc agcattcctg 1620 tctggagagc agaagaaagc tatcgtggac ctcctcttca agacgaaccg gaaagttacc 1680 gtgaaacagc tcaaagaaga ctatttcaaa aagattgaat gtttcgactc tgttgaaatc 1740 agcggagtgg aggatcgctt caacgcatcc ctgggaacgt atcacgatct cctgaaaatc 1800 attaaagaca aggacttcct ggacaatgag gagaacgagg acattcttga ggacattgtc 1860 ctcaccctta cgttgtttga agatagggag atgattgaag aacgcttgaa aacttacgct 1920 catctcttcg acgacaaagt catgaaacag ctcaagagac gccgatatac aggatggggg 1980 cggctgtcaa gaaaactgat caatggcatc cgagacaagc agagtggaaa gacaatcctg 2040 gattttctta agtccgatgg atttgccaac cggaacttca tgcagttgat ccatgatgac 2100 tctctcacct ttaaggagga catccagaaa gcacaagttt ctggccaggg ggacagtctt 2160 cacgagcaca tcgctaatct tgcaggtagc ccagctatca aaaagggaat actgcagacc 2220 gttaaggtcg tggatgaact cgtcaaagta atgggaaggc ataagcccga gaatatcgtt 2280 atcgagatgg cccgagagaa ccaaactacc cagaagggac agaagaacag tagggaaagg 2340 atgaagagga ttgaagaggg tataaaagaa ctggggtccc aaatccttaa ggaacaccca 2400 gttgaaaaca cccagcttca gaatgagaag ctctacctgt actacctgca gaacggcagg 2460 gacatgtacg tggatcagga actggacatc aaccggttgt ccgactacga cgtggatgct 2520 atcgtgcccc aaagctttct caaagatgat tctattgata ataaagtgtt gacaagatcc 2580 gataaaaata gagggaagag tgataacgtc ccctcagaag aagttgtcaa gaaaatgaaa 2640 aattattggc ggcagctgct gaacgccaaa ctgatcacac aacggaagtt cgataatctg 2700 actaaggctg aacgaggtgg cctgtctgag ttggataaag ccggcttcat caaaaggcag 2760 cttgttgaga cacgccagat caccaagcac gtggcccaaa ttctcgattc acgcatgaac 2820 accaagtacg atgaaaatga caaactgatt cgagaggtga aagttattac tctgaagtct 2880 aagctggtct cagatttcag aaaggacttt cagttttata aggtgagaga gatcaacaat 2940 taccaccatg cgcatgatgc ctacctgaat gcagtggtag gcactgcact tatcaaaaaa 3000 tatcccaagc tggaatctga atttgtttac ggagactata aagtgtacga tgttaggaaa 3060 atgatcgcaa agtctgagca ggaaataggc aaggccaccg ctaagtactt cttttacagc 3120 aatattatga attttttcaa gaccgagatt acactggcca atggagagat tcggaagcga 3180 ccacttatcg aaacaaacgg agaaacagga gaaatcgtgt gggacaaggg tagggatttc 3240 gcgacagtcc gcaaggtcct gtccatgccg caggtgaaca tcgttaaaaa gaccgaagta 3300 cagaccggag gcttctccaa ggaaagtatc ctcccgaaaa ggaacagcga caagctgatc 3360 gcacgcaaaa aagattggga ccccaagaaa tacggcggat tcgattctcc tacagtcgct 3420 tacagtgtac tggttgtggc caaagtggag aaagggaagt ctaaaaaact caaaagcgtc 3480 aaggaactgc tgggcatcac aatcatggag cgatccagct tcgagaaaaa ccccatcgac 3540 tttctcgaag cgaaaggata taaagaggtc aaaaaagacc tcatcattaa gctgcccaag 3600 tactctctct ttgagcttga aaacggccgg aaacgaatgc tcgctagtgc gggcgagctg 3660 cagaaaggta acgagctggc actgccctct aaatacgtta atttcttgta tctggccagc 3720 cactatgaaa agctcaaagg gtctcccgaa gataatgagc agaagcagct gttcgtggaa 3780 caacacaaac actaccttga tgagatcatc gagcaaataa gcgagttctc caaaagagtg 3840 atcctcgccg acgctaacct cgataaggtg ctttctgctt acaataagca cagggataag 3900 cccatcaggg agcaggcaga aaacattatc cacttgttta ctctgaccaa cttgggcgcg 3960 cctgcagcct tcaagtactt cgacaccacc atagacagaa agcggtacac ctctacaaag 4020 gaggtcctgg acgccacact gattcatcag tcaattacgg ggctctatga aacaagaatc 4080 gacctctctc agctcggtgg agac 4104 <210> 82 <211> 1368 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 82 Met Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp Ala Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965 970 975 Glu Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 980 985 990 Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 <210> 83 <211> 97 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 83 gaggtaccat agagtgaggc ggttttagag ctagaaatag caagttaaaa taaggctagt 60 ccgttatcaa cttgaaaaag tggcaccgag tcggtgc 97 <210> 84 <211> 97 <212> RNA <213> Artificial sequence <220> <223> Synthetic <400> 84 gagguaccau agagugaggc gguuuuagag cuagaaauag caaguuaaaa uaaggcuagu 60 ccguuaucaa cuugaaaaag uggcaccgag ucggugc 97 <210> 85 <211> 21 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 85 gaggtaccat agagtgaggc g 21 <210> 86 <211> 21 <212> RNA <213> Artificial sequence <220> <223> Synthetic <400> 86 gagguaccau agagugaggc g 21 <210> 87 <211> 99 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 87 accgaggcga ggatgaagcc gaggttttag agctagaaat agcaagttaa aataaggcta 60 gtccgttatc aacttgaaaa agtggcaccg agtcggtgc 99 <210> 88 <211> 99 <212> RNA <213> Artificial sequence <220> <223> Synthetic <400> 88 accgaggcga ggaugaagcc gagguuuuag agcuagaaau agcaaguuaa aauaaggcua 60 guccguuauc aacuugaaaa aguggcaccg agucggugc 99 <210> 89 <211> 23 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 89 accgaggcga ggatgaagcc gag 23 <210> 90 <211> 23 <212> RNA <213> Artificial sequence <220> <223> Synthetic <400> 90 accgaggcga ggaugaagcc gag 23 <210> 91 <211> 100 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 91 accgaagccg agaggatact gcaggtttta gagctagaaa tagcaagtta aaataaggct 60 agtccgttat caacttgaaa aagtggcacc gagtcggtgc 100 <210> 92 <211> 100 <212> RNA <213> Artificial sequence <220> <223> Synthesis <400> 92 accgaagccg agaggauacu gcagguuuua gagcuagaaa uagcaaguua aaauaaggcu 60 aguccguuau caacuugaaa aaguggcacc gagucggugc 100 <210> 93 <211> 24 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 93 accgaagccg agaggatact gcag 24 <210> 94 <211> 24 <212> RNA <213> Artificial sequence <220> <223> Synthesis <400> 94 accgaagccg agaggauacu gcag 24 <210> 95 <211> 1569 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 95 gacgcattgg acgattttga tctggatatg ctgggaagtg acgccctcga tgattttgac 60 cttgacatgc ttggttcgga tgcccttgat gactttgacc tcgacatgct cggcagtgac 120 gcccttgatg atttcgacct ggacatgctg attaactcta gaagttccgg atctccgaaa 180 aagaaacgca aagttggtag ccagtacctg cccgacaccg acgaccggca ccggatcgag 240 gaaaagcgga agcggaccta cgagacattc aagagcatca tgaagaagtc ccccttcagc 300 ggccccaccg accctagacc tccacctaga agaatcgccg tgcccagcag atccagcgcc 360 agcgtgccaa aacctgcccc ccagccttac cccttcacca gcagcctgag caccatcaac 420 tacgacgagt tccctaccat ggtgttcccc agcggccaga tctctcaggc ctctgctctg 480 gctccagccc ctcctcaggt gctgcctcag gctcctgctc ctgcaccagc tccagccatg 540 gtgtctgcac tggctcaggc accagcaccc gtgcctgtgc tggctcctgg acctccacag 600 gctgtggctc caccagcccc taaacctaca caggccggcg agggcacact gtctgaagct 660 ctgctgcagc tgcagttcga cgacgaggat ctgggagccc tgctgggaaa cagcaccgat 720 cctgccgtgt tcaccgacct ggccagcgtg gacaacagcg agttccagca gctgctgaac 780 cagggcatcc ctgtggcccc tcacaccacc gagcccatgc tgatggaata ccccgaggcc 840 atcacccggc tcgtgacagg cgctcagagg cctcctgatc cagctcctgc ccctctggga 900 gcaccaggcc tgcctaatgg actgctgtct ggcgacgagg acttcagctc tatcgccgat 960 atggatttct cagccttgct gggctctggc agcggcagcc gggattccag ggaagggatg 1020 tttttgccga agcctgaggc cggctccgct attagtgacg tgtttgaggg ccgcgaggtg 1080 tgccagccaa aacgaatccg gccatttcat cctccaggaa gtccatgggc caaccgccca 1140 ctccccgcca gcctcgcacc aacaccaacc ggtccagtac atgagccagt cgggtcactg 1200 accccggcac cagtccctca gccactggat ccagcgcccg cagtgactcc cgaggccagt 1260 cacctgttgg aggatcccga tgaagagacg agccaggctg tcaaagccct tcgggagatg 1320 gccgatactg tgattcccca gaaggaagag gctgcaatct gtggccaaat ggacctttcc 1380 catccgcccc caaggggcca tctggatgag ctgacaacca cacttgagtc catgaccgag 1440 gatctgaacc tggactcacc cctgaccccg gaattgaacg agattctgga taccttcctg 1500 aacgacgagt gcctcttgca tgccatgcat atcagcacag gactgtccat cttcgacaca 1560 tctctgttt 1569 <210> 96 <211> 523 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 96 Asp Ala Leu Asp Asp Phe Asp Leu Asp Met Leu Gly Ser Asp Ala Leu 1 5 10 15 Asp Asp Phe Asp Leu Asp Met Leu Gly Ser Asp Ala Leu Asp Asp Phe 20 25 30 Asp Leu Asp Met Leu Gly Ser Asp Ala Leu Asp Asp Phe Asp Leu Asp 35 40 45 Met Leu Ile Asn Ser Arg Ser Ser Gly Ser Pro Lys Lys Lys Arg Lys 50 55 60 Val Gly Ser Gln Tyr Leu Pro Asp Thr Asp Asp Arg His Arg Ile Glu 65 70 75 80 Glu Lys Arg Lys Arg Thr Tyr Glu Thr Phe Lys Ser Ile Met Lys Lys 85 90 95 Ser Pro Phe Ser Gly Pro Thr Asp Pro Arg Pro Pro Pro Arg Arg Ile 100 105 110 Ala Val Pro Ser Arg Ser Ser Ala Ser Val Pro Lys Pro Ala Pro Gln 115 120 125 Pro Tyr Pro Phe Thr Ser Ser Leu Ser Thr Ile Asn Tyr Asp Glu Phe 130 135 140 Pro Thr Met Val Phe Pro Ser Gly Gln Ile Ser Gln Ala Ser Ala Leu 145 150 155 160 Ala Pro Ala Pro Pro Gln Val Leu Pro Gln Ala Pro Ala Pro Ala Pro 165 170 175 Ala Pro Ala Met Val Ser Ala Leu Ala Gln Ala Pro Ala Pro Val Pro 180 185 190 Val Leu Ala Pro Gly Pro Pro Gln Ala Val Ala Pro Pro Ala Pro Lys 195 200 205 Pro Thr Gln Ala Gly Glu Gly Thr Leu Ser Glu Ala Leu Leu Gln Leu 210 215 220 Gln Phe Asp Asp Glu Asp Leu Gly Ala Leu Leu Gly Asn Ser Thr Asp 225 230 235 240 Pro Ala Val Phe Thr Asp Leu Ala Ser Val Asp Asn Ser Glu Phe Gln 245 250 255 Gln Leu Leu Asn Gln Gly Ile Pro Val Ala Pro His Thr Thr Glu Pro 260 265 270 Met Leu Met Glu Tyr Pro Glu Ala Ile Thr Arg Leu Val Thr Gly Ala 275 280 285 Gln Arg Pro Pro Asp Pro Ala Pro Ala Pro Leu Gly Ala Pro Gly Leu 290 295 300 Pro Asn Gly Leu Leu Ser Gly Asp Glu Asp Phe Ser Ser Ile Ala Asp 305 310 315 320 Met Asp Phe Ser Ala Leu Leu Gly Ser Gly Ser Gly Ser Arg Asp Ser 325 330 335 Arg Glu Gly Met Phe Leu Pro Lys Pro Glu Ala Gly Ser Ala Ile Ser 340 345 350 Asp Val Phe Glu Gly Arg Glu Val Cys Gln Pro Lys Arg Ile Arg Pro 355 360 365 Phe His Pro Pro Gly Ser Pro Trp Ala Asn Arg Pro Leu Pro Ala Ser 370 375 380 Leu Ala Pro Thr Pro Thr Gly Pro Val His Glu Pro Val Gly Ser Leu 385 390 395 400 Thr Pro Ala Pro Val Pro Gln Pro Leu Asp Pro Ala Pro Ala Val Thr 405 410 415 Pro Glu Ala Ser His Leu Leu Glu Asp Pro Asp Glu Glu Thr Ser Gln 420 425 430 Ala Val Lys Ala Leu Arg Glu Met Ala Asp Thr Val Ile Pro Gln Lys 435 440 445 Glu Glu Ala Ala Ile Cys Gly Gln Met Asp Leu Ser His Pro Pro Pro 450 455 460 Arg Gly His Leu Asp Glu Leu Thr Thr Thr Leu Glu Ser Met Thr Glu 465 470 475 480 Asp Leu Asn Leu Asp Ser Pro Leu Thr Pro Glu Leu Asn Glu Ile Leu 485 490 495 Asp Thr Phe Leu Asn Asp Glu Cys Leu Leu His Ala Met His Ile Ser 500 505 510 Thr Gly Leu Ser Ile Phe Asp Thr Ser Leu Phe 515 520 <210> 97 <211> 1939 <212> PRT <213> Artificial sequence <220> <223> Synthetic <400> 97 Met Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp Ala Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965 970 975 Glu Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 980 985 990 Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Lys Lys Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu Phe Ser Lys 1265 1270 1275 Arg Val Ile Leu Ala Asp Ala Asn Leu Asp Lys Val Leu Ser Ala 1280 1285 1290 Tyr Asn Lys His Arg Asp Lys Pro Ile Arg Glu Gln Ala Glu Asn 1295 1300 1305 Ile Ile His Leu Phe Thr Leu Thr Asn Leu Gly Ala Pro Ala Ala 1310 1315 1320 Phe Lys Tyr Phe Asp Thr Thr Ile Asp Arg Lys Arg Tyr Thr Ser 1325 1330 1335 Thr Lys Glu Val Leu Asp Ala Thr Leu Ile His Gln Ser Ile Thr 1340 1345 1350 Gly Leu Tyr Glu Thr Arg Ile Asp Leu Ser Gln Leu Gly Gly Asp 1355 1360 1365 Gly Thr Gly Gly Pro Pro Lys Lys Lys Arg Lys Val Ala Ala Ala 1370 1375 1380 Ser Arg Tyr Pro Arg Gly Asp Ala Leu Asp Asp Phe Asp Leu Asp 1385 1390 1395 Met Leu Gly Ser Asp Ala Leu Asp Asp Phe Asp Leu Asp Met Leu 1400 1405 1410 Gly Ser Asp Ala Leu Asp Asp Phe Asp Leu Asp Met Leu Gly Ser 1415 1420 1425 Asp Ala Leu Asp Asp Phe Asp Leu Asp Met Leu Ile Asn Ser Arg 1430 1435 1440 Ser Ser Gly Ser Pro Lys Lys Lys Arg Lys Val Gly Ser Gln Tyr 1445 1450 1455 Leu Pro Asp Thr Asp Asp Arg His Arg Ile Glu Glu Lys Arg Lys 1460 1465 1470 Arg Thr Tyr Glu Thr Phe Lys Ser Ile Met Lys Lys Ser Pro Phe 1475 1480 1485 Ser Gly Pro Thr Asp Pro Arg Pro Pro Pro Arg Arg Ile Ala Val 1490 1495 1500 Pro Ser Arg Ser Ser Ala Ser Val Pro Lys Pro Ala Pro Gln Pro 1505 1510 1515 Tyr Pro Phe Thr Ser Ser Leu Ser Thr Ile Asn Tyr Asp Glu Phe 1520 1525 1530 Pro Thr Met Val Phe Pro Ser Gly Gln Ile Ser Gln Ala Ser Ala 1535 1540 1545 Leu Ala Pro Ala Pro Pro Gln Val Leu Pro Gln Ala Pro Ala Pro 1550 1555 1560 Ala Pro Ala Pro Ala Met Val Ser Ala Leu Ala Gln Ala Pro Ala 1565 1570 1575 Pro Val Pro Val Leu Ala Pro Gly Pro Pro Gln Ala Val Ala Pro 1580 1585 1590 Pro Ala Pro Lys Pro Thr Gln Ala Gly Glu Gly Thr Leu Ser Glu 1595 1600 1605 Ala Leu Leu Gln Leu Gln Phe Asp Asp Glu Asp Leu Gly Ala Leu 1610 1615 1620 Leu Gly Asn Ser Thr Asp Pro Ala Val Phe Thr Asp Leu Ala Ser 1625 1630 1635 Val Asp Asn Ser Glu Phe Gln Gln Leu Leu Asn Gln Gly Ile Pro 1640 1645 1650 Val Ala Pro His Thr Thr Glu Pro Met Leu Met Glu Tyr Pro Glu 1655 1660 1665 Ala Ile Thr Arg Leu Val Thr Gly Ala Gln Arg Pro Pro Asp Pro 1670 1675 1680 Ala Pro Ala Pro Leu Gly Ala Pro Gly Leu Pro Asn Gly Leu Leu 1685 1690 1695 Ser Gly Asp Glu Asp Phe Ser Ser Ile Ala Asp Met Asp Phe Ser 1700 1705 1710 Ala Leu Leu Gly Ser Gly Ser Gly Ser Arg Asp Ser Arg Glu Gly 1715 1720 1725 Met Phe Leu Pro Lys Pro Glu Ala Gly Ser Ala Ile Ser Asp Val 1730 1735 1740 Phe Glu Gly Arg Glu Val Cys Gln Pro Lys Arg Ile Arg Pro Phe 1745 1750 1755 His Pro Pro Gly Ser Pro Trp Ala Asn Arg Pro Leu Pro Ala Ser 1760 1765 1770 Leu Ala Pro Thr Pro Thr Gly Pro Val His Glu Pro Val Gly Ser 1775 1780 1785 Leu Thr Pro Ala Pro Val Pro Gln Pro Leu Asp Pro Ala Pro Ala 1790 1795 1800 Val Thr Pro Glu Ala Ser His Leu Leu Glu Asp Pro Asp Glu Glu 1805 1810 1815 Thr Ser Gln Ala Val Lys Ala Leu Arg Glu Met Ala Asp Thr Val 1820 1825 1830 Ile Pro Gln Lys Glu Glu Ala Ala Ile Cys Gly Gln Met Asp Leu 1835 1840 1845 Ser His Pro Pro Pro Arg Gly His Leu Asp Glu Leu Thr Thr Thr 1850 1855 1860 Leu Glu Ser Met Thr Glu Asp Leu Asn Leu Asp Ser Pro Leu Thr 1865 1870 1875 Pro Glu Leu Asn Glu Ile Leu Asp Thr Phe Leu Asn Asp Glu Cys 1880 1885 1890 Leu Leu His Ala Met His Ile Ser Thr Gly Leu Ser Ile Phe Asp 1895 1900 1905 Thr Ser Leu Phe Pro Lys Lys Lys Arg Lys Val Arg Ser Lys Arg 1910 1915 1920 Pro Ala Ala Thr Lys Lys Ala Gly Gln Ala Lys Lys Lys Lys Leu 1925 1930 1935 Asp <210> 98 <211> 112 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 98 gaaacaagct atttgctgat ttgtattagg taccatagag tgaggcgagg atgaagccga 60 gaggatactg cagaggtctc tggtgcaatg tgtgtatgtg tgcgtttgtg tg 112 <210> 99 <211> 71 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 99 gaaacaagct atttgctgat ttgtattagg taccatagag tgaggcgagg atgaagccga 60 gaggatactg c 71 <210> 100 <211> 112 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 100 gaaacaagct atttgctgat ttgtattagg taccatagag tgaggcgagg atgaagccga 60 gaggatactg cagaggtctc tggtgcaatg tgtgtatgtg tgcgtttgtg tg 112 <210> 101 <211> 108 <212> DNA <213> Artificial sequence <220> <223> Synthetic <220> <221> A sequence whose biological properties cannot be described by the keywords in the feature table <222> (52)..(52) <223> n is a, c, g or t <400> 101 gaaacaagct atttgctgat ttgtattagg taccatagag tgaggcgagg angaagccga 60 gaggatactg cagaggtctc tggtgcaatg tgtgtatgtg tgcgtttg 108 <210> 102 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 102 ttccagtgtc gaatctgcat gcgcaacttc agccagcggg gaaacctggt gaggcatatc 60 cgcacccaca cgggagagaa gcct 84 <210> 103 <211> 87 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 103 tttgcctgcg atatttgtgg aaagaagttt gctctgagct tcaatctaac cagacacacc 60 aagattcata ctgggtccca gaaaccg 87 <210> 104 <211> 85 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 104 ttccagtgta ggatatgcat gaggaatttc tctcggagtg acaacttaac gcggcatata 60 aggacgcaca caggtgaaaa aacaa 85 <210> 105 <211> 87 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 105 tttgcatgcg acatctgtgg caaaaagttt gcggaccggt ctcaccttgc ccgacacaca 60 aaaatccata ccggcagtca aaagccc 87 <210> 106 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 106 tttcaatgtc gcatttgcat gcgaaacttc tcacagaagg cccatttgac tgcccatatt 60 cgtactcata ctggcgagaa acct 84 <210> 107 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 107 ttcgcttgcg atatatgtgg tcgtaagttt gcacggtcgg acaacctcac acgccacact 60 aagatacacc tgcggcagaa ggac 84 <210> 108 <211> 85 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 108 ttccagtgtc gaatctgcat gcgcaacttc agcccgaatg tccaacctga cacggcatat 60 ccgcacccac acgggagaga agcct 85 <210> 109 <211> 87 <212> DNA <213> Artificial sequence <220> <223> Synthesis <400> 109 tttgcctgcg atatttgtgg aaagaagttt gctgacaagc ggaccttaat ccgccacacc 60 aagattcata ctgggtccca gaaaccg 87 <210> 110 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 110 ttccagtgta ggatatgcat gaggaatttc tctcagcggg gaaatctagt gcgacatata 60 aggacgcaca caggtgaaaa acca 84 <210> 111 <211> 87 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 111 tttgcatgcg acatctgtgg caaaaagttt gcgctgagct tcaacttgac tcgtcacaca 60 aaaatccata ccggcagtca aaagccc 87 <210> 112 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 112 tttcaatgtc gcatttgcat gcgaaacttc tcacggagtg acaatcttac gagacatatt 60 cgtactcata ctggcgagaa acct 84 <210> 113 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 113 ttcgcttgcg atatatgtgg tcgtaagttt gcagaccgga gccacttagc caggcacact 60 aagatacacc tgcggcagaa ggac 84 <210> 114 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 114 ttccagtgtc gaatctgcat gcgcaacttc agcgaccgga gcgcgctggc acggcatatc 60 cgcacccaca cgggagagaa gcct 84 <210> 115 <211> 87 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 115 tttgcctgcg atatttgtgg aaagaagttt gctcgaagtg acaacttaac gcgccacacc 60 aagattcata ctgggtccca gaaaccg 87 <210> 116 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 116 ttccagtgta ggatatgcat gaggaatttc tctcagtcag gggacctcac tcgtcatata 60 aggacgcaca caggtgaaaa acca 84 <210> 117 <211> 87 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 117 tttgcatgcg acatctgtgg caaaaagttt gcggtacgac agacgcttaa acaacacaca 60 aaaatccata ccggcagtca aaagccc 87 <210> 118 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 118 tttcaatgtc gcatttgcat gcgaaacttc tcagccgctg gtaacttgac acgacatatt 60 cgtactcata ctggcgagaa acct 84 <210> 119 <211> 84 <212> DNA <213> Artificial sequence <220> <223> Synthetic <400> 119 ttcgcttgcg atatatgtgg tcgtaagttt gcaagatctg ataatctaac gcgtcacact 60 aagatacacc tgcggcagaa ggac 84 <210> 120 <211> 5 <212> PRT <213> Artificial sequence <220> <223> Synthetic polypeptide <400> 120 Thr Gly Glu Lys Pro 1 5 <210> 121 <211> 6 <212> PRT <213> Artificial sequence <220> <223> Synthetic polypeptide <400> 121 Thr Gly Ser Gln Lys Pro 1 5
Claims
1. A isolated nucleic acid comprising a transgene configured to express at least one DNA binding domain fused to at least one transcriptional regulator domain, wherein the at least one DNA binding domain binds to a target gene or a regulatory region of a target gene, and wherein the target gene encodes a voltage-gated sodium channel, and wherein the at least one DNA binding domain is a zinc finger protein comprising recognition helices 1-6 selected from: (a) Recognition helix 1: QRGNLVR (SEQ ID NO:17), recognition helix 2: LSFNLTR (SEQ ID NO:18), recognition helix 3: RSDNLTR (SEQ ID NO:19), recognition helix 4: DRSHLAR (SEQ ID NO:20), recognition helix 5: QKAHLTA (SEQ ID NO:21), recognition helix 6: RSDNLTR (SEQ ID NO:22); (b) Recognition helix 1: RSSNLTR (SEQ ID NO:29), recognition helix 2: DKRTLIR (SEQ ID NO:30), recognition helix 3: QRGNLVR (SEQ ID NO:31), recognition helix 4: LSFNLTR (SEQ ID NO:32), recognition helix 5: RSDNLTR (SEQ ID NO:33), recognition helix 6: DRSHLAR (SEQ ID NO:34); or (c) Recognition helix 1: DRSALAR (SEQ ID NO:41), recognition helix 2: RSDNLTR (SEQ ID NO:42), recognition helix 3: QSGDLTR (SEQ ID NO:43), recognition helix 4: VRQTLKQ (SEQ ID NO:44), recognition helix 5: AAGNLTR (SEQ ID NO:45), recognition helix 6: RSDNLTR (SEQ ID NO:46).
2. The isolated nucleic acid according to claim 1, wherein the transgene is flanked by inverted terminal repeats (ITRs) derived from adeno-associated virus (AAV).
3. The isolated nucleic acid according to claim 1, wherein the transcriptional regulator domain upregulates the expression of the target gene.
4. The isolated nucleic acid according to claim 1, wherein the at least one DNA binding domain is a zinc finger protein having at least 95% sequence identity with any one of SEQ ID NO:57, 59 and 61.
5. The isolated nucleic acid according to claim 1, wherein the at least one DNA binding domain is a zinc finger protein having at least 97% sequence identity with any one of SEQ ID NO:57, 59 and 61.
6. The isolated nucleic acid according to claim 1, wherein the at least one DNA binding domain is a zinc finger protein having at least 98% or at least 99% sequence identity with any one of SEQ ID NO:57, 59 and 61.
7. The isolated nucleic acid according to claim 1, wherein the at least one DNA-binding domain is a zinc finger protein having 100% sequence identity with any one of SEQ ID NOs: 57, 59, and 61.
8. The isolated nucleic acid according to any one of claims 1-7, wherein the at least one DNA-binding domain is a zinc finger protein encoded by a nucleic acid having the sequence shown in SEQ ID NO: 58, 60, or 62.
9. The isolated nucleic acid according to any one of claims 1-7, wherein the at least one DNA-binding domain is a zinc finger protein, and the zinc finger protein comprises recognition helices 1-6 selected from the following: (a) Recognition helix 1 encoded by SEQ ID NO: 11, recognition helix 2 encoded by SEQ ID NO: 12, recognition helix 3 encoded by SEQ ID NO: 13, recognition helix 4 encoded by SEQ ID NO: 14, recognition helix 5 encoded by SEQ ID NO: 15, recognition helix 6 encoded by SEQ ID NO: 16; (b) Recognition helix 1 encoded by SEQ ID NO: 23, recognition helix 2 encoded by SEQ ID NO: 24, recognition helix 3 encoded by SEQ ID NO: 25, recognition helix 4 encoded by SEQ ID NO: 26, recognition helix 5 encoded by SEQ ID NO: 27, recognition helix 6 encoded by SEQ ID NO: 28; (c) Recognition helix 1 encoded by SEQ ID NO: 35, recognition helix 2 encoded by SEQ ID NO: 36, recognition helix 3 encoded by SEQ ID NO: 37, recognition helix 4 encoded by SEQ ID NO: 38, recognition helix 5 encoded by SEQ ID NO: 39, recognition helix 6 encoded by SEQ ID NO:
40.
10. The isolated nucleic acid according to claim 9, wherein the at least one DNA-binding domain is a zinc finger protein, and the zinc finger protein comprises the recognition helix encoded by the nucleic acid of SEQ ID NO: 11, the recognition helix encoded by the nucleic acid of SEQ ID NO: 12, the recognition helix encoded by the nucleic acid of SEQ ID NO: 13, the recognition helix encoded by the nucleic acid of SEQ ID NO: 14, the recognition helix encoded by the nucleic acid of SEQ ID NO: 15, and / or the recognition helix encoded by the nucleic acid of SEQ ID NO:
16.
11. The isolated nucleic acid according to claim 9, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises a recognition helix encoded by the nucleic acid of SEQ ID NO: 23, a recognition helix encoded by the nucleic acid of SEQ ID NO: 24, a recognition helix encoded by the nucleic acid of SEQ ID NO: 25, a recognition helix encoded by the nucleic acid of SEQ ID NO: 26, a recognition helix encoded by the nucleic acid of SEQ ID NO: 27, and / or a recognition helix encoded by the nucleic acid of SEQ ID NO:
28.
12. The isolated nucleic acid according to claim 9, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises a recognition helix encoded by the nucleic acid of SEQ ID NO: 35, a recognition helix encoded by the nucleic acid of SEQ ID NO: 36, a recognition helix encoded by the nucleic acid of SEQ ID NO: 37, a recognition helix encoded by the nucleic acid of SEQ ID NO: 38, a recognition helix encoded by the nucleic acid of SEQ ID NO: 39, and / or a recognition helix encoded by the nucleic acid of SEQ ID NO:
40.
13. The isolated nucleic acid according to claim 1, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises a recognition helix consisting of SEQ ID NO: 17, a recognition helix consisting of SEQ ID NO: 18, a recognition helix consisting of SEQ ID NO: 19, a recognition helix consisting of SEQ ID NO: 20, a recognition helix consisting of SEQ ID NO: 21, and / or a recognition helix consisting of SEQ ID NO:
22.
14. The isolated nucleic acid according to claim 1, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises a recognition helix consisting of SEQ ID NO: 29, a recognition helix consisting of SEQ ID NO: 30, a recognition helix consisting of SEQ ID NO: 31, a recognition helix consisting of SEQ ID NO: 32, a recognition helix consisting of SEQ ID NO: 33, and / or a recognition helix consisting of SEQ ID NO:
34.
15. The isolated nucleic acid according to claim 1, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises a recognition helix consisting of SEQ ID NO: 41, a recognition helix consisting of SEQ ID NO: 42, a recognition helix consisting of SEQ ID NO: 43, a recognition helix consisting of SEQ ID NO: 44, a recognition helix consisting of SEQ ID NO: 45, and / or a recognition helix consisting of SEQ ID NO:
46.
16. The isolated nucleic acid according to any one of claims 1-7, wherein the at least one transcriptional regulator domain comprises VPR, Rta, p65 or Hsf1 transactivator or any combination thereof.
17. The isolated nucleic acid according to any one of claims 1-7, wherein the at least one transcriptional regulator domain is encoded by the nucleic acid sequence set forth in SEQ ID NO:
47.
18. The isolated nucleic acid according to any one of claims 1-7, wherein the at least one transcriptional regulator domain comprises the amino acid sequence set forth in SEQ ID NO:
48.
19. The isolated nucleic acid according to claim 2, wherein the ITR is AAV2 ITR.
20. The isolated nucleic acid according to claim 19, wherein the ITR is ΔTR and / or mTR.
21. The isolated nucleic acid according to any one of claims 1-7, wherein the transgene is operably linked to a promoter.
22. The isolated nucleic acid according to claim 21, wherein the promoter is a tissue-specific promoter.
23. The isolated nucleic acid according to claim 22, wherein the promoter is a neuronal promoter.
24. The isolated nucleic acid according to claim 23, wherein the promoter is selected from the promoters of phospho-activated glutaminase (PAG), vesicular glutamate transporter-1 (VGLUT1), glutamate decarboxylase 65 and 67 (GAD65, GAD67), synapsin I, a-CamKII, Dock10, Prox1, parvalbumin (PV), somatostatin (SST), cholecystokinin (CCK), calretinin (CR) or neuropeptide Y (NPY).
25. The isolated nucleic acid according to any one of claims 1-7, wherein the at least one DNA binding domain is fused to the at least one transcriptional regulator domain via a linker domain.
26. The isolated nucleic acid according to claim 25, wherein the linker domain is: (i) a flexible linker, or (ii) a cleavable linker.
27. The isolated nucleic acid according to claim 26, wherein the flexible linker comprises glycine.
28. The isolated nucleic acid according to any one of claims 1-7, wherein the transgene encodes 1 DNA binding domain, 2 DNA binding domains, 3 DNA binding domains, 4 DNA binding domains, 5 DNA binding domains, 6 DNA binding domains, 7 DNA binding domains, 8 DNA binding domains, 9 DNA binding domains or 10 DNA binding domains.
29. The isolated nucleic acid according to any one of claims 1-7, wherein the transgene encodes 1 transcriptional regulator domain, 2 transcriptional regulator domains, 3 transcriptional regulator domains, 4 transcriptional regulator domains, 5 transcriptional regulator domains, 6 transcriptional regulator domains, 7 transcriptional regulator domains, 8 transcriptional regulator domains, 9 transcriptional regulator domains or 10 transcriptional regulator domains.
30. The isolated nucleic acid according to any one of claims 1-7, wherein the at least one DNA binding domain binds to a nucleic acid of the sequence shown in any one of SEQ ID NOs: 5-7.
31. A recombinant adeno-associated virus (rAAV) comprising: (i) a nucleic acid comprising a transgene encoding at least one DNA-binding domain fused to at least one transcriptional regulator domain, wherein the at least one DNA-binding domain binds to a target gene or a regulatory region of a target gene, wherein the target gene encodes a voltage-gated sodium channel, and wherein the at least one DNA-binding domain is a zinc finger protein comprising recognition helices 1-6 selected from: (a) recognition helix 1: QRGNLVR (SEQ ID NO:17), recognition helix 2: LSFNLTR (SEQ ID NO:18), recognition helix 3: RSDNLTR (SEQ ID NO: 19), recognition helix 4: DRSHLAR (SEQ ID NO:20), recognition helix 5: QKAHLTA (SEQ ID NO:21), recognition helix 6: RSDNLTR (SEQ ID NO: 22); (b) recognition helix 1: RSSNLTR (SEQ ID NO:29), recognition helix 2: DKRTLIR (SEQ ID NO:30), recognition helix 3: QRGNLVR (SEQ ID NO: 31), recognition helix 4: LSFNLTR (SEQ ID NO:32), recognition helix 5: RSDNLTR (SEQ ID NO:33), recognition helix 6: DRSHLAR (SEQ ID NO: 34); or (c) recognition helix 1: DRSALAR (SEQ ID NO:41), recognition helix 2: RSDNLTR (SEQ ID NO:42), recognition helix 3: QSGDLTR (SEQ ID NO: 43), recognition helix 4: VRQTLKQ (SEQ ID NO:44), recognition helix 5: AAGNLTR (SEQ ID NO:45), recognition helix 6: RSDNLTR (SEQ ID NO: 46); and (ii) at least one capsid protein.
32. The rAAV of claim 31, wherein the transgene is flanked by inverted terminal repeats (ITRs) derived from adeno-associated virus (AAV).
33. The rAAV of claim 31, wherein the at least one transcriptional regulator domain upregulates the expression of the target gene.
34. The rAAV of claim 31, wherein the at least one DNA-binding domain is a zinc finger protein having at least 95% sequence identity to any one of SEQ ID NO:57, 59, and 61.
35. The rAAV of claim 31, wherein the at least one DNA-binding domain is a zinc finger protein having at least 97% sequence identity to any one of SEQ ID NO:57, 59, and 61.
36. The rAAV according to claim 31, wherein the at least one DNA binding domain is a zinc finger protein having at least 98% or at least 99% sequence identity to any one of SEQ ID NOs: 57, 59, and 61.
37. The rAAV according to claim 31, wherein the at least one DNA binding domain is a zinc finger protein having 100% sequence identity to any one of SEQ ID NOs: 57, 59, and 61.
38. The rAAV according to any one of claims 31-37, wherein the at least one DNA binding domain is a zinc finger protein encoded by a nucleic acid having the sequence shown in SEQ ID NO: 58, 60, or 62.
39. The rAAV according to any one of claims 31-37, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises recognition helices 1-6 selected from the following: (a) Recognition helix 1 encoded by SEQ ID NO: 11, recognition helix 2 encoded by SEQ ID NO: 12, recognition helix 3 encoded by SEQ ID NO: 13, recognition helix 4 encoded by SEQ ID NO: 14, recognition helix 5 encoded by SEQ ID NO: 15, recognition helix 6 encoded by SEQ ID NO: 16; (b) Recognition helix 1 encoded by SEQ ID NO: 23, recognition helix 2 encoded by SEQ ID NO: 24, recognition helix 3 encoded by SEQ ID NO: 25, recognition helix 4 encoded by SEQ ID NO: 26, recognition helix 5 encoded by SEQ ID NO: 27, recognition helix 6 encoded by SEQ ID NO: 28; (c) Recognition helix 1 encoded by SEQ ID NO: 35, recognition helix 2 encoded by SEQ ID NO: 36, recognition helix 3 encoded by SEQ ID NO: 37, recognition helix 4 encoded by SEQ ID NO: 38, recognition helix 5 encoded by SEQ ID NO: 39, recognition helix 6 encoded by SEQ ID NO:
40.
40. The rAAV according to claim 39, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises the recognition helix encoded by the nucleic acid of SEQ ID NO: 11, the recognition helix encoded by the nucleic acid of SEQ ID NO: 12, the recognition helix encoded by the nucleic acid of SEQ ID NO: 13, the recognition helix encoded by the nucleic acid of SEQ ID NO: 14, the recognition helix encoded by the nucleic acid of SEQ ID NO: 15, and / or the recognition helix encoded by the nucleic acid of SEQ ID NO:
16.
41. The rAAV according to claim 39, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises a recognition helix encoded by the nucleic acid of SEQ ID NO: 23, a recognition helix encoded by the nucleic acid of SEQ ID NO: 24, a recognition helix encoded by the nucleic acid of SEQ ID NO: 25, a recognition helix encoded by the nucleic acid of SEQ ID NO: 26, a recognition helix encoded by the nucleic acid of SEQ ID NO: 27, and / or a recognition helix encoded by the nucleic acid of SEQ ID NO:
28.
42. The rAAV according to claim 39, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises a recognition helix encoded by the nucleic acid of SEQ ID NO: 35, a recognition helix encoded by the nucleic acid of SEQ ID NO: 36, a recognition helix encoded by the nucleic acid of SEQ ID NO: 37, a recognition helix encoded by the nucleic acid of SEQ ID NO: 38, a recognition helix encoded by the nucleic acid of SEQ ID NO: 39, and / or a recognition helix encoded by the nucleic acid of SEQ ID NO:
40.
43. The rAAV according to claim 31, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises a recognition helix composed of SEQ ID NO: 17, a recognition helix composed of SEQ ID NO: 18, a recognition helix composed of SEQ ID NO: 19, a recognition helix composed of SEQ ID NO: 20, a recognition helix composed of SEQ ID NO: 21, and / or a recognition helix composed of SEQ ID NO:
22.
44. The rAAV according to claim 31, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises a recognition helix composed of SEQ ID NO: 29, a recognition helix composed of SEQ ID NO: 30, a recognition helix composed of SEQ ID NO: 31, a recognition helix composed of SEQ ID NO: 32, a recognition helix composed of SEQ ID NO: 33, and / or a recognition helix composed of SEQ ID NO:
34.
45. The rAAV according to claim 31, wherein the at least one DNA binding domain is a zinc finger protein, and the zinc finger protein comprises a recognition helix composed of SEQ ID NO: 41, a recognition helix composed of SEQ ID NO: 42, a recognition helix composed of SEQ ID NO: 43, a recognition helix composed of SEQ ID NO: 44, a recognition helix composed of SEQ ID NO: 45, and / or a recognition helix composed of SEQ ID NO:
46.
46. The rAAV according to any one of claims 31-37, wherein the at least one transcriptional regulator domain is a transactivator derived from VPR, Rta, p65, Hsf1, or any combination thereof.
47. The rAAV according to any one of claims 31-37, wherein the transgene encoding the at least one transcriptional regulatory domain comprises the sequence set forth in SEQ ID NO:
47.
48. The rAAV according to any one of claims 31-37, wherein the at least one transcriptional regulatory domain comprises the amino acid sequence set forth in SEQ ID NO:
48.
49. The rAAV according to any one of claims 31-37, wherein the at least one DNA binding domain is fused to the at least one transcriptional regulatory domain via a linker domain.
50. The rAAV according to claim 49, wherein the linker domain is: (i) a flexible linker, or (ii) a cleavable linker.
51. The rAAV according to claim 50, wherein the flexible linker comprises glycine.
52. The rAAV according to any one of claims 31-37, wherein the transgene encodes 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 DNA binding domains.
53. The rAAV according to any one of claims 31-37, wherein the transgene encodes 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 transcriptional regulatory domains.
54. The rAAV according to any one of claims 31-37, wherein the rAAV capsid serotype is selected from the group consisting of: AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAVrh8, AAV9, AAV10, AAVrh10, and AAV.PHPB.
55. The rAAV according to any one of claims 31-37, wherein the rAAV capsid serotype is AAV9.
56. The rAAV according to any one of claims 31-37, wherein the rAAV capsid serotype is AAV.PHBP.
57. The rAAV according to claim 32, wherein the ITR is AAV2 ITR.
58. The rAAV according to claim 57, wherein the ITR is ΔTR and / or mTR.
59. The rAAV according to any one of claims 31-37, wherein the transgene is operably linked to a promoter.
60. The rAAV according to claim 59, wherein the promoter is a tissue-specific promoter.
61. The rAAV according to claim 60, wherein the tissue-specific promoter is a neuronal promoter.
62. The rAAV according to claim 61, wherein the tissue-specific promoter is selected from the promoters of phospho-activated glutaminase (PAG), vesicular glutamate transporter-1 (VGLUT1), glutamate decarboxylase 65 and 67 (GAD65, GAD67), synapsin I, a-CamKII, Dock10, Prox1, parvalbumin (PV), somatostatin (SST), cholecystokinin (CCK), calretinin (CR), or neuropeptide Y (NPY).
63. The rAAV according to any one of claims 31-37, wherein the at least one DNA binding domain binds to a nucleic acid of a sequence shown in any one of SEQ ID NOs: 5-7.
64. Use of the isolated nucleic acid according to any one of claims 1-30 or the rAAV according to any one of claims 31-63 in the preparation of a medicament for increasing the expression of a target gene in a cell or a subject containing the target gene, wherein the target gene is SCN1A.
65. The use according to claim 64, wherein the cell or the subject is haploinsufficient for the target gene.
66. The use according to claim 64, wherein the cell is a neuron.
67. The use according to claim 64, wherein the cell is a GABAergic neuron.
68. A composition comprising the isolated nucleic acid according to any one of claims 1-30 or the rAAV according to any one of claims 31-63.
69. The composition according to claim 68, which further comprises a pharmaceutically acceptable carrier.
70. A kit comprising: A container that houses the isolated nucleic acid according to any one of claims 1-30 or the rAAV according to any one of claims 31-63.
71. The kit according to claim 70, wherein the kit further comprises a container that houses a pharmaceutically acceptable carrier.
72. The kit according to claim 71, wherein the isolated nucleic acid or the rAAV and the pharmaceutically acceptable carrier are housed in the same container.
73. The kit according to claim 70, wherein the container is a syringe.
74. A host cell comprising the isolated nucleic acid according to any one of claims 1-30 or the rAAV according to any one of claims 31-63.
75. The host cell according to claim 74, wherein the host cell is a eukaryotic cell.
76. The host cell according to claim 74, wherein the host cell is a mammalian cell.
77. The host cell according to claim 74, wherein the host cell is a human cell.
78. The host cell according to claim 74, wherein the host cell is a neuron.
79. The host cell according to claim 74, wherein the host cell is a GABAergic neuron.
Citation Information
Patent Citations
Gap variable stranded wire coating system and its combined guide roller
CN1152751C
Synchronization device
CN1299019A
Assembled concrete beam making bench in site
CN2486596Y
Dry-washing / water-washing two-purpose washing machine
CN2493636Y
Mutually-acting anti-theft safety system
CN2568488Y