Human t-cell lymphotropic virus type 1 targeting proteins and methods of use
Zinc finger proteins targeting the HTLV-I LTR effectively downregulate the HBZ gene, addressing the limitations of current treatments for HTLV-I diseases by inhibiting CCR4 and inducing apoptosis in infected cells, offering a novel therapeutic strategy.
Patent Information
- Application Number
- US18/854481
- Authority / Receiving Office
- US · United States
- Patent Type
- Applications(United States)
- Current Assignee / Owner
- Priority Date
- 2022-04-06
- Filing Date
- 2023-04-05
- Publication Date
- 2025-11-13
AI Technical Summary
Current treatments for Human T-lymphotropic virus type I (HTLV-I) associated diseases, such as acute T-cell leukemia/lymphoma (ATL), are ineffective, with no vaccine or commercially available therapy, and existing monoclonal antibody treatments show limited success, particularly in ATLs with CCR4 mutations.
Development of zinc finger proteins capable of binding to the long terminal repeat (LTR) of HTLV-I, specifically targeting the anti-sense HTLV-1 bZIP factor (HBZ) gene to downregulate its expression, using vectors, extracellular vesicles, and pharmaceutical compositions to deliver these proteins to infected cells.
The zinc finger proteins effectively inhibit HBZ expression, reducing CCR4 levels, inducing cell cycle arrest, apoptosis, and inhibiting proliferation of HTLV-I infected cells, providing a potential therapeutic approach for HTLV-I associated diseases.
Smart Images

Figure US20250346636A1-D00000_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 328,108, filed Apr. 6, 2022, which is hereby incorporated by reference in its entirety and for all purposes.STATEMENT AS TO RIGHTS TO INVENTIONS MADE UNDER FEDERALLY SPONSORED RESEARCH AND DEVELOPMENT
[0002] This invention was made with government support under RO1 M113407 awarded by the National Institutes of Health. The government has certain rights in the invention.REFERENCE TO A SEQUENCE LISTING, A TABLE OR A COMPUTER PROGRAM LISTING APPENDIX SUBMITTED AS AN ASCII TEXT FILE
[0003] The contents of the electronic sequence listing (048440-836001WO_ST26.xml; Size 133,584 bytes; and Date of Creation: Mar. 20, 2023) is hereby incorporated by reference in its entirety.BACKGROUND
[0004] Human T-lymphotropic virus type I (HTLV-I), a retrovirus, is transmitted by bodily fluids and establishes a life-long infection in patients. The virus infects primarily CD4+ T-cells in which the reverse transcribed genome integrates within the host cell to form a provirus. Viruses are predicted to cause about 15% of known cancers world-wide (1), and HTLV-I is the established etiological agent involved in the development of a group of blood-borne malignances. Through a complex interplay between viral factors over an extended incubation time, the virus has been linked to the transformation of CD4+ T-cells into a tumor state, resulting in acute T-cell leukemia / lymphoma (ATL). In its most aggressive form, acute ATL, the prognosis for the overall survival rate is ˜9 months. There remains no vaccine or treatment for HTLV-I, and, furthermore, ATL is refractory to chemotherapy and radiation therapy with no effective, commercially available alternative cancer treatment. The C-C Motif Chemokine Receptor 4 (CCR4) is upregulated on the surface of most ATLs (2), and a monoclonal antibody, mogamulizumab, has been used in clinical trials in CCR4-positive ATL patients with limited improvement in disease outcomes (3). However, a sub-class of ATLs with gain-of-function CCR4 mutations substantially improved the antibody's treatment response (4). Nonetheless, the overall lack of effective approaches to inhibit ATL urges the development of novel therapeutic strategies.
[0005] HTLV-I has ˜9 kb genome flanked by long terminal repeats (LTRs) at the 5′ and 3′ ends that serve as promoters to drive sense and anti-sense expression, respectively. The HTLV-I transactivator protein Tax is expressed from the 5′ LTR, along with other accessory and structural genes involved in productive viral replication, and is a well-established factor in clonal expansion and oncogenic transformation (5). However, Tax is highly immunogenic resulting in cytotoxic CD8+ T-cell clearance of Tax-positive cells, and in ATL is generally lowly expressed or silent as a result of gene mutation, 5′LTR truncation, or promoter epigenetic hypermethylation (6).
[0006] Recently, the anti-sense HTLV-1 bZIP factor (HBZ) gene expressed from the 3′LTR has been realized as playing an underappreciated role in oncogenesis as it suppresses apoptosis (7), induces genetic instability (8), and results in T-cell lymphomas in HBZ transgenic mice (9). Importantly, the HBZ RNA and protein have been implicated in various proliferative and pathological roles in ATL (10), such as the up-regulation of CCR4 that augments the tumor's migration and proliferation (11). Furthermore, all primary ATL samples are positive for HBZ expression (12), and the selective inhibition of HBZ reduced proliferation in a range of HTLV-I cell lines (13,14), presenting a potential common molecular target for cancer intervention.
[0007] Provided herein, inter alia, are solutions to these and other problems in the art.BRIEF SUMMARY
[0008] Provided herein, inter alia, are proteins including zinc finger domains capable of binding a sequence within the long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-I). The proteins provided herein including embodiments thereof are contemplated to be effective for downregulating expression of the HTLV-1 bZIP factor (HBZ) gene. Applicant has further discovered that proteins provided herein including embodiments thereof may be effective for treating and / or preventing HTLV-1 associated diseases (e.g. adult T-cell leukemia, etc.). Thus, in an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:27.
[0009] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:25.
[0010] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:28.
[0011] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:32.
[0012] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:31.
[0013] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:30.
[0014] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:24
[0015] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:26
[0016] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:29.
[0017] In another aspect a nucleic acid encoding the protein provided herein including embodiments thereof is provided.
[0018] In an aspect a vector including the nucleic acid provided herein including embodiments thereof is provided.
[0019] In another aspect is provided an extracellular vesicle (EV) including a nucleic acid encoding the protein provided herein including embodiments thereof.
[0020] In an aspect is provided a pharmaceutical composition including the protein provided herein including embodiments thereof, the nucleic acid provided herein including embodiments thereof, the vector provided herein including embodiments thereof, or the EV provided herein including embodiments thereof.
[0021] In another aspect is provided a cell including the protein provided herein including embodiments thereof, the nucleic acid provided herein including embodiments thereof, the vector provided herein including embodiments thereof, or the EV provided herein including embodiments thereof.
[0022] In another aspect is provided a method of treating a human T-cell lymphotropic virus type 1 (HTLV-1) associated disease in a subject in need thereof, including administering to the subject an effective amount of the protein provided herein including embodiments thereof, the nucleic acid provided herein including embodiments thereof, the vector provided herein including embodiments thereof, or the EV provided herein including embodiments thereof.BRIEF DESCRIPTION OF THE DRAWINGS
[0023] FIG. 1: Schematic of the HTLV-I genome and ZFP target sites. The 5′ LTR and 3′ LTRs flank the ˜9 kb integrated HTLV-I genome and the 3′ LTR drives the expression of the anti-sense HBZ gene. The representative target sites of a series of ZFP within the LTR are indicated (arrows, ZFP2 to ZFP10). Transcription factor SpI binding sites, the transcription start site (TSS) in the 3′ LTR, and the HBZ coding sequence are as labeled.
[0024] FIGS. 2A-2E: Screening of ZFP repressors that inhibit HTLV-1 LTR expression. (FIG. 2A) HEK293 cells were transfected with a vector that contains a HTLV-1 LTR bidirectionally driving the expression Rluc (anti-sense) and Fluc (sense) luciferase. A mutated Rluc translational start ensures that expression of Rluc only occurs if the 5′ HBZ sequence within the LTR is spliced onto the reporter. A series of HTLV-I ZFP-KRAB repressors (2-10) were transfected with the reporter vector and 48 hrs post-transfection the levels of luciferase were determined. (FIG. 2B) HEK293 cells were transfected with a vector containing the HTLV-1 3′-LTR driving the expression of the HBZ-3×FLAG with the ZFP vectors, and 48 hrs post-transfection the levels of HBZ RNA were assessed. Both spliced (HBZsp) and unspliced (e.g. nascent) HBZ RNA (HBZusp) was detected. For (FIG. 2A) and (FIG. 2B), error bars represent standard deviation from samples treated in triplicate from two independent experiments. The levels of luciferase or HBZ RNA was made relative to a ZFP-HIV-KRAB control, set a 100%. (FIG. 2C) HEK293 cells were transfected as described in (FIG. 2B) and the HBZ-3×FLAG and ZFPs were detected through their Flag and myc tags, respectively. A Rluc expression vector or untreated cells (mock) were included as ZFP and HBZ detection controls, respectively. Alpha-tubulin was detected as a loading control. The RNA levels were determined for (FIG. 2D) spliced (HBZsp) and nascent HBZ RNA (HBZusp), and (FIG. 2E) KRAB, ZFP3, or ZFP5.
[0025] FIGS. 3A-3B: Anti-proliferative effects of the anti-HBZ ZFP repressors. TL-Om1 cells were electroporated with an (FIG. 3A) 2 μg ‘low’ dose or (FIG. 3B) 4 μg ‘high’ dose of mRNA expressing the ZFP5-KRAB or ZFP5-KRAB-meCP2 and outgrowth was assessed up to day 21 through proliferation (top panel), viability (middle panel), or cell count (bottom panel). The ZFP-HIV-KRAB or GFP mRNAs were included as negative controls. Error bars represent standard deviation from samples treated in triplicate.
[0026] FIGS. 4A-4C: Anti-HTLV-I ZFPs reduce HBZ-induced CCR4 levels. TL-Om1 cells were electroporated with 2 μg of ZFP5-KRAB or ZFP5-KRAB-meCP2 mRNA, and the levels of (FIG. 4A) HBZ spliced RNA, (FIG. 4B) CCR4 RNA, (FIG. 4C) or surface CCR4 receptor was assessed at 24 hrs and 48 hrs post-electroporation. Cells treated with a ZFP-HIV-KRAB mRNA or untreated cells (mock) were included as negative controls. Error bars represent standard deviation from samples treated in triplicate and p-values were determined by one-way ANOVA analysis (Dunnett's post-test) when compared to the ZFP-HIV-control (*p<0.05, **p<0.01 ***p<0.001, ****p<0.0001).
[0027] FIGS. 5A-5D: Anti-HBZ ZFPs cause cell cycle arrest and apoptosis. (FIG. 5A) TL-Om1 cells were electroporated with 2 μg of mRNA expressing the ZFP5-KRAB or ZFP5-KRAB-meCP2 and the percentage of cell cycle phase was assessed at 24 hrs post-electroporation. (FIG. 5B) The levels of E2F1 mRNA were assessed at 24 hrs and 48 hrs post-electroporation. Cells treated with a ZFP-HIV-KRAB mRNA or untreated (mock) were included as negative controls. For (FIG. 5C), the samples were made relative to the ZFP-HIV-KRAB set at 100%. To assess the induction of apoptosis, TL-Om1 cells were electroporated with a (FIG. 5C) 2 μg ‘low’ dose or (FIG. 5D) 4 μg ‘high’ dose of mRNA and Annexin V and PI detected at 48 hrs and 72 hrs post-electroporation. ZFP-HIV-KRAB was used as a negative control. For (FIG. 5A) and (FIG. 5B), error bars represent standard deviation from samples treated in triplicate. For (FIG. 5C) and (FIG. 5D), the line represents the mean from samples treated in triplicate. The p-values were determined by one-way ANOVA analysis (Dunnett's post-test) when compared against ZFP-HIV-control (*p<0.05, **p<0.01 ***p<0.001, ****p<0.0001).
[0028] FIGS. 6A-6B: Anti-HTLV-I ZFP repressors inhibit the LTRs from multiple HTLV-I genotypes. (FIG. 6A) A schematic of the vector that contains a HTLV-1 LTR bidirectionally driving the expression Rluc (anti-sense) and Fluc (sense) luciferase. The LTR upstream of the HBZ start was replaced with sequences from different HTLV-I genotypes (a-g). The country of origins, accession numbers, genotypes, and ZFP5 target site sequences are indicated. Mismatches are in bold. (FIG. 6B) HEK293 cells were transfected with an LTR(a-g) spliced reporter vector with the ZFP5-KRAB and ZFP5-KRAB-meCP2 vectors, and 48 hrs post-transfection the levels of luciferase was determined. Error bars represent standard deviation from samples treated in triplicate. The levels of luciferase were made relative to a ZFP-HIV-KRAB control set a 100%.
[0029] FIGS. 7A-7D: Verification of HTLV-1 ZFP repressor activity and expression. (FIG. 7A) Schematic of the ZFP expression vector. CMV=cytomegalovirus promoter, NLS=nuclear localization signal, KRAB=kruppel-associated box, PA=polyA transcription terminator. Generic (KRAB) or ZFP specific (ZFP3 / 5) primer binding sites for detection of the expressed ZFP RNA are indicated. (FIG. 7B) HEK293 cells were transfected with a vector that contains a HTLV-1 LTR bidirectionally driving the expression Rluc (anti-sense) and Fluc (sense) luciferase. A series of HTLV-I ZFP-KRAB (2-10) were transfected with the reporter vector and 48 hrs post-transfection the levels of luciferase were determined. (FIG. 7C, FIG. 7D) HEK293 were transfected with a vector containing the HTLV-I 3′-LTR driving the expression of the HBZ-3×FLAG with the ZFP expression vectors, and at 48 hrs post-transfection the levels of HBZ RNA were assessed. (FIG. 7C) Both spliced (HBZsp), unspliced HBZ RNA (HBZusp), (FIG. 7D) KRAB, or ZFP3, ZFP5, RNA was determined. For (FIG. 7B-7D), error bars represent standard deviation from samples treated in triplicate from two independent experiments. For (FIG. 7B), the levels of luciferase or HBZ RNA were made relative to a ZFP-HIV-KRAB control set a 100%.
[0030] FIGS. 8A-8C: Assessing anti-HTLV-I DNA vectors for anti-proliferative effects. TL-Om1 cells were electroporated with DNA vectors expressing the ZFP5-KRAB or ZFP6-KRAB and outgrowth measured up to day 24 through (FIG. 8A) proliferation, (FIG. 8B) viability or (FIG. 8C) cell count. The ZFP-HIV-KRAB or GFP vectors were included as negative controls. Error bars represent standard deviation from samples treated in triplicate.
[0031] FIGS. 9A-9D: Screening of ZFP repressors with alternative repressor domains. (FIG. 9A) Schematic of the ZFP expression vectors with alternative repressor domains. CMV=cytomegalovirus promoter, NLS=nuclear localization signal, KRAB=kruppel-associated box, ZIM3=KRAB(ZIM3), meCP2=methyl CpG binding protein 2, PA=polyA transcription terminator. (FIG. 9B) HEK293 were transfected with a vector containing the HTLV-1 LTR bi-directional reporter to measure Fluc (sense) or the HBZ(spliced)-Rluc (anti-sense) activity with the ZFP5 variant vectors. At 48 hrs post-transfection the levels of luciferase activity were assessed. The ZFP5 variants were generated by fusing a KRAB, KRAB(ZIM3), KRAB-meCP2, PAM. A ZFP5 without a KRAB domain was also included (−). The levels of ZFP and HBZ (FIG. 9C) RNA or (FIG. 9D) protein were determined after transfecting HEK293 cells with an LTR-HBZ and the ZFP5 variants vectors. For (FIG. 9B) and (FIG. 9C), the ZFP5 variants were made relative to a control ZFP-HIV-KRAB, which was set a 100%. Error bars represent standard deviation from samples treated in triplicate. The levels of luciferase or HBZ RNA were made relative to a ZFP-HIV-KRAB control set a 100%. For (FIG. 9D), the HBZ and ZFPs were detected through a FLAG tag and myc tag, respectively. Untreated cells (mock) were included as ZFP and HBZ detection controls. Alpha-tubulin was detected as a loading control.
[0032] FIGS. 10A-10F: The anti-HTLV-I ZFPs do not affect a non-HTLV-I transformed T-cell line. Jurkat cells were electroporated with an (FIG. 10A) 2 μg ‘low’ dose or (FIG. 10B) 4 μg ‘high’ dose of mRNA expressing the ZFP5-KRAB or ZFP5-KRAB-meCP2 and outgrowth measured up to day 21 through proliferation (top panel), viability (middle panel) or cell count (bottom panel). (FIG. 10C) HEK293 cells stably expressing GFP from a LTR from HIV-1 was transfected with the ZFP5-KRAB, ZFP5-KRAB-meCP2 and ZFP-HIV-KRAB expression vectors, and 72 hrs post-transfection the levels of GFP were assessed by flow cytometry. An empty vector (pUC19) was included as a negative control. Short hairpin RNAs (shRNAs) targeted to the HIV-1 promoter (shRNA-362) and GFP (shRNA-GFP) were included as positive controls. ATL55T(+) cells were electroporated with 4 μg of ZFP5-KRAB and the levels of (FIG. 10D) HBZ and TAX RNA was assessed at 24 hrs post-electroporation. (FIG. 10E) ATL55T(+) cell line proliferation and (FIG. 10F) cell counts were assessed at day 3 and 6. The ZFP-HIV-KRAB or GFP mRNAs were included as negative controls. Error bars represent standard deviation from samples treated in triplicate.
[0033] FIGS. 11A-11C: Detection of HBZ and anti-HTLV-I ZFP molecules. TL-Om1 cells were electroporated with 2 μg or 4 μg of ZFP mRNA and the (FIG. 11A) RNA (KRAB) or (FIG. 11B) protein (anti-myc) was assessed. Untreated (mock) cells were included as a ZFP detection control. Alpha-tubulin was detected as a loading control. (FIG. 11C) TL-Om1 cells were electroporated with 2 μg of mRNA and the ZFP (KRAB), HBZsp, or HBZusp RNA was detected at 24, 48, and 72 hrs post-electroporation. A ZFP-HIV-KRAB mRNA was included as a negative control. Error bars represent standard deviation from samples treated in triplicate. The levels of HBZ RNA were made relative to a ZFP-HIV-KRAB control set a 100%.
[0034] FIGS. 12A-12C: TL-Om1 cells were electroporated with 4 μg (or 2 μg as indicated as ‘low’) of ZFP5-KRAB or ZFP5-KRAB-meCP2 mRNA, and the levels of (FIG. 12A) HBZ spliced RNA, (FIG. 12B) CCR4 RNA (24 hrs only), (FIG. 12C) or surface CCR4 receptor was assessed at 24 hrs and 48 hrs post-electroporation. Cells treated with the ZFP-HIV-KRAB mRNA or untreated cells (mock) were included as negative controls. Error bars represent standard deviation from samples treated in triplicate and p-values were determined by one-way ANOVA analysis (Dunnett's post-test) when compared to the ZFP-HIV-control (*p<0.05, **p<0.01).
[0035] FIGS. 13A-13C: ZFP5-KRAB-meCP2 is a more potent inhibitor of the HTLV-I LTR. (FIG. 13A) Jurkat cells were selected to stably express the HBZ gene expressed off a HTLV-I 3′ LTR in-frame with an internal ribosomal entry site (IRES) and a GFP-puromycin fusion protein (GFP-puro). (FIG. 13B) The Jurkat cells containing the LTR-HBZ-IRES-GFP construct were electroporated with 2 μg of ZFP5-KRAB or ZFP5-KRAB-meCP2 mRNA, and the percentage of GFP negative cells was assessed by flow cytometry at day 1, 2 or 4 post-electroporation. (FIG. 13C) Data from FIG. 13B represented as the percentage of GFP positive cells as assessed by flow cytometry at day 1, 2 or 4 post-electroporation. Error bars represent standard deviation from samples treated in triplicate. Cells treated with the ZFP-HIV-KRAB mRNA were included as a control.
[0036] FIG. 14: Anti-HTLV-I ZFP induce caspase activity. TL-Om1 cells were electroporated with 2 μg ‘low’ or 4 μg ‘high’ of ZFP5-KRAB or ZFP5-KRAB-meCP2 mRNA, and the levels of caspase 3 / 7 activity was assessed 24 hrs post-electroporation. Cells treated with the ZFP-HIV-KRAB mRNA or untreated cells (mock) were included as negative controls. Error bars represent standard deviation from samples treated in triplicate.
[0037] FIG. 15: Effect of ZFP repressor on the Fluc levels from a vector with an LTR from different HTLV-I genotypes. HEK293 cells were transfected with an LTR(a-g) spliced reporter vector with the ZFP5-KRAB and ZFP5-KRAB-meCP2 vectors, and 48 hrs post-transfection the levels of Fluc luciferase were determined. Error bars represent standard deviation from samples treated in triplicate. The levels of luciferase were made relative to a ZFP-HIV-KRAB control set a 100%.
[0038] FIGS. 16A-16B: Schematic for the development of anti-HTLV-1 EV HBZ CCR4 targeted therapy. (FIG. 16A) Stable HEK293 cells are transduced to express the EXOtic EV producer machinery including Connexion (CX43)(7), the HTLV-1 epigenetic repressor, ZFP5-KRAB / meCP2-CD mRNA (ZFP5-KrMe-CD), CD63-L7ae or CD63-anti-CCR4 for CCR4 targeted EVs. Over-expression of ZFP5-KrMe-CD results in expression and de novo packaging of ZFP5-KRAB / meCP2 protein (8). (FIG. 16B) Three different EVs are generated and tested in this proposal containing ZFP5 fused to KRAB and meCP2; the untargeted EV-a (ZFP5-KrMe), and the CCR4 targeted EVs; EV-b which consists of the PTGFRN CCR4 scFV fusion (ZFP5-KrMe-PTGFRN-R4) and EV-c which consists of CD63 fused to CCR4 (ZFP5-KrMe-CD63-R4). The EVs (EV-a-c) become taken up by HTLV-1 infected T-cells and deliver the HTLV-1 HBZ epigenetic repressor (ZFP5-KrMe-CD) mRNA and corresponding proteins (ZFP5-KrMe) both packaged into the EVs. The ZFP5-KrMe protein translocates to the nucleus where it binds and epigenetically inhibits the HBZ promoter which leads to death of the HTLV-1 HBZ driven oncogenic T-cell.
[0039] FIG. 17: Receptor targeted exosomes. Schematic of the CD63 receptor and example insertion sites of an scFv or nanobody (Ex1.1, Ex2.2, Ex2.3, or Ex2.4).
[0040] FIG. 18: Model for EV treatment of HTLV-1 infected NOD SCID ß2m mouse. Human CD34+ cells from cord blood are injected at day 1 or 2 after birth and following total body irradiation at 100 cGy. After 12 weeks engraftment the mice will be injected with HTLV-1 (MOI=5.0) infected donor matched CD4+ T-cells. The infection with HTLV-1 is monitored on weeks 4 and 8 post-infection for HTLV-1 infection by ELISA (p19) (Ji, 2020 #4460) and qRT-PCR for HBZ and Tax mRNA expression. Following detectable infection ˜week 8, EVs are administered R.O. every week thereafter until week 14. At week 14 and bi-weekly blood draws will be carried out to measure anti-HTLV-1 effects of the EV treatment and HTLV-1 persistence. The mice are euthanized and analysed at 18 weeks post-viral infection for tissue harvest and analysis (˜32 weeks post-transplantation).
[0041] FIGS. 19A-19B: LTR-targeted ZFP repressors reduce chromatin accessibility. TL-Om1 cells were electroporated with 4 μg of mRNA expressing the ZFP5-KRAB or ZFP5-KRAB-meCP2 and at 24 hrs the cells were subjected to ATAC-seq to assess chromatin accessibility. (FIG. 19A) Integrated genomic viewer (IGV) of the HTLV-I genome displaying accessibility. (FIG. 19B) Enrichment plot of nucleosome-free regions across HTLV-I's LTR. The read counts are the average of triplicate treated cells.
[0042] FIGS. 20A-20B: Specificity of the ZFP-KRAB vectors. (FIG. 20A) HEK293 cells were transfected with the HTLV-I 3′-LTR driving the expression of the HBZ-3×FLAG with the ZFP5-KRAB vector, and 48 hrs post-transfection the levels of HBZ RNA and protein were assessed. (FIG. 20B) Jurkat cells were electroporated with 2 μg of mRNA expressing the ZFP3-KRAB or ZFP-HIV-KRAB and proliferation was assessed at day 3. Error bars represent standard deviation from samples treated in triplicate.
[0043] FIGS. 21A-21B: Anti-HTLV-I ZFPs effects in TL-Om1 cells. (FIG. 21A) The levels of HBZ and TAX RNA was determined for MT-2, MT-4, Jurkat and TL-Om1 cells. (FIG. 21B) TL-Om1 cells were electroporated with a 2 μg ‘low’ dose or 4 μg ‘high’ dose of mRNA expressing the ZFP5-KRAB or ZFP5-KRAB-meCP and the number of viable cells per ml was determined using flow cytometry at day 2 and 5 (top panels), and day 3 and 6 (bottom panels). The ZFP-HIV-KRAB was included as negative controls. Error bars represent standard deviation from samples treated in triplicate.
[0044] FIGS. 22A-22C: Pathway analysis on a ATL cell line treated with anti-HTLV ZFPs. TL-Om1 cells were electroporated with 4 μg of (FIG. 21A) ZFP5-KRAB, (FIG. 21B) ZFP5-KRAB-meCP2, or (FIG. 21C) ZFP-HIV-KRAB mRNA and subjected to ATAC-seq. KEGG pathway analysis was performed for the ZFPs and each compared to mock treated cells. Dot size corresponds to gene ratio. Moreover, adjusted p values are also indicated.
[0045] FIG. 23: Reduced viability with ZFP5-HTLV treatment in ATL55T(+) cells compared to control.
[0046] FIG. 24: ATAC-seq reads reduced at a known enhancer site within SRF-ERK1 site in the HTLV ZFP treated samples compared to controls.DETAILED DESCRIPTION
[0047] While various embodiments and aspects of the present invention are shown and described herein, it will be obvious to those skilled in the art that such embodiments and aspects are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention.
[0048] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described. All documents, or portions of documents, cited in the application including, without limitation, patents, patent applications, articles, books, manuals, and treatises are hereby expressly incorporated by reference in their entirety for any purpose.
[0049] The abbreviations used herein have their conventional meaning within the chemical and biological arts. The chemical structures and formulae set forth herein are constructed according to the standard rules of chemical valency known in the chemical arts.
[0050] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by a person of ordinary skill in the art. See, e.g., Singleton et al., DICTIONARY OF MICROBIOLOGY AND MOLECULAR BIOLOGY 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., MOLECULAR CLONING, A LABORATORY MANUAL, Cold Springs Harbor Press (Cold Springs Harbor, NY 1989). Any methods, devices and materials similar or equivalent to those described herein can be used in the practice of this invention. The following definitions are provided to facilitate understanding of certain terms used frequently herein and are not meant to limit the scope of the present disclosure.
[0051] “Nucleic acid” refers to nucleotides (e.g., deoxyribonucleotides or ribonucleotides) and polymers thereof in either single-, double- or multiple-stranded form, or complements thereof; or nucleosides (e.g., deoxyribonucleosides or ribonucleosides). In embodiments, “nucleic acid” does not include nucleosides. The terms “polynucleotide,”“oligonucleotide,”“oligo” or the like refer, in the usual and customary sense, to a linear sequence of nucleotides. The term “nucleoside” refers, in the usual and customary sense, to a glycosylamine including a nucleobase and a five-carbon sugar (ribose or deoxyribose). Non limiting examples, of nucleosides include, cytidine, uridine, adenosine, guanosine, thymidine and inosine. The term “nucleotide” refers, in the usual and customary sense, to a single unit of a polynucleotide, i.e., a monomer. Nucleotides can be ribonucleotides, deoxyribonucleotides, or modified versions thereof. Examples of polynucleotides contemplated herein include single and double stranded DNA, single and double stranded RNA, and hybrid molecules having mixtures of single and double stranded DNA and RNA. Examples of nucleic acid, e.g. polynucleotides contemplated herein include any types of RNA, e.g. mRNA, siRNA, miRNA, and guide RNA and any types of DNA, genomic DNA, plasmid DNA, and minicircle DNA, and any fragments thereof. The term “duplex” in the context of polynucleotides refers, in the usual and customary sense, to double strandedness. Nucleic acids can be linear or branched. For example, nucleic acids can be a linear chain of nucleotides or the nucleic acids can be branched, e.g., such that the nucleic acids comprise one or more arms or branches of nucleotides. Optionally, the branched nucleic acids are repetitively branched to form higher ordered structures such as dendrimers and the like.
[0052] As may be used herein, the terms “nucleic acid,”“nucleic acid molecule,”“nucleic acid oligomer,”“oligonucleotide,”“nucleic acid sequence,”“nucleic acid fragment” and “polynucleotide” are used interchangeably and are intended to include, but are not limited to, a polymeric form of nucleotides covalently linked together that may have various lengths, either deoxyribonucleotides or ribonucleotides, or analogs, derivatives or modifications thereof. Different polynucleotides may have different three-dimensional structures, and may perform various functions, known or unknown. Non-limiting examples of polynucleotides include a gene, a gene fragment, an exon, an intron, intergenic DNA (including, without limitation, heterochromatic DNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, a ribozyme, cDNA, a recombinant polynucleotide, a branched polynucleotide, a plasmid, a vector, isolated DNA of a sequence, isolated RNA of a sequence, a nucleic acid probe, and a primer. For example, the nucleic acid provided herein may be part of a vector. In embodiments, the nucleic acid provided herein may be part of a lentiviral vector, which may be transduced into a cell. Polynucleotides useful in the methods of the disclosure may comprise natural nucleic acid sequences and variants thereof, artificial nucleic acid sequences, or a combination of such sequences.
[0053] The terms also encompass nucleic acids containing known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, and non-naturally occurring, which have similar binding properties as the reference nucleic acid, and which are metabolized in a manner similar to the reference nucleotides. Examples of such analogs include, without limitation, phosphodiester derivatives including, e.g., phosphoramidate, phosphorodiamidate, phosphorothioate (also known as phosphothioate having double bonded sulfur replacing oxygen in the phosphate), phosphorodithioate, phosphonocarboxylic acids, phosphonocarboxylates, phosphonoacetic acid, phosphonoformic acid, methyl phosphonate, boron phosphonate, or O-methylphosphoroamidite linkages (see Eckstein, OLIGONUCLEOTIDES AND ANALOGUES: A PRACTICAL APPROACH, Oxford University Press) as well as modifications to the nucleotide bases such as in 5-methyl cytidine or pseudouridine; and peptide nucleic acid backbones and linkages. Other analog nucleic acids include those with positive backbones; non-ionic backbones, modified sugars, and non-ribose backbones (e.g. phosphorodiamidate morpholino oligos or locked nucleic acids (LNA) as known in the art), including those described in U.S. Pat. Nos. 5,235,033 and 5,034,506, and Chapters 6 and 7, ASC Symposium Series 580, CARBOHYDRATE MODIFICATIONS IN ANTISENSE RESEARCH, Sanghui & Cook, eds. Nucleic acids containing one or more carbocyclic sugars are also included within one definition of nucleic acids. Modifications of the ribose-phosphate backbone may be done for a variety of reasons, e.g., to increase the stability and half-life of such molecules in physiological environments or as probes on a biochip. Mixtures of naturally occurring nucleic acids and analogs can be made; alternatively, mixtures of different nucleic acid analogs, and mixtures of naturally occurring nucleic acids and analogs may be made. In embodiments, the internucleotide linkages in DNA are phosphodiester, phosphodiester derivatives, or a combination of both.
[0054] Nucleic acids can include nonspecific sequences. As used herein, the term “nonspecific sequence” refers to a nucleic acid sequence that contains a series of residues that are not designed to be complementary to or are only partially complementary to any other nucleic acid sequence. By way of example, a nonspecific nucleic acid sequence is a sequence of nucleic acid residues that does not function as an inhibitory nucleic acid when contacted with a cell or organism.
[0055] A polynucleotide is typically composed of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (uracil (U) for thymine (T) when the polynucleotide is RNA). Thus, the term “polynucleotide sequence” is the alphabetical representation of a polynucleotide molecule; alternatively, the term may be applied to the polynucleotide molecule itself. This alphabetical representation can be input into databases in a computer having a central processing unit and used for bioinformatics applications such as functional genomics and homology searching. Polynucleotides may optionally include one or more non-standard nucleotide(s), nucleotide analog(s) and / or modified nucleotides.
[0056] The term “complement,” as used herein, refers to a nucleotide (e.g., RNA or DNA) or a sequence of nucleotides capable of base pairing with a complementary nucleotide or sequence of nucleotides. As described herein and commonly known in the art the complementary (matching) nucleotide of adenosine is thymidine and the complementary (matching) nucleotide of guanosine is cytosine. Thus, a complement may include a sequence of nucleotides that base pair with corresponding complementary nucleotides of a second nucleic acid sequence. The nucleotides of a complement may partially or completely match the nucleotides of the second nucleic acid sequence. Where the nucleotides of the complement completely match each nucleotide of the second nucleic acid sequence, the complement forms base pairs with each nucleotide of the second nucleic acid sequence. Where the nucleotides of the complement partially match the nucleotides of the second nucleic acid sequence only some of the nucleotides of the complement form base pairs with nucleotides of the second nucleic acid sequence. Examples of complementary sequences include coding and a non-coding sequences, wherein the non-coding sequence contains complementary nucleotides to the coding sequence and thus forms the complement of the coding sequence. A further example of complementary sequences are sense and antisense sequences, wherein the sense sequence contains complementary nucleotides to the antisense sequence and thus forms the complement of the antisense sequence.
[0057] As described herein the complementarity of sequences may be partial, in which only some of the nucleic acids match according to base pairing, or complete, where all the nucleic acids match according to base pairing. Thus, two sequences that are complementary to each other, may have a specified percentage of nucleotides that are the same (i.e., about 60% identity, preferably 65%, 70%, 75%, 75%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity over a specified region).
[0058] The term “amino acid” refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids. Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, γ-carboxyglutamate, and O-phosphoserine. Amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an α carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e.g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid. Amino acid mimetics refers to chemical compounds that have a structure that is different from the general chemical structure of an amino acid, but that functions in a manner similar to a naturally occurring amino acid. The terms “non-naturally occurring amino acid” and “unnatural amino acid” refer to amino acid analogs, synthetic amino acids, and amino acid mimetics which are not found in nature.
[0059] The term “amino acid side chain” refers to the functional substituent contained on amino acids. For example, an amino acid side chain may be the side chain of a naturally occurring amino acid. Naturally occurring amino acids are those encoded by the genetic code (e.g., alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine), as well as those amino acids that are later modified, e.g., hydroxyproline, γ-carboxyglutamate, and O-phosphoserine. In embodiments, the amino acid side chain may be a non-natural amino acid side chain. In embodiments, the amino acid side chain is H,
[0060] Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes.
[0061] The terms “polypeptide,”“peptide” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues, wherein the polymer may In embodiments be conjugated to a moiety that does not consist of amino acids. The terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymers.
[0062] A “fusion protein” refers to a chimeric protein encoding two or more separate protein sequences that are recombinantly expressed as a single moiety. Because the different proteins in fusion proteins may affect the functionality of other proteins under certain circumstances, peptide linkers may be used between different proteins within the same fusion protein. These peptide linkers may have a flexible structure and separate the proteins within the fusion protein so that each protein in the fusion proteins substantially retains its function. Peptide linkers are known in the art and described, for example, in Chen et al, Adv Drug Deliv Rev, 65(10); 1357-1369 (2013).
[0063] An amino acid or nucleotide base “position” is denoted by a number that sequentially identifies each amino acid (or nucleotide base) in the reference sequence based on its position relative to the N-terminus (or 5′-end). Due to deletions, insertions, truncations, fusions, and the like that must be taken into account when determining an optimal alignment, in general the amino acid residue number in a test sequence determined by simply counting from the N-terminus will not necessarily be the same as the number of its corresponding position in the reference sequence. For example, in a case where a variant has a deletion relative to an aligned reference sequence, there will be no amino acid in the variant that corresponds to a position in the reference sequence at the site of deletion. Where there is an insertion in an aligned reference sequence, that insertion will not correspond to a numbered amino acid position in the reference sequence. In the case of truncations or fusions there can be stretches of amino acids in either the reference or aligned sequence that do not correspond to any amino acid in the corresponding sequence.
[0064] The terms “numbered with reference to” or “corresponding to,” when used in the context of the numbering of a given amino acid or polynucleotide sequence, refers to the numbering of the residues of a specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence. An amino acid residue in a protein “corresponds” to a given residue when it occupies the same essential structural position within the protein as the given residue. One skilled in the art will immediately recognize the identity and location of residues corresponding to a specific position in a protein in other proteins with different numbering systems. For example, by performing a simple sequence alignment with a protein the identity and location of residues corresponding to specific positions of the protein are identified in other protein sequences aligning to the protein. For example, a selected residue in a selected protein corresponds to glutamic acid at position 138 when the selected residue occupies the same essential spatial or other structural relationship as a glutamic acid at position 138. In some embodiments, where a selected protein is aligned for maximum homology with a protein, the position in the aligned selected protein aligning with glutamic acid 138 is the to correspond to glutamic acid 138. Instead of a primary sequence alignment, a three dimensional structural alignment can also be used, e.g., where the structure of the selected protein is aligned for maximum correspondence with the glutamic acid at position 138, and the overall structures compared. In this case, an amino acid that occupies the same essential position as glutamic acid 138 in the structural model is the to correspond to the glutamic acid 138 residue.
[0065] “Conservatively modified variants” applies to both amino acid and nucleic acid sequences. With respect to particular nucleic acid sequences, “conservatively modified variants” refers to those nucleic acids that encode identical or essentially identical amino acid sequences. Because of the degeneracy of the genetic code, a number of nucleic acid sequences will encode any given protein. For instance, the codons GCA, GCC, GCG and GCU all encode the amino acid alanine. Thus, at every position where an alanine is specified by a codon, the codon can be altered to any of the corresponding codons described without altering the encoded polypeptide. Such nucleic acid variations are “silent variations,” which are one species of conservatively modified variations. Every nucleic acid sequence herein which encodes a polypeptide also describes every possible silent variation of the nucleic acid. One of skill will recognize that each codon in a nucleic acid (except AUG, which is ordinarily the only codon for methionine, and TGG, which is ordinarily the only codon for tryptophan) can be modified to yield a functionally identical molecule. Accordingly, each silent variation of a nucleic acid which encodes a polypeptide is implicit in each described sequence.
[0066] As to amino acid sequences, one of skill will recognize that individual substitutions, deletions or additions to a nucleic acid, peptide, polypeptide, or protein sequence which alters, adds or deletes a single amino acid or a small percentage of amino acids in the encoded sequence is a “conservatively modified variant” where the alteration results in the substitution of an amino acid with a chemically similar amino acid. Conservative substitution tables providing functionally similar amino acids are well known in the art. Such conservatively modified variants are in addition to and do not exclude polymorphic variants, interspecies homologs, and alleles of the disclosure.
[0067] The following eight groups each contain amino acids that are conservative substitutions for one another:
[0068] 1) Alanine (A), Glycine (G);
[0069] 2) Aspartic acid (D), Glutamic acid (E);
[0070] 3) Asparagine (N), Glutamine (Q);
[0071] 4) Arginine (R), Lysine (K);
[0072] 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V);
[0073] 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W);
[0074] 7) Serine (S), Threonine (T); and
[0075] 8) Cysteine (C), Methionine (M)
[0076] (see, e.g., Creighton, Proteins (1984)).
[0077] The terms “identical” or percent “identity,” in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (i.e., about 60% identity, preferably 65%, 70%, 75%, 75%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity over a specified region, when compared and aligned for maximum correspondence over a comparison window or designated region) as measured using a BLAST or BLAST 2.0 sequence comparison algorithms with default parameters described below, or by manual alignment and visual inspection (see, e.g., NCBI web site http: / / www.ncbi.nlm.nih.gov / BLAST / or the like). Such sequences are then said to be “substantially identical.” This definition also refers to, or may be applied to, the compliment of a test sequence. The definition also includes sequences that have deletions and / or additions, as well as those that have substitutions. The preferred algorithms can account for gaps and the like. Preferably, identity exists over a region that is at least about 25 amino acids or nucleotides in length, or more preferably over a region that is 50-100 amino acids or nucleotides in length.
[0078] “Percentage of sequence identity” is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity.
[0079] An amino acid or nucleotide base “position” is denoted by a number that sequentially identifies each amino acid (or nucleotide base) in the reference sequence based on its position relative to the N-terminus (or 5′-end). Due to deletions, insertions, truncations, fusions, and the like that must be taken into account when determining an optimal alignment, in general the amino acid residue number in a test sequence determined by simply counting from the N-terminus will not necessarily be the same as the number of its corresponding position in the reference sequence. For example, in a case where a variant has a deletion relative to an aligned reference sequence, there will be no amino acid in the variant that corresponds to a position in the reference sequence at the site of deletion. Where there is an insertion in an aligned reference sequence, that insertion will not correspond to a numbered amino acid position in the reference sequence. In the case of truncations or fusions there can be stretches of amino acids in either the reference or aligned sequence that do not correspond to any amino acid in the corresponding sequence.
[0080] A “comparison window”, as used herein, includes reference to a segment of any one of the number of contiguous positions selected from the group consisting of, e.g., a full length sequence or from 20 to 600, about 50 to about 200, or about 100 to about 150 amino acids or nucleotides in which a sequence may be compared to a reference sequence of the same number of contiguous positions after the two sequences are optimally aligned. Methods of alignment of sequences for comparison are well-known in the art. Optimal alignment of sequences for comparison can be conducted, e.g., by the local homology algorithm of Smith and Waterman (1970) Adv. Appl. Math. 2:482c, by the homology alignment algorithm of Needleman and Wunsch (1970) J. Mol. Biol. 48:443, by the search for similarity method of Pearson and Lipman (1988) Proc. Nat'l. Acad. Sci. USA 85:2444, by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI), or by manual alignment and visual inspection (see, e.g., Ausubel et al., Current Protocols in Molecular Biology (1995 supplement)).
[0081] An example of an algorithm that is suitable for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al. (1977) Nuc. Acids Res. 25:3389-3402, and Altschul et al. (1990) J. Mol. Biol. 215:403-410, respectively. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ). This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a word length (W) of 11, an expectation (E) or 10, M=5, N=−4 and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word length of 3, and expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff (1989) Proc. Natl. Acad. Sci. USA 89:10915) alignments (B) of 50, expectation (E) of 10, M=5, N=−4, and a comparison of both strands.
[0082] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5787). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the smallest sum probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.2, more preferably less than about 0.01, and most preferably less than about 0.001.
[0083] For specific proteins described herein, the named protein includes any of the protein's naturally occurring forms, variants or homologs that maintain activity of the protein (e.g., within at least 50%, 75%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to the native protein). In some embodiments, variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring form. In other embodiments, the protein is the protein as identified by its NCBI sequence reference. In other embodiments, the protein is the protein as identified by its NCBI sequence reference, homolog or functional fragment thereof.
[0084] The term “HBZ protein” or “HBZ” as used herein includes any of the recombinant or naturally-occurring forms of HTLV-1 basic zipper factor (HBZ), or variants or homologs thereof that maintain HBZ activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to HBZ). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring HBZ protein. In embodiments, the HBZ protein is substantially identical to the protein identified by the UniProt reference number POC746 or a variant or homolog having substantial identity thereto.
[0085] The term “meCP2 protein” or “meCP2” as used herein includes any of the recombinant or naturally-occurring forms of methyl CpG binding protein 2 (meCP2), also known as demethylase, DMTase, or variants or homologs thereof that maintain meCP2 activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to meCP2). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring meCP2 protein. In embodiments, the meCP2 protein is substantially identical to the protein identified by the UniProt reference number Q9UBB5 or a variant or homolog having substantial identity thereto. In embodiments, the meCP2 protein includes a sequence having at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the sequence of SEQ ID NO:125. In embodiments, the meCP2 protein includes a sequence having at least 80% sequence identity to the sequence of SEQ ID NO:125. In embodiments, the meCP2 protein includes a sequence having at least 90% sequence identity to the sequence of SEQ ID NO:125. In embodiments, the meCP2 protein includes a sequence having at least 95% sequence identity to the sequence of SEQ ID NO:125. In embodiments, the meCP2 protein includes a sequence having at least 96% sequence identity to the sequence of SEQ ID NO:125. In embodiments, the meCP2 protein includes a sequence having at least 97% sequence identity to the sequence of SEQ ID NO:125. In embodiments, the meCP2 protein includes a sequence having at least 98% sequence identity to the sequence of SEQ ID NO:125. In embodiments, the meCP2 protein includes a sequence having at least 99% sequence identity to the sequence of SEQ ID NO:125. In embodiments, the meCP2 protein includes the sequence of SEQ ID NO:125. In embodiments, the meCP2 protein is the sequence of SEQ ID NO:125.
[0086] The term “DNA methyltransferase” or “DNA methyltransferase protein” as provided herein refers to an enzyme that catalyzes the transfer of a methyl group to DNA. Non-limiting examples of DNA methyltransferases include Dnmt1, Dnmt3A, and Dnmt3B. In aspects, the DNA methyltransferase is mammalian DNA methyltransferase. In aspects, the DNA methyltransferase is human DNA methyltransferase. In aspects, the DNA methyltransferase is mouse DNA methyltransferase. In aspects, the DNA methyltransferase is a bacterial cytosine methyltransferase and / or a bacterial non-cytosine methyltransferase. Depending on the specific DNA methyltransferase, different regions of DNA are methylated. For example, Dnmt3A typically targets CpG dinucleotides for methylation. Through DNA methylation, DNA methyltransferases can modify the activity of a DNA segment (e.g., gene expression) without altering the DNA sequence. In aspects, DNA methylation results in repression of gene transcription and / or modulation of methylation sensitive transcription factors or CTCF. As described herein, fusion proteins may include one or more (e.g., two) DNA methyltransferases. When a DNA methyltransferase is included as part of a fusion protein, the DNA methyltransferase may be referred to as a “DNA methyltransferase domain.”
[0087] A “Dnmt3A”, “Dnmt3a,”“DNA (cytosine-5)-methyltransferase 3A” or “DNA methyltransferase 3a” protein as referred to herein includes any of the recombinant or naturally-occurring forms of the Dnmt3A enzyme or variants or homologs thereof that maintain Dnmt3A enzyme activity (e.g. within at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to Dnmt3A). In aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g., a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring Dnmt3A protein. In aspects, the Dnmt3A protein is substantially identical to the protein identified by the UniProt reference number Q9Y6K1 or a variant or homolog having substantial identity thereto.
[0088] The term “Kruppel associated box domain” or “KRAB domain” as provided herein refers to a category of transcriptional repression domains present in approximately 400 human zinc finger protein-based transcription factors. KRAB domains typically include about 45 to about 75 amino acid residues. A description of KRAB domains, including their function and use, may be found, for example, in Ecco, G., Imbeault, M., Trono, D., KRAB zinc finger proteins, Development 144, 2017; Lambert et al. The human transcription factors, Cell 172, 2018; Gilbert et al., Cell (2013); and Gilbert et al., Cell (2014). In embodiments, the KRAB domain includes a sequence having at least 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to the sequence of SEQ ID NO:123. In embodiments, the KRAB domain includes a sequence having at least 80% sequence identity to the sequence of SEQ ID NO:123. In embodiments, the KRAB domain includes a sequence having at least 90% sequence identity to the sequence of SEQ ID NO:123. In embodiments, the KRAB domain includes a sequence having at least 95% sequence identity to the sequence of SEQ ID NO:123. In embodiments, the KRAB domain includes a sequence having at least 96% sequence identity to the sequence of SEQ ID NO:123. In embodiments, the KRAB domain includes a sequence having at least 97% sequence identity to the sequence of SEQ ID NO:123. In embodiments, the KRAB domain includes a sequence having at least 98% sequence identity to the sequence of SEQ ID NO:123. In embodiments, the KRAB domain includes a sequence having at least 99% sequence identity to the sequence of SEQ ID NO:123. In embodiments, the KRAB domain includes the sequence of SEQ ID NO:123. In embodiments, the KRAB domain is the sequence of SEQ ID NO:123.
[0089] The term “CD63 protein” or “CD63” as used herein includes any of the recombinant or naturally-occurring forms of CD63, also known as Granulophysin, Lysosomal-associated membrane protein 3, LAMP-3, Lysosome integral membrane protein 1, Limp1, or variants or homologs thereof that maintain CD63 activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to CD63). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring CD63 protein. In embodiments, the CD63 protein is substantially identical to the protein identified by the UniProt reference number P08962 or a variant or homolog having substantial identity thereto.
[0090] The term “PTGFRN protein” or “PTGFRN” as used herein includes any of the recombinant or naturally-occurring forms of Prostaglandin F2 receptor negative regulator (PTGFRN), also known as CD9 partner 1, EWI motif-containing protein F, CD315, or variants or homologs thereof that maintain PTGFRN activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to PTGFRN). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring PTGFRN protein. In embodiments, the PTGFRN protein is substantially identical to the protein identified by the UniProt reference number Q9P2B2 or a variant or homolog having substantial identity thereto.
[0091] The term “CD9 protein” or “CD9” as used herein includes any of the recombinant or naturally-occurring forms of CD9, also known as MIC3, or TSPAN29, or variants or homologs thereof that maintain CD9 activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to CD9). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring CD9 protein. In embodiments, the CD9 protein is substantially identical to the protein identified by the UniProt reference number P21926 or a variant or homolog having substantial identity thereto.
[0092] The term “CCR4 protein” or “CCR4” as used herein includes any of the recombinant or naturally-occurring forms of C-C chemokine receptor type 4 (CCR4), also known as K5-5, CD194, or variants or homologs thereof that maintain CCR4 activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to CCR4). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring CCR4 protein. In embodiments, the CCR4 protein is substantially identical to the protein identified by the UniProt reference number P51679 or a variant or homolog having substantial identity thereto.
[0093] The term “CD4 protein” or “CD4” as used herein includes any of the recombinant or naturally-occurring forms of CD4, also known as T-cell surface glycoprotein CD4, T-cell surface antigen T4 / Leu-3 or variants or homologs thereof that maintain CD4 activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to CD4). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring CD4 protein. In embodiments, the CD4 protein is substantially identical to the protein identified by the UniProt reference number P01730 or a variant or homolog having substantial identity thereto.
[0094] The term “OX40 protein” or “OX40” as used herein includes any of the recombinant or naturally-occurring forms of OX40, also known as tumor necrosis factor receptor superfamily member 4 (TNFRSF4), ACT35 antigen, TAX transcriptionally-activated glycoprotein 1 receptor, CD134, or variants or homologs thereof that maintain OX40 activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to OX40). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring OX40 protein. In embodiments, the OX40 protein is substantially identical to the protein identified by the UniProt reference number P43489 or a variant or homolog having substantial identity thereto.
[0095] The term “CD5 protein” or “CD5” as used herein includes any of the recombinant or naturally-occurring forms of CD5, also known as T-cell surface glycoprotein CD5, lymphocyte antigen T1 / Leu-1, or variants or homologs thereof that maintain CD5 activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to CD5). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring CD5 protein. In embodiments, the CD5 protein is substantially identical to the protein identified by the UniProt reference number P06127 or a variant or homolog having substantial identity thereto.
[0096] The term “CD25 protein” or “CD25” as used herein includes any of the recombinant or naturally-occurring forms of CD25, also known as Interleukin-2 receptor subunit alpha, TAC antigen, p55, IL-2-RA, IL2-RA, or variants or homologs thereof that maintain CD25 activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to CD25). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring CD25 protein. In embodiments, the CD25 protein is substantially identical to the protein identified by the UniProt reference number P01589 or a variant or homolog having substantial identity thereto.
[0097] The term “lactadherin protein” or “lactadherin” as used herein includes any of the recombinant or naturally-occurring forms of lactadherin, also known as breast epithelial antigen BA46, IMIFG, MFGM, milk fat globule-EGF factor 8, SED1, or variants or homologs thereof that maintain lactadherin activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to lactadherin). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring lactadherin protein. In embodiments, the lactadherin protein is substantially identical to the protein identified by the UniProt reference number Q08431 or a variant or homolog having substantial identity thereto.
[0098] The term “CD37 protein” or “CD37” as used herein includes any of the recombinant or naturally-occurring forms of CD37, also known as leukocyte antigen CD37, tetraspanin-26, or variants or homologs thereof that maintain CD37 activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to CD37). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring CD37 protein. In embodiments, the CD37 protein is substantially identical to the protein identified by the UniProt reference number P11049 or a variant or homolog having substantial identity thereto.
[0099] The term “LAMP-1 protein” or “LAMP-1” as used herein includes any of the recombinant or naturally-occurring forms of LAMP-1, also known lysosome-associated membrane glycoprotein 1, CD107a, or variants or homologs thereof that maintain LAMP-1 activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to LAMP-1). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring LAMP-1 protein. In embodiments, the LAMP-1 protein is substantially identical to the protein identified by the UniProt reference number P11279 or a variant or homolog having substantial identity thereto.
[0100] The term “LAMP-2A protein” or “LAMP-2A” as used herein includes any of the recombinant or naturally-occurring forms of LAMP-2A, also known lysosome-associated membrane glycoprotein 2, CD107b, LGP-96, LAMP-2, or variants or homologs thereof that maintain LAMP-2A activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to LAMP-2A). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring LAMP-2A protein. In embodiments, the LAMP-2A protein is substantially identical to the protein identified by the UniProt reference number P13473 or a variant or homolog having substantial identity thereto.
[0101] The term “CD70 protein” or “CD70” as used herein includes any of the recombinant or naturally-occurring forms of CD70, also known as CD27 ligand, tumor necrosis factor ligand superfamily member 7, or variants or homologs thereof that maintain CD70 activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to CD70). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring CD70 protein. In embodiments, the CD70 protein is substantially identical to the protein identified by the UniProt reference number P32970 or a variant or homolog having substantial identity thereto.
[0102] The term “IL15RA protein” or “IL15RA” as used herein includes any of the recombinant or naturally-occurring forms of IL15RA, also known as CD215, soluble interleukin-15 receptor subunit alpha, IL-15 receptor subunit alpha, tumor necrosis factor ligand superfamily member 7, or variants or homologs thereof that maintain IL15RA activity (e.g. within at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99% or 100% activity compared to IL15RA). In some aspects, the variants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99% or 100% amino acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 50, 100, 150 or 200 continuous amino acid portion) compared to a naturally occurring IL15RA protein. In embodiments, the IL15RA protein is substantially identical to the protein identified by the UniProt reference number Q13261 or a variant or homolog having substantial identity thereto.
[0103] The term “antibody” refers to a polypeptide encoded by an immunoglobulin gene or functional fragments thereof that specifically binds and recognizes an antigen. The recognized immunoglobulin genes include the kappa, lambda, alpha, gamma, delta, epsilon, and mu constant region genes, as well as the myriad immunoglobulin variable region genes. Light chains are classified as either kappa or lambda. Heavy chains are classified as gamma, mu, alpha, delta, or epsilon, which in turn define the immunoglobulin classes, IgG, IgM, IgA, IgD and IgE, respectively.
[0104] The phrase “specifically (or selectively) binds” to an antibody or “specifically (or selectively) immunoreactive with,” when referring to a protein or peptide, refers to a binding reaction that is determinative of the presence of the protein, often in a heterogeneous population of proteins and other biologics. Thus, under designated immunoassay conditions, the specified antibodies bind to a particular protein at least two times the background and more typically more than 10 to 100 times background. Specific binding to an antibody under such conditions requires an antibody that is selected for its specificity for a particular protein. For example, polyclonal antibodies can be selected to obtain only a subset of antibodies that are specifically immunoreactive with the selected antigen and not with other proteins. This selection may be achieved by subtracting out antibodies that cross-react with other molecules. A variety of immunoassay formats may be used to select antibodies specifically immunoreactive with a particular protein. For example, solid-phase ELISA immunoassays are routinely used to select antibodies specifically immunoreactive with a protein (see, e.g., Harlow & Lane, Using Antibodies, A Laboratory Manual (1998) for a description of immunoassay formats and conditions that can be used to determine specific immunoreactivity).
[0105] An exemplary immunoglobulin (antibody) structural unit comprises a tetramer. Each tetramer is composed of two identical pairs of polypeptide chains, each pair having one “light” (about 25 kDa) and one “heavy” chain (about 50-70 kDa). The N-terminus of each chain defines a variable region of about 100 to 110 or more amino acids primarily responsible for antigen recognition. The terms “variable heavy chain” or “VH,” refers to the variable region of an immunoglobulin heavy chain, including an Fv, scFv, dsFv or Fab; while the terms “variable light chain” or “VL” refers to the variable region of an immunoglobulin light chain, including of an Fv, scFv, dsFv or Fab.
[0106] Examples of antibody functional fragments include, but are not limited to, complete antibody molecules, antibody fragments, such as Fv, single chain Fv (scFv), complementarity determining regions (CDRs), VL (light chain variable region), VH (heavy chain variable region), Fab, F(ab)2′ and any combination of those or any other functional portion of an immunoglobulin peptide capable of binding to target antigen (see, e.g., Fundamental Immunology (Paul ed., 4th ed. 2001). As appreciated by one of skill in the art, various antibody fragments can be obtained by a variety of methods, for example, digestion of an intact antibody with an enzyme, such as pepsin; or de novo synthesis. Antibody fragments are often synthesized de novo either chemically or by using recombinant DNA methodology. Thus, the term antibody includes antibody fragments either produced by the modification of whole antibodies, or those synthesized de novo using recombinant DNA methodologies (e.g., single chain Fv) or those identified using phage display libraries (see, e.g., McCafferty et al., (1990) Nature 348:552). The term “antibody” also includes bivalent or bispecific molecules, diabodies, triabodies, and tetrabodies. Bivalent and bispecific molecules are described in, e.g., Kostelny et al. (1992) J. Immunol. 148:1547, Pack and Pluckthun (1992) Biochemistry 31:1579, Hollinger et al. (1993), PNAS. USA 90:6444, Gruber et al. (1994) J Immunol. 152:5368, Zhu et al. (1997) Protein Sci. 6:781, Hu et al. (1996) Cancer Res. 56:3055, Adams et al. (1993) Cancer Res. 53:4026, and McCartney, et al. (1995) Protein Eng. 8:301.
[0107] A single-chain variable fragment (scFv) is typically a fusion protein of the variable regions of the heavy (VH) and light chains (VL) of immunoglobulins, connected with a short linker peptide of 10 to about 25 amino acids. The linker may usually be rich in glycine for flexibility, as well as serine or threonine for solubility. The linker can either connect the N-terminus of the VH with the C-terminus of the VL, or vice versa.
[0108] The epitope of a mAb is the region of its antigen to which the mAb binds. Two antibodies bind to the same or overlapping epitope if each competitively inhibits (blocks) binding of the other to the antigen. That is, a 1×, 5×, 10×, 20× or 100× excess of one antibody inhibits binding of the other by at least 30% but preferably 50%, 75%, 90% or even 99% as measured in a competitive binding assay (see, e.g., Junghans et al., Cancer Res. 50:1495, 1990). Alternatively, two antibodies have the same epitope if essentially all amino acid mutations in the antigen that reduce or eliminate binding of one antibody reduce or eliminate binding of the other. Two antibodies have overlapping epitopes if some amino acid mutations that reduce or eliminate binding of one antibody reduce or eliminate binding of the other.
[0109] A “ligand” refers to an agent, e.g., a polypeptide or other molecule, capable of binding to a receptor or antibody, antibody variant, antibody region or fragment thereof.
[0110] The term “gene” means the segment of DNA involved in producing a protein; it includes regions preceding and following the coding region (leader and trailer) as well as intervening sequences (introns) between individual coding segments (exons). The leader, the trailer as well as the introns include regulatory elements that are necessary during the transcription and the translation of a gene. Further, a “protein gene product” is a protein expressed from a particular gene.
[0111] The terms “plasmid”, “vector” or “expression vector” refer to a nucleic acid molecule that encodes for genes and / or regulatory elements necessary for the expression of genes. Expression of a gene from a plasmid can occur in cis or in trans. If a gene is expressed in cis, the gene and the regulatory elements are encoded by the same plasmid. Expression in trans refers to the instance where the gene and the regulatory elements are encoded by separate plasmids.
[0112] As used herein, the term “construct” is intended to mean any recombinant nucleic acid molecule. In embodiments, a construct includes an expression cassette, plasmid, cosmid, virus, autonomously replicating polynucleotide molecule, phage, or linear or circular, single-stranded or double-stranded, DNA or RNA polynucleotide molecule. A construct may be derived from any source, capable of genomic integration or autonomous replication, including a nucleic acid molecule where one or more nucleic acid sequences has been linked in a functionally operative manner, e.g., operably linked.
[0113] The terms “operably linked” or “functionally linked”, are interchangeable and denote a physical or functional linkage between two or more elements, e.g., polypeptide sequences or polynucleotide sequences, which permits them to operate in their intended fashion. For example, an operable linkage between a polynucleotide of interest and a regulatory sequence (for example, a promoter, an LTR, a sequence within an LTR) is functional link that allows for expression of the polynucleotide of interest. In this sense, the term “operably linked” refers to the positioning of a regulatory region (e.g. an LTR, a sequence within an LTR) and a coding sequence (e.g. polynucleotide encoding a gene editing agent, etc.) to be transcribed so that the regulatory region is effective for regulating transcription or translation of the coding sequence of interest. In some embodiments disclosed herein, the term “operably linked” denotes a configuration in which a regulatory sequence is placed at an appropriate position relative to a sequence that encodes a polypeptide or functional RNA such that the control sequence directs or regulates the expression or cellular localization of the mRNA encoding the polypeptide, the polypeptide, and / or the functional RNA. Thus, operably linked elements may be contiguous or non-contiguous. In addition, in the context of a polypeptide, “operably linked” refers to a physical linkage (e.g., directly or indirectly linked) between amino acid sequences (e.g., different segments, modules, or domains) to provide for a described activity of the polypeptide. In the present disclosure, various segments, regions, or domains of the engineered antibodies disclosed herein may be operably linked to retain proper folding, processing, targeting, expression, binding, and other functional properties of the engineered antibodies in the cell. Operably linked regions, domains, and segments of the engineered antibodies of the disclosure may be contiguous or non-contiguous (e.g., linked to one another through a linker).
[0114] The terms “transfection”, “transduction”, “transfecting” or “transducing” can be used interchangeably and are defined as a process of introducing a nucleic acid molecule or a protein to a cell. Nucleic acids are introduced to a cell using non-viral or viral-based methods. The nucleic acid molecules may be gene sequences encoding complete proteins or functional portions thereof. Non-viral methods of transfection include any appropriate transfection method that does not use viral DNA or viral particles as a delivery system to introduce the nucleic acid molecule into the cell. Exemplary non-viral transfection methods include calcium phosphate transfection, liposomal transfection, nucleofection, sonoporation, transfection through heat shock, magnetifection and electroporation. In some embodiments, the nucleic acid molecules are introduced into a cell using electroporation following standard procedures well known in the art. For viral-based methods of transfection any useful viral vector may be used in the methods described herein. Examples for viral vectors include, but are not limited to retroviral, adenoviral, lentiviral and adeno-associated viral vectors. In some embodiments, the nucleic acid molecules are introduced into a cell using a lentiviral vector following standard procedures well known in the art. The terms “transfection” or “transduction” also refer to introducing proteins into a cell from the external environment. Typically, transduction or transfection of a protein relies on attachment of a peptide or protein capable of crossing the cell membrane to the protein of interest. See, e.g., Ford et al. (2001) Gene Therapy 8:1-4 and Prochiantz (2007) Nat. Methods 4:119-20.
[0115] “Transduce” or “transduction” are used according to their plain ordinary meanings and refer to the process by which one or more foreign nucleic acids (i.e. DNA not naturally found in the cell) are introduced into a cell. Typically, transduction occurs by introduction of a virus or viral vector (e.g. a CMV vector, a lentivirus vector, etc.) into the cell.
[0116] As used herein, the term “promoter” refers to a sequence of DNA which proteins bind to initiate gene expression. For example, transcription factors may bind a promoter region of a gene to transcribe RNA from DNA. In embodiments, the HTLV-1 LRT functions as a promoter for the HBZ gene.
[0117] “Contacting” is used in accordance with its plain ordinary meaning and refers to the process of allowing at least two distinct species (e.g. chemical compounds including biomolecules or cells) to become sufficiently proximal to react, interact or physically touch. It should be appreciated; however, the resulting reaction product can be produced directly from a reaction between the added reagents or from an intermediate from one or more of the added reagents that can be produced in the reaction mixture.
[0118] The term “contacting” may include allowing two species to react, interact, or physically touch, wherein the two species may be, for example, a nucleic acid as provided herein and a cell. In embodiments contacting includes, for example, allowing a nucleic acid as described herein to interact with a cell. Thus, in embodiments, contacting includes allowing a nucleic acid to interact with a cell, thereby resulting in transduced cell. In embodiments contacting includes, for example, allowing a pharmaceutical composition as described herein to interact with a cell.
[0119] A “cell” as used herein, refers to a cell carrying out metabolic or other function sufficient to preserve or replicate its genomic DNA. A cell can be identified by well-known methods in the art including, for example, presence of an intact membrane, staining by a particular dye, ability to produce progeny or, in the case of a gamete, ability to combine with a second gamete to produce a viable offspring. Cells may include prokaryotic and eukaryotic cells. Prokaryotic cells include but are not limited to bacteria. Eukaryotic cells include but are not limited to yeast cells and cells derived from plants and animals, for example mammalian, insect (e.g., spodoptera) and human cells. Cells may be useful when they are naturally nonadherent or have been treated not to adhere to surfaces, for example by trypsinization.
[0120] The terms “virus” or “virus particle” are used according to its plain ordinary meaning within Virology and refers to a virion including the viral genome (e.g. DNA, RNA, single strand, double strand), viral capsid and associated proteins, and in the case of enveloped viruses (e.g. herpesvirus), an envelope including lipids and optionally components of host cell membranes, and / or viral proteins.
[0121] The term “replicate” is used in accordance with its plain ordinary meaning and refers to the ability of a cell or virus to produce progeny. A person of ordinary skill in the art will immediately understand that the term replicate when used in connection with DNA, refers to the biological process of producing two identical replicas of DNA from one original DNA molecule.
[0122] In the context of a virus, the term “replicate” includes the ability of a virus to replicate (duplicate the viral genome and packaging said genome into viral particles) in a host cell and subsequently release progeny viruses from the host cell, which results in the lysis of the host cell.
[0123] The term “recombinant” when used with reference, e.g., to a cell, nucleic acid, protein, or vector, indicates that the cell, nucleic acid, protein or vector, has been modified by the introduction of a heterologous nucleic acid or protein or the alteration of a native nucleic acid or protein, or that the cell is derived from a cell so modified. Thus, for example, recombinant cells express proteins that are not found within the native (non-recombinant) form of the cell.
[0124] The term “isolated”, when applied to a nucleic acid or protein, denotes that the nucleic acid or protein is essentially free of other cellular components with which it is associated in the natural state. It can be, for example, in a homogeneous state and may be in either a dry or aqueous solution. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. A protein that is the predominant species present in a preparation is substantially purified.
[0125] The term “heterologous” when used with reference to portions of a nucleic acid indicates that the nucleic acid comprises two or more subsequences that are not found in the same relationship to each other in nature. For instance, the nucleic acid is typically recombinantly produced, having two or more sequences from unrelated genes arranged to make a new functional nucleic acid, e.g., a promoter from one source and a coding region from another source. Similarly, a heterologous protein indicates that the protein comprises two or more subsequences that are not found in the same relationship to each other in nature (e.g., a fusion protein).
[0126] The term “exogenous” refers to a molecule or substance (e.g., a compound, nucleic acid or protein) that originates from outside a given cell or organism. For example, an “exogenous promoter” as referred to herein is a promoter that does not originate from the cell or organism it is expressed by. Conversely, the term “endogenous” or “endogenous promoter” refers to a molecule or substance that is native to, or originates within, a given cell or organism.
[0127] The term “inhibition”, “inhibit”, “inhibiting” and the like in reference to a protein-inhibitor interaction means negatively affecting (e.g. decreasing) the activity or function of the protein relative to the activity or function of the protein in the absence of the inhibitor. In aspects inhibition means negatively affecting (e.g. decreasing) the concentration or levels of the protein relative to the concentration or level of the protein in the absence of the inhibitor. In aspects inhibition refers to reduction of a disease or symptoms of disease. In aspects, inhibition refers to a reduction in the activity of a particular protein target. Thus, inhibition includes, at least in part, partially or totally blocking stimulation, decreasing, preventing, or delaying activation, or inactivating, desensitizing, or down-regulating signal transduction or enzymatic activity or the amount of a protein. In aspects, inhibition refers to a reduction of activity of a target protein resulting from a direct interaction (e.g. an inhibitor binds to the target protein). In aspects, inhibition refers to a reduction of activity of a target protein from an indirect interaction (e.g. an inhibitor binds to a protein that activates the target protein, thereby preventing target protein activation).
[0128] The terms “inhibitor,”“repressor” or “antagonist” or “downregulator” interchangeably refer to a substance capable of detectably decreasing the expression or activity of a given gene or protein. The antagonist can decrease expression or activity 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 90% or more in comparison to a control in the absence of the antagonist. In certain instances, expression or activity is 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold or lower than the expression or activity in the absence of the antagonist.
[0129] The term “expression” includes any step involved in the production of the polypeptide including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, and secretion. Expression can be detected using conventional techniques for detecting protein (e.g., ELISA, Western blotting, flow cytometry, immunofluorescence, immunohistochemistry, etc.).
[0130] “Biological sample” or “sample” refer to materials obtained from or derived from a subject or patient. A biological sample includes sections of tissues such as biopsy and autopsy samples, and frozen sections taken for histological purposes. Such samples include bodily fluids such as blood and blood fractions or products (e.g., serum, plasma, platelets, red blood cells, and the like), sputum, tissue, cultured cells (e.g., primary cultures, explants, and transformed cells) stool, urine, synovial fluid, joint tissue, synovial tissue, synoviocytes, fibroblast-like synoviocytes, macrophage-like synoviocytes, immune cells, hematopoietic cells, fibroblasts, macrophages, T cells, etc. A biological sample is typically obtained from a eukaryotic organism, such as a mammal such as a primate e.g., chimpanzee or human; cow; dog; cat; a rodent, e.g., guinea pig, rat, mouse; rabbit; or a bird; reptile; or fish.
[0131] “Control” or “control experiment” is used in accordance with its plain ordinary meaning and refers to an experiment in which the subjects or reagents of the experiment are treated as in a parallel experiment except for omission of a procedure, reagent, or variable of the experiment. In some instances, the control is used as a standard of comparison in evaluating experimental effects. In some embodiments, a control is the measurement of the activity of a protein in the absence of a compound as described herein (including embodiments and examples).
[0132] A “control” or “standard control” refers to a sample, measurement, or value that serves as a reference, usually a known reference, for comparison to a test sample, measurement, or value. For example, a test sample can be taken from a patient suspected of having a given disease (e.g. cancer) and compared to a known normal (non-diseased) individual (e.g. a standard control subject). A standard control can also represent an average measurement or value gathered from a population of similar individuals (e.g. standard control subjects) that do not have a given disease (i.e. standard control population), e.g., healthy individuals with a similar medical background, same age, weight, etc. A standard control value can also be obtained from the same individual, e.g. from an earlier-obtained sample from the patient prior to disease onset. For example, a control can be devised to compare therapeutic benefit based on pharmacological data (e.g., half-life) or therapeutic measures (e.g., comparison of side effects). Controls are also valuable for determining the significance of data. For example, if values for a given parameter are widely variant in controls, variation in test samples will not be considered as significant. One of skill will recognize that standard controls can be designed for assessment of any number of parameters (e.g. RNA levels, protein levels, specific cell types, specific bodily fluids, specific tissues, etc).
[0133] One of skill in the art will understand which standard controls are most appropriate in a given situation and be able to analyze data based on comparisons to standard control values. Standard controls are also valuable for determining the significance (e.g. statistical significance) of data. For example, if values for a given parameter are widely variant in standard controls, variation in test samples will not be considered as significant.
[0134] “Patient”, “subject” or “subject in need thereof” refers to a living organism suffering from or prone to a disease or condition that can be treated by administration of a pharmaceutical composition as provided herein. Non-limiting examples include humans, other mammals, bovines, rats, mice, dogs, monkeys, goat, sheep, cows, deer, and other non-mammalian animals. In some embodiments, a patient is human.
[0135] The terms “disease” or “condition” refer to a state of being or health status of a patient or subject capable of being treated with the compounds or methods provided herein. The disease may be a human T-cell lymphotropic virus type 1 (HTLV-1) associated disease. The HTLV-1 associated disease may be adult T-cell leukemia, adult T-cell lymphoma, HTLV-1 associated myelopathy, tropical spastic paraparesis, or HTLV-1 infection.
[0136] The term “associated” or “associated with” in the context of a substance or substance activity or function associated with a disease means that the disease (e.g. adult T-cell leukemia, adult T-cell lymphoma, HTLV-1 Associated Myelopathy, Tropical spastic paraparesis, HTLV-1 infection) is caused by (in whole or in part), or a symptom of the disease is caused by (in whole or in part) the substance or substance activity or function. For example, an HTLV-1 associated disease may be caused by HTVL-1 infection. As used herein, what is described as being associated with a disease, if a causative agent, could be a target for treatment of the disease.
[0137] The term “aberrant” as used herein refers to different from normal. When used to describe enzymatic activity or protein function, aberrant refers to activity or function that is greater or less than a normal control or the average of normal non-diseased control samples. Aberrant activity may refer to an amount of activity that results in a disease, wherein returning the aberrant activity to a normal or non-disease-associated amount (e.g. by administering a compound or using a method as described herein), results in reduction of the disease or one or more disease symptoms.
[0138] The terms “treating”, or “treatment” refers to any indicia of success in the therapy or amelioration of an injury, disease, pathology or condition, including any objective or subjective parameter such as abatement; remission; diminishing of symptoms or making the injury, pathology or condition more tolerable to the patient; slowing in the rate of degeneration or decline; making the final point of degeneration less debilitating; improving a patient's physical or mental well-being. The treatment or amelioration of symptoms can be based on objective or subjective parameters; including the results of a physical examination, neuropsychiatric exams, and / or a psychiatric evaluation. The term “treating” and conjugations thereof, may include prevention of an injury, pathology, condition, or disease. In embodiments, treating is preventing. In embodiments, treating does not include preventing.
[0139] “Treating” or “treatment” as used herein (and as well-understood in the art) also broadly includes any approach for obtaining beneficial or desired results in a subject's condition, including clinical results. Beneficial or desired clinical results can include, but are not limited to, alleviation or amelioration of one or more symptoms or conditions, diminishment of the extent of a disease, stabilizing (i.e., not worsening) the state of disease, prevention of a disease's transmission or spread, delay or slowing of disease progression, amelioration or palliation of the disease state, diminishment of the reoccurrence of disease, and remission, whether partial or total and whether detectable or undetectable. In other words, “treatment” as used herein includes any cure, amelioration, or prevention of a disease. Treatment may prevent the disease from occurring; inhibit the disease's spread; relieve the disease's symptoms, fully or partially remove the disease's underlying cause, shorten a disease's duration, or do a combination of these things. Thus in the disclosed method, treatment can refer to a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 90%, or 100% reduction in the severity of an established disease, condition, or symptom of the disease or condition. For example, a method for treating a disease is considered to be a treatment if there is a 10% reduction in one or more symptoms of the disease in a subject as compared to a control. Thus the reduction can be a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 90%, 100%, or any percent reduction in between 10% and 100% as compared to native or control levels. It is understood that treatment does not necessarily refer to a cure or complete ablation of the disease, condition, or symptoms of the disease or condition. Further, as used herein, references to decreasing, reducing, or inhibiting include a change of 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 90% or greater as compared to a control level and such terms can include but do not necessarily include complete elimination.
[0140] “Treating” and “treatment” as used herein include prophylactic treatment. Treatment methods include administering to a subject a therapeutically effective amount of an active agent. The administering step may consist of a single administration or may include a series of administrations. The length of the treatment period depends on a variety of factors, such as the severity of the condition, the age of the patient, the concentration of active agent, the activity of the compositions used in the treatment, or a combination thereof. It will also be appreciated that the effective dosage of an agent used for the treatment or prophylaxis may increase or decrease over the course of a particular treatment or prophylaxis regime. Changes in dosage may result and become apparent by standard diagnostic assays known in the art. In some instances, chronic administration may be required. For example, the compositions are administered to the subject in an amount and for a duration sufficient to treat the patient. In embodiments, the treating or treatment is not prophylactic treatment.
[0141] The term “prevent” refers to a decrease in the occurrence of disease symptoms in a patient. As indicated above, the prevention may be complete (no detectable symptoms) or partial, such that fewer symptoms are observed than would likely occur absent treatment.
[0142] As used herein, the term “administering” is used in accordance with its plain and ordinary meaning and includes oral administration, administration as a suppository, topical contact, intravenous, parenteral, intraperitoneal, intramuscular, intralesional, intrathecal, intranasal or subcutaneous administration, or the implantation of a slow-release device, e.g., a mini-osmotic pump, to a subject. Administration is by any route, including parenteral and transmucosal (e.g., buccal, sublingual, palatal, gingival, nasal, vaginal, rectal, or transdermal). Parenteral administration includes, e.g., intravenous, intramuscular, intra-arteriole, intradermal, subcutaneous, intraperitoneal, intraventricular, and intracranial. Other modes of delivery include, but are not limited to, the use of liposomal formulations, intravenous infusion, transdermal patches, etc. In embodiments, the administering does not include administration of any active agent other than the recited active agent.
[0143] “Co-administer” it is meant that a composition described herein is administered at the same time, just prior to, or just after the administration of one or more additional therapies. The compounds provided herein can be administered alone or can be coadministered to the patient. Co-administration is meant to include simultaneous or sequential administration of the compounds individually or in combination (more than one compound). Thus, the preparations can also be combined, when desired, with other active substances (e.g., to reduce metabolic degradation). The compositions of the present disclosure can be delivered transdermally, by a topical route, or formulated as applicator sticks, solutions, suspensions, emulsions, gels, creams, ointments, pastes, jellies, paints, powders, and aerosols.
[0144] “Pharmaceutically acceptable excipient” and “pharmaceutically acceptable carrier” refer to a substance that aids the administration of an active agent to and absorption by a subject and can be included in the compositions of the present disclosure without causing a significant adverse toxicological effect on the patient. Non-limiting examples of pharmaceutically acceptable excipients include water, NaCl, normal saline solutions, lactated Ringer's, normal sucrose, normal glucose, binders, fillers, disintegrants, lubricants, coatings, sweeteners, flavors, salt solutions (such as Ringer's solution), alcohols, oils, gelatins, carbohydrates such as lactose, amylose or starch, fatty acid esters, hydroxymethycellulose, polyvinyl pyrrolidine, and colors, and the like. Such preparations can be sterilized and, if desired, mixed with auxiliary agents such as lubricants, preservatives, stabilizers, wetting agents, emulsifiers, salts for influencing osmotic pressure, buffers, coloring, and / or aromatic substances and the like that do not deleteriously react with the compounds of the disclosure. One of skill in the art will recognize that other pharmaceutical excipients are useful in the present disclosure.
[0145] A “therapeutic agent” as used herein refers to an agent (e.g., compound or composition described herein) that when administered to a subject will have the intended prophylactic effect, e.g., preventing or delaying the onset (or reoccurrence) of an injury, disease, pathology or condition, or reducing the likelihood of the onset (or reoccurrence) of an injury, disease, pathology, or condition, or their symptoms or the intended therapeutic effect, e.g., treatment or amelioration of an injury, disease, pathology or condition, or their symptoms including any objective or subjective parameter of treatment such as abatement; remission; diminishing of symptoms or making the injury, pathology or condition more tolerable to the patient; slowing in the rate of degeneration or decline; making the final point of degeneration less debilitating; or improving a patient's physical or mental well-being.
[0146] It is understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes in light thereof will be suggested to persons skilled in the art and are to be included within the spirit and purview of this application and scope of the appended claims. All publications, patents, and patent applications cited herein are hereby incorporated by reference in their entirety for all purposes.Zinc Finger Containing Proteins
[0147] Provided herein, inter alia, are compositions including a protein having a zinc finger domain where the zinc finger domain binds a sequence within the long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1). Applicant has discovered that binding of the zinc finger domain to the sequence within the HTLV-1 LTR potently suppresses HTLV-1 bZIP factor (HBZ) expression. The term “zinc finger domain” refers to a protein, or a domain within a larger protein, that binds DNA in a sequence-specific manner through one or more zinc fingers. Zinc fingers are regions of amino acid sequences whose structure is typically stabilized through coordination of a metal (e.g. a zinc ion). In embodiments, a zinc finger may adopt a structure including an antiparallel β sheet followed by an α helix. In embodiments, a zinc finger includes an antiparallel β sheet including two β strands followed by an α helix. Any of the zinc finger domains described herein may include 1, 2, 3, 4, 5, 6 or more zinc fingers, each zinc finger having a recognition helix region that binds a sequence within the LTR of HTLV-1. In embodiments, the zinc finger domain includes 4, 5 or 6 zinc fingers. In embodiments, the zinc finger domain includes 4 zinc fingers. In embodiments, the zinc finger domain includes 5 zinc fingers. In embodiments, the zinc finger domain includes 6 zinc fingers. In embodiments, the individual zinc fingers include zinc finger recognition helix regions (e.g. recognition helix regions), wherein the zinc finger recognition helix regions are designated F1, F2, F3, F4, F5 and F6, and include the amino acid sequences of the recognition helix regions as shown in Table 4. As used herein, zinc finger recognition helix region (e.g. recognition helix region), refers to a subportion of the zinc finger that makes specific contacts with a target nucleic acid sequence (e.g. a sequence within the HTLV-1 LTR). For example, a zinc finger recognition helix region may be a sequence within an α-helix structure within the zinc finger that makes specific contacts with a target nucleic acid sequence (e.g. a sequence within the HTLV-1 LTR).
[0148] In embodiments, the zinc finger domain is non-naturally occurring in that it is engineered to bind to a target site of choice. There is generally a wide range of sequence variation in the amino acids of the known zinc finger domains. In embodiments, a zinc finger domain has a sequence of the form X3-Cys-X2-4-Cys-Xu-His-X3-5-His-X4, wherein X is any amino acid (e.g., X24 indicates an oligopeptide 2-4 amino acids in length). In embodiments, only the two consensus histidine residues and two consensus cysteine residues bound to the central zinc atom are invariant. Of the remaining residues, typically three to five are highly conserved, while there may be significant variation among the other residues. Despite the wide range of sequence variation in zinc finger domains, zinc finger domains of this type generally have a similar three dimensional structure. However, there is a wide range of binding specificities among the different zinc finger domains, i.e., different zinc fingers may bind double stranded polynucleotides having a wide range of nucleotides sequences. In embodiments, the zinc finger domain is the C2H2 type. In embodiments, the zinc finger domain is the CCHC type. In embodiments, the zinc finger domain is the PHD type. In embodiments, the zinc finger domain is the RING type.
[0149] In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases, about 4 bases, about 5 bases, about 6 bases, about 7 bases, about 8 bases, about 9 bases, about 10 bases, about 11 bases, about 12 bases, about 13 bases, about 14 bases, about 15 bases, about 16 bases, about 18 bases, about 20 bases, about 22 bases, about 24 bases, about 26 bases, about 28 bases, about 30 bases, about 32 bases, about 34 bases, about 36 bases, about 38 bases, or about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 3 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 4 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 5 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 6 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 7 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 8 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 9 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 10 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 12 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 14 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 16 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 18 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 20 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 22 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 24 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 26 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 28 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 30 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 32 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 34 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 36 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 38 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes (e.g. binds to) a derivative of the target sequence which has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identify to the target sequence (e.g. a sequence within the HTLV-1 LTR).
[0150] In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 6 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 9 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 12 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 15 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 18 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 21 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 24 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 27 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 30 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 33 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 36 bases to about 40 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR).
[0151] In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 36 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 33 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 30 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 27 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 24 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 21 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 18 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 15 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 12 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 9 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR). In embodiments, the zinc finger domain recognizes with specificity (e.g. specifically binds) about 3 bases to about 6 bases of a recognized sequence (e.g. a sequence within the HTLV-1 LTR).
[0152] Long terminal repeats (LTRs) are used according to their plain and ordinary meaning and the art. Thus, LTR's may contain identical sequences of DNA or RNA that repeat tens, and more often hundreds or thousands of times found at either end of viral retroviral genome or proviral DNA that is formed by reverse transcription of retroviral RNA. LTRs may be used by viruses to insert their genetic material into the host genomes. The LTRs may be partially transcribed into an RNA intermediate, followed by reverse transcription into complementary DNA (cDNA) and ultimately dsDNA (double-stranded DNA) with full LTRs. The LTRs may then mediate integration of the retroviral DNA via an LTR specific integrase into another region of the host chromosome. In the proviral latency, once the provirus has been integrated, the LTR on the 5′ end may serve as the promoter for the entire retroviral genome, while the LTR at the 3′ end may provide for nascent viral RNA polyadenylation and encodes some accessory proteins. In embodiments, the protein provided herein including embodiments thereof targets (or binds to) a sequence within the 5′ LTR, 3′ LTR or both. In embodiments, the protein provided herein including embodiments thereof binds to a sequence within the 3′LTR.
[0153] Thus, in an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:27. In embodiments, the sequence has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% nucleic acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 5, 10, 11, 12, 13, 14, or 15 continuous nucleic acid portion) of SEQ ID NO:27.
[0154] In embodiments, the sequence within the HTLV-1 LTR has at least 75% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has at least 80% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has at least 85% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has at least 90% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has at least 95% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has at least 98% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has at least 99% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR includes the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR is the sequence of SEQ ID NO:27.
[0155] In embodiments, the sequence within the HTLV-1 LTR has about 75% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has about 80% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has about 85% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has about 90% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has about 95% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has about 98% sequence identity to the sequence of SEQ ID NO:27. In embodiments, the sequence within the HTLV-1 LTR has about 99% sequence identity to the sequence of SEQ ID NO:27.
[0156] In embodiments, the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
[0157] In embodiments, the zinc finger domain includes six zinc finger recognition helix regions designated F1 to F6, wherein F1 includes SEQ ID NO:51, F2 includes SEQ ID NO:52, F3 includes SEQ ID NO:53, F4 includes SEQ ID NO:54, F5 includes SEQ ID NO:55 and F6 includes SEQ ID NO:56. In embodiments, the F1 is SEQ ID NO:51, F2 is SEQ ID NO:52, F3 is SEQ ID NO:53, F4 is SEQ ID NO:54, F5 is SEQ ID NO:55 and F6 is SEQ ID NO:56.
[0158] In embodiments, the zinc finger domain has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:4. In embodiments, the zinc finger domain has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, or 160 continuous amino acid portion) of SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 75% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 80% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 85% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 90% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 91% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 92% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 93% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 94% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 95% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 96% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 97% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 98% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having at least 99% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes the sequence of SEQ ID NO:4. In embodiments, the zinc finger domain is the sequence of SEQ ID NO:4.
[0159] In embodiments, the zinc finger domain includes a sequence having about 75% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 80% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 85% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 90% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 91% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 92% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 93% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 94% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 95% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 96% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 97% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 98% sequence identity to SEQ ID NO:4. In embodiments, the zinc finger domain includes a sequence having about 99% sequence identity to SEQ ID NO:4.
[0160] In embodiments, the zinc finger domain has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. A “non-contiguous sequence” as provided herein refers to a sequence including one or more sequence fragments having no sequence identity to the indicated sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:4 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:4 through a sequence fragment having no sequence identity to SEQ ID NO:4. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:4 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:4.
[0161] In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:4. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:4.
[0162] The sequence of SEQ ID NO:4 encodes a non-naturally occurring peptide sequence, which may be referred to herein as “HTLV-ZFP-5” or “ZFP-5”.
[0163] In embodiments, the protein further includes a transcriptional repressor. The term “transcriptional repressor” refers to a protein that decreases gene transcription of a gene or set of genes. For example, transcriptional repressors may be DNA-binding proteins that bind to promoter-proximal elements, including the HTLV-1 LTR or sequences within the HTLV-1 LTR. The transcriptional repressors used in the fusion proteins described herein include, but are not limited to, Kruppel associated box (KRAB) domains, methyl CpG binding protein 2 (meCP2), DNA methyltransferase (DNMT) domains and derivatives or functional fragments thereof.
[0164] In embodiments, the transcriptional repressor includes a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof. In embodiments, the transcriptional repressor includes a KRAB domain. In embodiments, the transcriptional repressor includes meCP2 or a fragment thereof. In embodiments, the transcriptional repressor includes a DNMT domain. In embodiments, the transcriptional repressor includes a KRAB domain and meCP2 or a fragment thereof.
[0165] In embodiments, the protein of the present disclosure includes further components, including, but are not limited to, a cell-penetrating peptide (e.g. a TAT peptide or a derivative thereof) and / or one or more nuclear localization signals. In embodiments, the protein includes a peptide that promotes stabilization of the protein and / or enhances protein isolation (e.g. myc-tag sequence).
[0166] Cell-penetrating peptides (CPPs) generally are short peptides that can facilitate cellular intake / uptake of various molecular equipment (e.g. a protein). The cargo is associated with the CPPs either through chemical linkage via covalent bonds or through non-covalent interactions. The function of the CPPs is to deliver the cargo into cells. Any peptide that is known to be capable of facilitating cellular uptake or have cell-penetrating activity can be used in the composition and methods of the disclosure. In embodiments, the CPP is trans-activating transcriptional activator (Tat) or a derivative thereof. In embodiments, Tat enhances the cellular intake / uptake of the protein into the cells. Thus, in embodiments, the protein provided herein further includes Tat. In embodiments, Tat includes a sequence having at least 80% sequence identity to SEQ ID NO:120. In embodiments, Tat includes a sequence having at least 90% sequence identity to SEQ ID NO: 120. In embodiments, Tat includes a sequence having at least 95% sequence identity to SEQ ID NO:120. In embodiments, Tat includes a sequence having at least 98% sequence identity to SEQ ID NO: 120. In embodiments, Tat includes a sequence having at least 99% sequence identity to SEQ ID NO:120. In embodiments, Tat includes the sequence of SEQ ID NO:20. In embodiments, Tat is SEQ ID NO:120.
[0167] A nuclear localization signal or sequence (NLS) is an amino acid sequence that tags a protein for import into the cell nucleus by nuclear transport. Any peptides that are known to be capable of nuclear localization activity can be used in the composition and methods provided herein including embodiments thereof. In embodiments, the protein provided herein includes one or more NLSs. In embodiments, the protein provided herein includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS. In embodiments, the NLS includes the sequence having at least 90% sequence identity to SEQ ID NO:121. In embodiments, the NLS includes the sequence of SEQ ID NO:121. In embodiments, the NLS is the sequence of SEQ ID NO:121. In embodiments, the NLS includes the sequence having at least 80% sequence identity to SEQ ID NO:124. In embodiments, the NLS includes the sequence having at least 90% sequence identity to SEQ ID NO:124. In embodiments, the NLS includes the sequence having at least 95% sequence identity to SEQ ID NO:124. In embodiments, the NLS includes the sequence having at least 98% sequence identity to SEQ ID NO:124. In embodiments, the NLS includes the sequence having at least 99% sequence identity to SEQ ID NO:124. In embodiments, the NLS includes the sequence of SEQ ID NO:124. In embodiments, the NLS is the sequence of SEQ ID NO:124.
[0168] In embodiments, the protein provided herein includes one or more additional sequences such as a myc-tag sequence. A myc tag is a polypeptide protein tag derived from the c-myc gene product. In embodiments, the myc tag is used for affinity chromatography (e.g. to isolate the protein provided herein including embodiments thereof from a non-homogenous composition). In embodiments, the Myc tag includes a sequence having at least 80% sequence identity to SEQ ID NO:122. In embodiments, the Myc tag includes a sequence having at least 90% sequence identity to SEQ ID NO:122. In embodiments, the Myc tag includes a sequence having at least 95% sequence identity to SEQ ID NO: 122. In embodiments, the Myc tag includes a sequence having at least 98% sequence identity to SEQ ID NO:122. In embodiments, the Myc tag includes a sequence having at least 99% sequence identity to SEQ ID NO:122. In embodiments, the Myc tag includes SEQ ID NO:122. In embodiments, the Myc tag is the sequence of SEQ ID NO:122.
[0169] Thus, in embodiments, the protein further includes a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof. In embodiments, the protein further includes a nuclear localization signal. In embodiments, the protein further includes a a Tat domain. In embodiments, the protein further includes a Myc tag.
[0170] In embodiments, the protein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO: 13. In embodiments, the protein has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, or 300 continuous amino acid portion) compared to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 75% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 80% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 85% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 90% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 91% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 92% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 93% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 94% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 95% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 96% sequence identity to SEQ ID NO: 13. In embodiments, the protein includes a sequence having at least 97% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 98% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having at least 99% sequence identity to SEQ ID NO:13. In embodiments, the protein includes the sequence of SEQ ID NO:13. In embodiments, the protein is the sequence of SEQ ID NO:13.
[0171] In embodiments, the protein includes a sequence having about 75% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 80% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 85% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 90% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 91% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 92% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 93% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 94% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 95% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 96% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 97% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 98% sequence identity to SEQ ID NO:13. In embodiments, the protein includes a sequence having about 99% sequence identity to SEQ ID NO:13.
[0172] In embodiments, the protein has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:13 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:13 through a sequence fragment having no sequence identity to SEQ ID NO:13. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:13 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:13.
[0173] In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 300 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 290 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 280 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 270 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 260 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 250 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 240 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 230 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 220 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 210 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 200 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 190 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 180 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:13. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:13.
[0174] In embodiments, the protein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:20. In embodiments, the protein has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, or 220 continuous amino acid portion) compared to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 75% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 80% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 85% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 90% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 91% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 92% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 93% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 94% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 95% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 96% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 97% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 98% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having at least 99% sequence identity to SEQ ID NO:20. In embodiments, the protein includes the sequence of SEQ ID NO:20. In embodiments, the protein is the sequence of SEQ ID NO:20.
[0175] In embodiments, the protein includes a sequence having about 75% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 80% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 85% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 90% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 91% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 92% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 93% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 94% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 95% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 96% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 97% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 98% sequence identity to SEQ ID NO:20. In embodiments, the protein includes a sequence having about 99% sequence identity to SEQ ID NO:20.
[0176] In embodiments, the protein has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:20 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:20 through a sequence fragment having no sequence identity to SEQ ID NO:20. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:20 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:20.
[0177] In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 220 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 210 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 200 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 190 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 180 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:20. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:20.
[0178] In embodiments, the protein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:21. In embodiments, the protein has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, or 300 continuous amino acid portion) compared to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 75% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 80% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 85% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 90% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 91% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 92% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 93% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 94% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 95% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 96% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 97% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 98% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having at least 99% sequence identity to SEQ ID NO:21. In embodiments, the protein includes the sequence of SEQ ID NO:21. In embodiments, the protein is the sequence of SEQ ID NO:21.
[0179] In embodiments, the protein includes a sequence having about 75% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 80% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 85% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 90% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 91% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 92% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 93% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 94% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 95% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 96% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 97% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 98% sequence identity to SEQ ID NO:21. In embodiments, the protein includes a sequence having about 99% sequence identity to SEQ ID NO:21.
[0180] In embodiments, the protein has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:21 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:21 through a sequence fragment having no sequence identity to SEQ ID NO:21. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:21 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:21.
[0181] In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 330 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 320 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 310 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 300 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 290 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 280 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 270 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 260 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 250 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 240 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 230 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 220 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 210 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 200 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 190 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 180 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:21. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:21.
[0182] In embodiments, the protein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:22. In embodiments, the protein has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 420, 440, 460, 480, 500, 520, 540, 560, 580, or 600 continuous amino acid portion) compared to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 75% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 80% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 85% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 90% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 91% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 92% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 93% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 94% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 95% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 96% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 97% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 98% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having at least 99% sequence identity to SEQ ID NO:22. In embodiments, the protein includes the sequence of SEQ ID NO:22. In embodiments, the protein is the sequence of SEQ ID NO:22.
[0183] In embodiments, the protein includes a sequence having about 75% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 80% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 85% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 90% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 91% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 92% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 93% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 94% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 95% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 96% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 97% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 98% sequence identity to SEQ ID NO:22. In embodiments, the protein includes a sequence having about 99% sequence identity to SEQ ID NO:22.
[0184] In embodiments, the protein has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:22 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:22 through a sequence fragment having no sequence identity to SEQ ID NO:22. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:22 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:22.
[0185] In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 600 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 590 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 580 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 570 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 560 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 550 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 540 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 530 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 520 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 510 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 500 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 490 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 480 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 470 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 460 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 450 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 440 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 430 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 420 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 410 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 400 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 390 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 380 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 370 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 360 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 350 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 340 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 330 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 320 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 310 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 300 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 290 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 280 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 270 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 260 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 250 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 240 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 230 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 220 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 210 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 200 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 190 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 180 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:22. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:22.
[0186] In embodiments, the protein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:23. In embodiments, the protein has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 420, 440, 460, 480, 500, 520, 540, 560, 580, 600, 620, 640, 660, 680, 700, 720, 740, 760, 780, or 800 continuous amino acid portion) compared to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 75% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 80% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 85% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 90% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 91% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 92% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 93% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 94% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 95% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 96% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 97% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 98% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having at least 99% sequence identity to SEQ ID NO:23. In embodiments, the protein includes the sequence of SEQ ID NO:23. In embodiments, the protein is the sequence of SEQ ID NO:23.
[0187] In embodiments, the protein includes a sequence having about 75% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 80% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 85% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 90% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 91% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 92% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 93% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 94% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 95% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 96% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 97% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 98% sequence identity to SEQ ID NO:23. In embodiments, the protein includes a sequence having about 99% sequence identity to SEQ ID NO:23.
[0188] In embodiments, the protein has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:23 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:23 through a sequence fragment having no sequence identity to SEQ ID NO:23. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:23 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:23.
[0189] In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 810 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 800 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 790 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 780 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 770 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 760 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 750 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 740 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 730 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 720 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 710 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 700 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 690 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 680 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 670 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 660 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 650 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 640 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 630 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 620 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 610 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 600 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 590 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 580 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 570 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 560 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 550 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 540 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 530 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 520 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 510 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 500 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 490 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 480 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 470 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 460 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 450 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 440 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 430 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 420 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 410 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 400 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 390 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 380 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 370 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 360 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 350 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 340 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 330 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 320 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 310 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 300 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 290 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 280 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 270 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 260 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 250 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 240 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 230 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 220 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 210 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 200 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 190 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 180 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:23. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:23.
[0190] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:25. In embodiments, the sequence has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% nucleic acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 5, 10, 11, 12, 13, 14, or 15 continuous nucleic acid portion) of SEQ ID NO:25.
[0191] In embodiments, the sequence within the HTLV-1 LTR has at least 75% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has at least 80% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has at least 85% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has at least 90% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has at least 95% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has at least 98% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has at least 99% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR includes the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR is the sequence of SEQ ID NO:25.
[0192] In embodiments, the sequence within the HTLV-1 LTR has about 75% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has about 80% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has about 85% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has about 90% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has about 95% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has about 98% sequence identity to the sequence of SEQ ID NO:25. In embodiments, the sequence within the HTLV-1 LTR has about 99% sequence identity to the sequence of SEQ ID NO:25.
[0193] In embodiments, the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
[0194] In embodiments, the zinc finger domain includes six zinc finger recognition helix regions designated F1 to F6, wherein F1 includes SEQ ID NO:39, F2 includes SEQ ID NO:40, F3 includes SEQ ID NO:41, F4 includes SEQ ID NO:42, F5 includes SEQ ID NO:43 and F6 includes SEQ ID NO:44. In embodiments, F1 is SEQ ID NO:39, F2 is SEQ ID NO:40, F3 is SEQ ID NO:41, F4 is SEQ ID NO:42, F5 is SEQ ID NO:43 and F6 is SEQ ID NO:44.
[0195] In embodiments, the zinc finger domain has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:2. In embodiments, the zinc finger domain has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, or 160 continuous amino acid portion) of SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having at least 75% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain has includes a sequence having at least 80% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain has includes a sequence having at least 85% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain has includes a sequence having at least 90% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having at least 91% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having at least 92% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having at least 93% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having at least 94% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain has includes a sequence having at least 95% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having at least 96% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having at least 97% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having at least 98% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having at least 99% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes the sequence of SEQ ID NO:2. In embodiments, the zinc finger domain is the sequence of SEQ ID NO:2.
[0196] In embodiments, the zinc finger domain includes a sequence having about 75% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 80% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 85% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 90% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 91% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 92% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 93% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 94% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 95% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 96% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 97% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 98% sequence identity to SEQ ID NO:2. In embodiments, the zinc finger domain includes a sequence having about 99% sequence identity to SEQ ID NO:2.
[0197] In embodiments, the zinc finger domain has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:2 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:2 through a sequence fragment having no sequence identity to SEQ ID NO:2. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:2 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:2.
[0198] In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:2. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:2.
[0199] The sequence of SEQ ID NO:2 encodes a non-naturally occurring peptide sequence, which may be referred to herein as “HTLV-ZFP-3” or “ZFP-3”.
[0200] In embodiments, the protein further includes a transcriptional repressor. In embodiments, the transcriptional repressor includes a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof. In embodiments, the transcriptional repressor includes a KRAB domain. In embodiments, the transcriptional repressor includes meCP2 or a fragment thereof. In embodiments, the transcriptional repressor includes a DNMT domain. In embodiments, the transcriptional repressor includes a KRAB domain and meCP2 or a fragment thereof.
[0201] In embodiments, the protein further includes a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof. In embodiments, the protein further includes a nuclear localization signal. In embodiments, the protein further includes a a Tat domain. In embodiments, the protein further includes a Myc tag.
[0202] In embodiments, the protein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO: 11. In embodiments, the protein has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, or 300 continuous amino acid portion) compared to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 75% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 80% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 85% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 90% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 91% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 92% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 93% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 94% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 95% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 96% sequence identity to SEQ ID NO: 11. In embodiments, the protein includes a sequence having at least 97% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 98% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having at least 99% sequence identity to SEQ ID NO:11. In embodiments, the protein includes the sequence of SEQ ID NO:11. In embodiments, the protein is the sequence of SEQ ID NO:11.
[0203] In embodiments, the protein includes a sequence having about 75% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 80% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 85% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 90% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 91% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 92% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 93% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 94% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 95% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 96% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 97% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 98% sequence identity to SEQ ID NO:11. In embodiments, the protein includes a sequence having about 99% sequence identity to SEQ ID NO:11.
[0204] In embodiments, the protein has a sequence with the percentage sequence identity as disclosed above, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:11 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:11 through a sequence fragment having no sequence identity to SEQ ID NO: 11. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:11 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:11.
[0205] In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 300 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 290 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 280 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 270 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 260 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 250 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 240 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 230 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 220 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 210 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 200 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 190 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 180 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:11. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:11.
[0206] In embodiments, the protein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO: 19. In embodiments, the protein has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, or 220 continuous amino acid portion) compared to SEQ ID NO:19. In embodiments, the protein includes a sequence having at least 75% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having at least 80% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having at least 85% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having at least 90% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having at least 91% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having at least 92% sequence identity to SEQ ID NO: 19. In embodiments, the protein includes a sequence having at least 93% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having at least 94% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having at least 95% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having at least 96% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having at least 97% sequence identity to SEQ ID NO: 19. In embodiments, the protein includes a sequence having at least 98% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having at least 99% sequence identity to SEQ ID NO:19. In embodiments, the protein includes the sequence of SEQ ID NO:19. In embodiments, the protein is the sequence of SEQ ID NO:19.
[0207] In embodiments, the protein includes a sequence having about 75% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 80% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 85% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 90% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 91% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 92% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 93% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 94% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 95% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 96% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 97% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 98% sequence identity to SEQ ID NO:19. In embodiments, the protein includes a sequence having about 99% sequence identity to SEQ ID NO:19.
[0208] In embodiments, the protein has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:19 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:19 through a sequence fragment having no sequence identity to SEQ ID NO:19. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:19 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:19.
[0209] In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 220 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 210 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 200 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 190 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 180 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:19. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:19.
[0210] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:28. In embodiments, the sequence has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% nucleic acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 5, 10, 15 or 20 continuous nucleic acid portion) of SEQ ID NO:28.
[0211] In embodiments, the sequence within the HTLV-1 LTR has at least 75% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has at least 80% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has at least 85% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has at least 90% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has at least 95% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has at least 98% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has at least 99% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR includes the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR is the sequence of SEQ ID NO:28.
[0212] In embodiments, the sequence within the HTLV-1 LTR has about 75% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has about 80% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has about 85% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has about 90% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has about 95% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has about 98% sequence identity to the sequence of SEQ ID NO:28. In embodiments, the sequence within the HTLV-1 LTR has about 99% sequence identity to the sequence of SEQ ID NO:28.
[0213] In embodiments, the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
[0214] In embodiments, the zinc finger domain includes six zinc finger recognition helix regions designated F1 to F6, wherein F1 includes SEQ ID NO:57, F2 includes SEQ ID NO:58, F3 includes SEQ ID NO:59, F4 includes SEQ ID NO:60, F5 includes SEQ ID NO:61 and F6 includes SEQ ID NO:62. In embodiments, F1 is SEQ ID NO:57, F2 is SEQ ID NO:58, F3 is SEQ ID NO:59, F4 is SEQ ID NO:60, F5 is SEQ ID NO:61 and F6 is SEQ ID NO:62.
[0215] In embodiments, the zinc finger domain has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:5. In embodiments, the zinc finger domain has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, or 170 continuous amino acid portion) of SEQ ID NO: 5. In embodiments, the zinc finger domain has at least 75% sequence identity to SEQ ID NO: 5. In embodiments, the zinc finger domain has at least 80% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain has at least 85% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain has at least 90% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having at least 91% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having at least 92% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having at least 93% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having at least 94% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain has at least 95% sequence identity to SEQ ID NO: 5. In embodiments, the zinc finger domain has at least 96% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having at least 97% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain has at least 98% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain has at least 99% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes the sequence of SEQ ID NO:5. In embodiments, the zinc finger domain is the sequence of SEQ ID NO:5.
[0216] In embodiments, the zinc finger domain includes a sequence having about 75% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 80% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 85% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 90% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 91% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 92% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 93% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 94% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 95% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 96% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 97% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 98% sequence identity to SEQ ID NO:5. In embodiments, the zinc finger domain includes a sequence having about 99% sequence identity to SEQ ID NO:5.
[0217] In embodiments, the zinc finger domain has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:5 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:5 through a sequence fragment having no sequence identity to SEQ ID NO:5. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:5 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:5.
[0218] In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:5. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:5.
[0219] The sequence of SEQ ID NO:5 encodes a non-naturally occurring peptide sequence, which may be referred to herein as “HTLV-ZFP-6” or “ZFP-6”.
[0220] In embodiments, the protein further includes a transcriptional repressor. In embodiments, the transcriptional repressor includes a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof. In embodiments, the transcriptional repressor includes a KRAB domain. In embodiments, the transcriptional repressor includes meCP2 or a fragment thereof. In embodiments, the transcriptional repressor includes a DNMT domain.
[0221] In embodiments, the protein includes a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof. In embodiments, the protein further includes a nuclear localization signal. In embodiments, the protein further includes a Tat domain. In embodiments, the protein further includes a Myc tag.
[0222] In embodiments, the protein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO: 14. In embodiments, the protein has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, or 300 continuous amino acid portion) compared to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 75% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 80% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 85% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 90% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 91% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 92% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 93% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 94% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 95% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 96% sequence identity to SEQ ID NO: 14. In embodiments, the protein includes a sequence having at least 97% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 98% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having at least 99% sequence identity to SEQ ID NO:14. In embodiments, the protein includes the sequence of SEQ ID NO:14. In embodiments, the protein is the sequence of SEQ ID NO:14.
[0223] In embodiments, the protein includes a sequence having about 75% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 80% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 85% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 90% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 91% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 92% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 93% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 94% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 95% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 96% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 97% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 98% sequence identity to SEQ ID NO:14. In embodiments, the protein includes a sequence having about 99% sequence identity to SEQ ID NO:14.
[0224] In embodiments, the protein has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:14 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:14 through a sequence fragment having no sequence identity to SEQ ID NO:14. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:14 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:14.
[0225] In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 300 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 290 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 280 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 270 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 260 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 250 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 240 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 230 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 220 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 210 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 200 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 190 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 180 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:14. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:14.
[0226] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:32. In embodiments, the sequence has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% nucleic acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 5, 10, 15 or 20 continuous nucleic acid portion) of SEQ ID NO:32.
[0227] In embodiments, the sequence within the HTLV-1 LTR has at least 75% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has at least 80% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has at least 85% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has at least 90% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has at least 95% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has at least 98% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has at least 99% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR includes the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR is the sequence of SEQ ID NO:32.
[0228] In embodiments, the sequence within the HTLV-1 LTR has about 75% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has about 80% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has about 85% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has about 90% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has about 95% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has about 98% sequence identity to the sequence of SEQ ID NO:32. In embodiments, the sequence within the HTLV-1 LTR has about 99% sequence identity to the sequence of SEQ ID NO:32.
[0229] In embodiments, the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
[0230] In embodiments, the zinc finger domain includes six zinc finger recognition helix regions designated F1 to F6, wherein F1 includes SEQ ID NO:81, F2 includes SEQ ID NO:82, F3 includes SEQ ID NO:83, F4 includes SEQ ID NO:84, F5 includes SEQ ID NO:85 and F6 includes SEQ ID NO:86. In embodiments, F1 is SEQ ID NO:81, F2 is SEQ ID NO:82, F3 is SEQ ID NO:83, F4 is SEQ ID NO:84, F5 is SEQ ID NO:85 and F6 is SEQ ID NO:86.
[0231] In embodiments, the zinc finger domain has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:9. In embodiments, the zinc finger domain has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, or 160 continuous amino acid portion) of SEQ ID NO:9. In embodiments, the zinc finger domain has at least 75% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain has at least 80% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain has at least 85% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain has at least 90% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having at least 91% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having at least 92% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having at least 93% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having at least 94% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain has at least 95% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain has at least 96% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having at least 97% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain has at least 98% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain has at least 99% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes the sequence of SEQ ID NO:9. In embodiments, the zinc finger domain is the sequence of SEQ ID NO:9.
[0232] In embodiments, the zinc finger domain includes a sequence having about 75% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 80% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 85% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 90% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 91% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 92% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 93% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 94% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 95% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 96% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 97% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 98% sequence identity to SEQ ID NO:9. In embodiments, the zinc finger domain includes a sequence having about 99% sequence identity to SEQ ID NO:9.
[0233] In embodiments, the zinc finger domain has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:9 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:9 through a sequence fragment having no sequence identity to SEQ ID NO:9. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:9 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:9.
[0234] In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:9. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:9.
[0235] The sequence of SEQ ID NO:9 encodes a non-naturally occurring peptide sequence, which may be referred to herein as “HTLV-ZFP-10” or “ZFP-10”.
[0236] In embodiments, the protein further includes a transcriptional repressor. In embodiments, the transcriptional repressor includes a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof. In embodiments, the transcriptional repressor includes a KRAB domain. In embodiments, the transcriptional repressor includes meCP2 or a fragment thereof. In embodiments, the transcriptional repressor includes a DNMT domain. In embodiments, the transcriptional repressor includes a KRAB domain and meCP2 or a fragment thereof.
[0237] In embodiments, the protein further includes a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof. In embodiments, the protein further includes a nuclear localization signal. In embodiments, the protein further includes a Tat domain. In embodiments, the protein further includes a Myc tag.
[0238] In embodiments, the protein has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, or 300 continuous amino acid portion) compared to SEQ ID NO:18. In embodiments, the protein includes a sequence having at least 75% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having at least 80% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having at least 85% sequence identity to SEQ ID NO: 18. In embodiments, the protein includes a sequence having at least 90% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having at least 91% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having at least 92% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having at least 93% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having at least 94% sequence identity to SEQ ID NO: 18. In embodiments, the protein includes a sequence having at least 95% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having at least 96% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having at least 97% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having at least 98% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having at least 99% sequence identity to SEQ ID NO: 18. In embodiments, the protein includes the sequence of SEQ ID NO:18. In embodiments, the protein is the sequence of SEQ ID NO:18.
[0239] In embodiments, the protein includes a sequence having about 75% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 80% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 85% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 90% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 91% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 92% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 93% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 94% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 95% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 96% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 97% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 98% sequence identity to SEQ ID NO:18. In embodiments, the protein includes a sequence having about 99% sequence identity to SEQ ID NO:18.
[0240] In embodiments, the protein has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:18 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:18 through a sequence fragment having no sequence identity to SEQ ID NO:18. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:18 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:18.
[0241] In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 300 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 290 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 280 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 270 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 260 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 250 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 240 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 230 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 220 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 210 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 200 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 190 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 180 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:18. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:18.
[0242] In another aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:31. In embodiments, the sequence has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% nucleic acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 5, 10, 11, 12, 13, 14, or 15 continuous nucleic acid portion) of SEQ ID NO:31.
[0243] In embodiments, the sequence within the HTLV-1 LTR has at least 75% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has at least 80% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has at least 85% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has at least 90% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has at least 95% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has at least 98% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has at least 99% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR includes the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR is the sequence of SEQ ID NO:31.
[0244] In embodiments, the sequence within the HTLV-1 LTR has about 75% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has about 80% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has about 85% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has about 90% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has about 95% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has about 98% sequence identity to the sequence of SEQ ID NO:31. In embodiments, the sequence within the HTLV-1 LTR has about 99% sequence identity to the sequence of SEQ ID NO:31.
[0245] In embodiments, the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
[0246] In embodiments, the zinc finger domain includes six zinc finger recognition helix regions designated F1 to F6, wherein F1 includes SEQ ID NO:75, F2 includes SEQ ID NO:76, F3 includes SEQ ID NO:77, F4 includes SEQ ID NO:78, F5 includes SEQ ID NO:79 and F6 includes SEQ ID NO:80. In embodiments, the F1 is SEQ ID NO:75, F2 is SEQ ID NO:76, F3 is SEQ ID NO:77, F4 is SEQ ID NO:78, F5 is SEQ ID NO:79 and F6 is SEQ ID NO:80.
[0247] In embodiments, the zinc finger domain has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:8. In embodiments, the zinc finger domain has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, or 160 continuous amino acid portion) of SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 75% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 80% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 85% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 90% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 91% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 92% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 93% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 94% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 95% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 96% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 97% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 98% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having at least 99% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes the sequence of SEQ ID NO: 8. In embodiments, the zinc finger domain is the sequence of SEQ ID NO:8.
[0248] In embodiments, the zinc finger domain includes a sequence having about 75% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 80% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 85% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 90% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 91% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 92% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 93% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 94% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 95% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 96% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 97% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 98% sequence identity to SEQ ID NO:8. In embodiments, the zinc finger domain includes a sequence having about 99% sequence identity to SEQ ID NO:8.
[0249] In embodiments, the zinc finger domain has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:8 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:8 through a sequence fragment having no sequence identity to SEQ ID NO:8. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:8 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:8.
[0250] In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:8. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:8.
[0251] The sequence of SEQ ID NO:8 encodes a non-naturally occurring peptide sequence, which may be referred to herein as “HTLV-ZFP-9” or “ZFP-9”.
[0252] In embodiments, the protein further includes a transcriptional repressor. In embodiments, the transcriptional repressor includes a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof. In embodiments, the transcriptional repressor includes a KRAB domain. In embodiments, the transcriptional repressor includes meCP2 or a fragment thereof. In embodiments, the transcriptional repressor includes a KRAB domain and meCP2 or a fragment thereof.
[0253] In embodiments, the protein further includes a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof. In embodiments, the protein further includes a nuclear localization signal. In embodiments, the protein further includes a a Tat domain. In embodiments, the protein further includes a Myc tag.
[0254] In embodiments, the protein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO: 17. In embodiments, the protein has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, or 300 continuous amino acid portion) compared to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 75% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 80% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 85% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 90% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 91% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 92% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 93% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 94% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 95% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 96% sequence identity to SEQ ID NO: 17. In embodiments, the protein includes a sequence having at least 97% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 98% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having at least 99% sequence identity to SEQ ID NO:17. In embodiments, the protein includes the sequence of SEQ ID NO:17. In embodiments, the protein is the sequence of SEQ ID NO:17.
[0255] In embodiments, the protein includes a sequence having about 75% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 80% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 85% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 90% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 91% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 92% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 93% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 94% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 95% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 96% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 97% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 98% sequence identity to SEQ ID NO:17. In embodiments, the protein includes a sequence having about 99% sequence identity to SEQ ID NO:17.
[0256] In embodiments, the protein has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:17 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:17 through a sequence fragment having no sequence identity to SEQ ID NO:17. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:17 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:17.
[0257] In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 300 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 290 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 280 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 270 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 260 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 250 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 240 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 230 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 220 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 210 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 200 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 190 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:17. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:17.
[0258] In another aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:30. In embodiments, the sequence has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% nucleic acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 5, 10, 11, 12, 13, 14, or 15 continuous nucleic acid portion) of SEQ ID NO:30.
[0259] In embodiments, the sequence within the HTLV-1 LTR has at least 75% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has at least 80% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has at least 85% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has at least 90% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has at least 95% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has at least 98% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has at least 99% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR includes the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR is the sequence of SEQ ID NO:30.
[0260] In embodiments, the sequence within the HTLV-1 LTR has about 75% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has about 80% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has about 85% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has about 90% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has about 95% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has about 98% sequence identity to the sequence of SEQ ID NO:30. In embodiments, the sequence within the HTLV-1 LTR has about 99% sequence identity to the sequence of SEQ ID NO:30.
[0261] In embodiments, the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
[0262] In embodiments, the zinc finger domain includes six zinc finger recognition helix regions designated F1 to F6, wherein F1 includes SEQ ID NO:69, F2 includes SEQ ID NO:70, F3 includes SEQ ID NO:71, F4 includes SEQ ID NO:72, F5 includes SEQ ID NO:73 and F6 includes SEQ ID NO:74. In embodiments, the F1 is SEQ ID NO:69, F2 is SEQ ID NO:70, F3 is SEQ ID NO:71, F4 is SEQ ID NO:72, F5 is SEQ ID NO:73 and F6 is SEQ ID NO:74.
[0263] In embodiments, the zinc finger domain has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:7. In embodiments, the zinc finger domain has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, or 160 continuous amino acid portion) of SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 75% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 80% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 85% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 90% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 91% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 92% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 93% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 94% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 95% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 96% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 97% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 98% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having at least 99% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes the sequence of SEQ ID NO:7. In embodiments, the zinc finger domain is the sequence of SEQ ID NO:7.
[0264] In embodiments, the zinc finger domain includes a sequence having about 75% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 80% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 85% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 90% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 91% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 92% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 93% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 94% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 95% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 96% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 97% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 98% sequence identity to SEQ ID NO:7. In embodiments, the zinc finger domain includes a sequence having about 99% sequence identity to SEQ ID NO:7.
[0265] In embodiments, the zinc finger domain has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:7 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:7 through a sequence fragment having no sequence identity to SEQ ID NO:7. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:7 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:7.
[0266] In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:7. In embodiments, the zinc finger domain has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:7.
[0267] The sequence of SEQ ID NO:7 encodes a non-naturally occurring peptide sequence, which may be referred to herein as “HTLV-ZFP-8” or “ZFP-8”.
[0268] In embodiments, the protein further includes a transcriptional repressor. In embodiments, the transcriptional repressor includes a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof. In embodiments, the transcriptional repressor includes a KRAB domain. In embodiments, the transcriptional repressor includes meCP2 or a fragment thereof. In embodiments, the transcriptional repressor includes a KRAB domain and meCP2 or a fragment thereof.
[0269] In embodiments, the protein further includes a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof. In embodiments, the protein further includes a nuclear localization signal. In embodiments, the protein further includes a a Tat domain. In embodiments, the protein further includes a Myc tag.
[0270] In embodiments, the protein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO: 16. In embodiments, the protein has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, 160, 180, 200, 220, 240, 260, 280, or 300 continuous amino acid portion) compared to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 75% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 80% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 85% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 90% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 91% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 92% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 93% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 94% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 95% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 96% sequence identity to SEQ ID NO: 16. In embodiments, the protein includes a sequence having at least 97% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 98% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having at least 99% sequence identity to SEQ ID NO:16. In embodiments, the protein includes the sequence of SEQ ID NO:16. In embodiments, the protein is the sequence of SEQ ID NO:16.
[0271] In embodiments, the protein includes a sequence having about 75% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 80% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 85% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 90% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 91% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 92% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 93% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 94% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 95% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 96% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 97% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 98% sequence identity to SEQ ID NO:16. In embodiments, the protein includes a sequence having about 99% sequence identity to SEQ ID NO:16.
[0272] In embodiments, the protein has a sequence with the percentage sequence identity as disclosed in the above paragraphs, and the sequence having the percentage sequence identity as disclosed above is a non-contiguous sequence. In embodiments, the non-contiguous sequence is a sequence including a first sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:16 connected to a second sequence fragment having at least the percentage sequence identity as disclosed above to SEQ ID NO:16 through a sequence fragment having no sequence identity to SEQ ID NO:16. In embodiments, the non-contiguous sequence is a sequence including a plurality of sequence fragments having at least the percentage sequence identity as disclosed above to SEQ ID NO:16 connected through a plurality of sequence fragments having no sequence identity to SEQ ID NO:16.
[0273] In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 300 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 290 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 280 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 270 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 260 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 250 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 240 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 230 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 220 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 210 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 200 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 190 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 180 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 170 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 160 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 150 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 140 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 130 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 120 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 110 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 100 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 90 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 80 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 70 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 60 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 50 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 40 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 30 continuous amino acids of SEQ ID NO:16. In embodiments, the protein has the percentage sequence identity as disclosed in the above paragraphs to at least 20 continuous amino acids of SEQ ID NO:16.
[0274] In an aspect is provided a protein including a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:24. In embodiments, the sequence has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% nucleic acid sequence identity across the whole sequence or a portion of the sequence (e.g. a 5, 10, 11, 12, 13, 14, or 15 continuous nucleic acid portion) of SEQ ID NO:24.
[0275] In embodiments, the sequence within the HTLV-1 LTR has at least 75% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has at least 80% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has at least 85% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has at least 90% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has at least 95% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has at least 98% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has at least 99% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR includes the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR is the sequence of SEQ ID NO:24.
[0276] In embodiments, the sequence within the HTLV-1 LTR has about 75% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has about 80% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has about 85% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has about 90% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has about 95% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has about 98% sequence identity to the sequence of SEQ ID NO:24. In embodiments, the sequence within the HTLV-1 LTR has about 99% sequence identity to the sequence of SEQ ID NO:24.
[0277] In embodiments, the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
[0278] In embodiments, the zinc finger domain includes six zinc finger recognition helix regions designated F1 to F6, wherein F1 includes SEQ ID NO:33, F2 includes SEQ ID NO:34, F3 includes SEQ ID NO:35, F4 includes SEQ ID NO:36, F5 includes SEQ ID NO:37 and F6 includes SEQ ID NO:38. In embodiments, the F1 is SEQ ID NO:33, F2 is SEQ ID NO:34, F3 is SEQ ID NO:35, F4 is SEQ ID NO:36, F5 is SEQ ID NO:37 and F6 is SEQ ID NO:38.
[0279] In embodiments, the zinc finger domain has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity to the sequence of SEQ ID NO:1. In embodiments, the zinc finger domain has at least 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity across the whole sequence or a portion of the sequence (e.g. a 20, 40, 60, 80, 100, 120, 140, or 160 continuous amino acid portion) of SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 75% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 80% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 85% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 90% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 91% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 92% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 93% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 94% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 95% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 96% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 97% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 98% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having at least 99% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes the sequence of SEQ ID NO: 1. In embodiments, the zinc finger domain is the sequence of SEQ ID NO:1.
[0280] In embodiments, the zinc finger domain includes a sequence having about 75% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having about 80% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having about 85% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having about 90% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having about 91% sequence identity to SEQ ID NO:1. In embodiments, the zinc finger domain includes a sequence having about 92% sequence identity to SEQ ID NO:1. In embodi...
Examples
embodiments
[0344]Embodiment 1. A protein comprising a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:27.
[0345]Embodiment 2. The protein of embodiment 1, wherein the sequence within the HTLV-1 LTR comprises the sequence of SEQ ID NO:27.
[0346]Embodiment 3. The protein of embodiment 1 or 2, wherein the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
[0347]Embodiment 4. The protein of any one of embodiments 1-3, wherein the zinc finger domain comprises six zinc finger recognition helix regions designated F1 to F6, wherein F1 comprises SEQ ID NO:51, F2 comprises SEQ ID NO:52, F3 comprises SEQ ID NO:53, F4 comprises SEQ ID NO:54, F5 comprises SEQ ID NO:55 and F6 comprises SEQ ID NO:56.
[0348]Embodiment 5. The protein of any one of embodiments 1-4, wherein the zinc finger domain comprises a s...
example 1
Targeted Zinc-Finger Repressors to the Oncogenic HBZ Gene Inhibit Acute T-Cell Leukeamia (ATL) Proliferation
Introduction
[0476]Human T-lymphotropic virus type I (HTLV-I) largely infects CD4+ T-cells resulting in a latent, life-long infection in patients. Crosstalk between oncogenic viral factors results in the transformation of the host cell into an aggressive cancer, acute T-cell leukemia / lymphoma (ATL). ATL has a very poor prognosis with no currently available effective treatments, urging the development of novel therapeutic strategies. Recent evidence exploring the mechanisms contributing to ATL highlights the viral anti-sense gene HTLV-1 bZIP factor (HBZ) as a tumor driver and a potential therapeutic target. The cys2his2 zinc-finger proteins (ZFPs) are abundant endogenous regulatory proteins that bind specific DNA motifs to control gene expression. As a result of well-characterized rules for DNA motif recognition, custom zinc-finger arrays can be generated to target unique sequen...
example 2
EV Delivery of a Zinc Finger Protein to Direct Killing of Human T-Cell Leukemia Virus Type 1 Transformed Cancer Cells
Introduction
[0522]HTLV-1 infects T-cells (Yoshie, 2008 #4489) and the persistent expression of the HTLV-1 HBZ gene plays a part in the oncogenic transformation and maintenance of HTLV-1-infected cells in vivo, while also inducing increased CCR4 expression known to augment disease pathology (Matsuoka, 2011 #4488). A methodology that can target the specific inhibition of HBZ can lead to a loss of those cells transformed by HTLV-1 and presumably a cure for HTLV-1 associated disease. We show that HTLV-1 transformed T-cells can be specifically targeted and killed by a newly developed anti-HTLV HBZ gene targeted zinc finger protein repressor containing a fusion of KRAB and meCP2 epigenetic regulatory proteins (ZFP5-KrMe) delivered to virus transformed CCR4 over-expressing T-cells by targeted extracellular vesicles. We develop and characterize a highly innovative next-genera...
Claims
1. A protein comprising a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:27.
2. The protein of claim 1, wherein the sequence within the HTLV-1 LTR comprises the sequence of SEQ ID NO:27.
3. The protein of claim 1, wherein the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
4. The protein of claim 1, wherein the zinc finger domain comprises six zinc finger recognition helix regions designated F1 to F6, wherein F1 comprises SEQ ID NO:51, F2 comprises SEQ ID NO:52, F3 comprises SEQ ID NO:53, F4 comprises SEQ ID NO:54, F5 comprises SEQ ID NO:55 and F6 comprises SEQ ID NO:56.
5. The protein of claim 1, wherein the zinc finger domain comprises a sequence having at least 75% sequence identity to SEQ ID NO:4.
6. The protein of claim 5, wherein the zinc finger domain comprises the sequence of SEQ ID NO:4.
7. The protein of claim 1, wherein the protein further comprises a transcriptional repressor.
8. The protein of claim 7, wherein the transcriptional repressor comprises a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof.
9. The protein of claim 8, wherein the transcriptional repressor comprises a KRAB domain.
10. The protein of claim 8, wherein the transcriptional repressor comprises a KRAB domain and meCP2.
11. The protein of claim 1, further comprising a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof.
12. The protein of claim 1, comprising a sequence having at least 75% sequence identity to SEQ ID NO:13, 20, 21, 22, or 23.
13. The protein of claim 12, comprising the sequence of SEQ ID NO:13, 20, 21, 22, or 23.
14. A protein comprising a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:25.
15. The protein of claim 14, wherein the sequence within the HTLV-1 LTR comprises the sequence of SEQ ID NO:25.
16. The protein of claim 14, wherein the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
17. The protein of claim 14, wherein the zinc finger domain comprises six zinc finger recognition helix regions designated F1 to F6, wherein F1 comprises SEQ ID NO:39, F2 comprises SEQ ID NO:40, F3 comprises SEQ ID NO:41, F4 comprises SEQ ID NO:42, F5 comprises SEQ ID NO:43 and F6 comprises SEQ ID NO:44.
18. The protein of claim 14, wherein the zinc finger domain comprises a sequence having at least 75% sequence identity to SEQ ID NO:2.
19. The protein of claim 18, wherein the zinc finger domain comprises the sequence of SEQ ID NO:2.
20. The protein of claim 14, wherein the protein further comprises a transcriptional repressor.
21. The protein of claim 20, wherein the transcriptional repressor comprises a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof.
22. The protein of claim 21, wherein the transcriptional repressor comprises a KRAB domain.
23. The protein of claim 21, wherein the transcriptional repressor comprises a KRAB domain and mcCP2.
24. The protein of claim 14, further comprising a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof.
25. The protein of claim 14, comprising a sequence having at least 75% sequence identity to SEQ ID NO:11 or 19.
26. The protein of claim 25, comprising the sequence of SEQ ID NO:11 or 19.
27. A protein comprising a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:28.
28. The protein of claim 27, wherein the sequence within the HTLV-1 LTR comprises SEQ ID NO:28.
29. The protein of claim 27, wherein the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
30. The protein of claim 27, wherein the zinc finger domain comprises six zinc finger recognition helix regions designated F1 to F6, wherein F1 comprises SEQ ID NO:57, F2 comprises SEQ ID NO:58, F3 comprises SEQ ID NO:59, F4 comprises SEQ ID NO:60, F5 comprises SEQ ID NO:61 and F6 comprises SEQ ID NO:62.
31. The protein of claim 27, wherein the zinc finger domain comprises a sequence having at least 75% sequence identity to SEQ ID NO:5.
32. The protein of claim 31, wherein the zinc finger domain comprises the sequence of SEQ ID NO:5.
33. The protein of claim 27, wherein the protein further comprises a transcriptional repressor.
34. The protein of claim 33, wherein the transcriptional repressor comprises a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof.
35. The protein of claim 34, wherein the transcriptional repressor comprises a KRAB domain.
36. The protein of claim 34, wherein the transcriptional repressor comprises a KRAB domain and meCP2.
37. The protein of claim 27, further comprising a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof.
38. The protein of claim 27, comprising a sequence having at least 75% sequence identity to SEQ ID NO:14.
39. The protein of claim 38, comprising the sequence of SEQ ID NO:14.
40. A protein comprising a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:32.
41. The protein of claim 40, wherein the sequence within the HTLV-1 LTR comprises the sequence of SEQ ID NO:32.
42. The protein of claim 40, wherein the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
43. The protein of claim 40, wherein the zinc finger domain comprises six zinc finger recognition helix regions designated F1 to F6, wherein F1 comprises SEQ ID NO:81, F2 comprises SEQ ID NO:82, F3 comprises SEQ ID NO:83, F4 comprises SEQ ID NO:84, F5 comprises SEQ ID NO:85 and F6 comprises SEQ ID NO:86.
44. The protein of claim 40, wherein the zinc finger domain comprises a sequence having at least 75% sequence identity to SEQ ID NO:9.
45. The protein of claim 44, wherein the zinc finger domain comprises the sequence of SEQ ID NO:9.
46. The protein of claim 40, wherein the protein further comprises a transcriptional repressor.
47. The protein of claim 46, wherein the transcriptional repressor comprises a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof.
48. The protein of claim 47, wherein the transcriptional repressor comprises a KRAB domain.
49. The protein of claim 47, wherein the transcriptional repressor comprises a KRAB domain and meCP2.
50. The protein of claim 40, further comprising a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof.
51. The protein of claim 40, comprising a sequence having at least 75% sequence identity to SEQ ID NO:18.
52. The protein of claim 51, comprising the sequence of SEQ ID NO:18.
53. A protein comprising a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:31.
54. The protein of claim 53, wherein the sequence within the HTLV-1 LTR comprises the sequence of SEQ ID NO:31.
55. The protein of claim 53, wherein the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
56. The protein of claim 53, wherein the zinc finger domain comprises six zinc finger recognition helix regions designated F1 to F6, wherein F1 comprises SEQ ID NO:75, F2 comprises SEQ ID NO:76, F3 comprises SEQ ID NO:77, F4 comprises SEQ ID NO:78, F5 comprises SEQ ID NO:79 and F6 comprises SEQ ID NO:80.
57. The protein of claim 53, wherein the zinc finger domain comprises a sequence having at least 75% sequence identity to SEQ ID NO:8.
58. The protein of claim 57, wherein the zinc finger domain comprises the sequence of SEQ ID NO:8.
59. The protein of claim 53, wherein the protein further comprises a transcriptional repressor.
60. The protein of claim 59, wherein the transcriptional repressor comprises a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof.
61. The protein of claim 60, wherein the transcriptional repressor comprises a KRAB domain.
62. The protein of claim 60, wherein the transcriptional repressor comprises a KRAB domain and meCP2.
63. The protein of claim 53, further comprising a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof.
64. The protein of claim 53, comprising a sequence having at least 75% sequence identity to SEQ ID NO:17.
65. The protein of claim 64, comprising the sequence of SEQ ID NO:17.
66. A protein comprising a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:30.
67. The protein of claim 66, wherein the sequence within the HTLV-1 LTR comprises the sequence of SEQ ID NO:30.
68. The protein of claim 66, wherein the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
69. The protein of claim 66, wherein the zinc finger domain comprises six zinc finger recognition helix regions designated F1 to F6, wherein F1 comprises SEQ ID NO:69, F2 comprises SEQ ID NO:70, F3 comprises SEQ ID NO:71, F4 comprises SEQ ID NO:72, F5 comprises SEQ ID NO:73 and F6 comprises SEQ ID NO:74.
70. The protein of claim 66, wherein the zinc finger domain comprises a sequence having at least 75% sequence identity to SEQ ID NO:7.
71. The protein of claim 70, wherein the zinc finger domain comprises the sequence of SEQ ID NO:7.
72. The protein of claim 66, wherein the protein further comprises a transcriptional repressor.
73. The protein of claim 72, wherein the transcriptional repressor comprises a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof.
74. The protein of claim 73, wherein the transcriptional repressor comprises a KRAB domain.
75. The protein of claim 73, wherein the transcriptional repressor comprises a KRAB domain and meCP2.
76. The protein of claim 66, further comprising a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof.
77. The protein of claim 66, comprising a sequence having at least 75% sequence identity to SEQ ID NO:16.
78. The protein of claim 77, comprising the sequence of SEQ ID NO:16.
79. A protein comprising a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:24.
80. The protein of claim 79, wherein the sequence within the HTLV-1 LTR comprises the sequence of SEQ ID NO:24.
81. The protein of claim 79, wherein the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
82. The protein of claim 79, wherein the zinc finger domain comprises six zinc finger recognition helix regions designated F1 to F6, wherein F1 comprises SEQ ID NO:33, F2 comprises SEQ ID NO:34, F3 comprises SEQ ID NO:35, F4 comprises SEQ ID NO:36, F5 comprises SEQ ID NO:37 and F6 comprises SEQ ID NO:38.
83. The protein of claim 79, wherein the zinc finger domain comprises a sequence having at least 75% sequence identity to SEQ ID NO:1.
84. The protein of claim 83, wherein the zinc finger domain comprises the sequence of SEQ ID NO:1.
85. The protein of claim 79, wherein the protein further comprises a transcriptional repressor.
86. The protein of claim 85, wherein the transcriptional repressor comprises a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof.
87. The protein of claim 86, wherein the transcriptional repressor comprises a KRAB domain.
88. The protein of claim 86, wherein the transcriptional repressor comprises a KRAB domain and meCP2.
89. The protein of claim 79, further comprising a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof.
90. The protein of claim 79, comprising a sequence having at least 75% sequence identity to SEQ ID NO:10.
91. The protein of claim 90, comprising the sequence of SEQ ID NO:10.
92. A protein comprising a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:26.
93. The protein of claim 92, wherein the sequence within the HTLV-1 LTR comprises the sequence of SEQ ID NO:26.
94. The protein of claim 92, wherein the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
95. The protein of claim 92, wherein the zinc finger domain comprises six zinc finger recognition helix regions designated F1 to F6, wherein F1 comprises SEQ ID NO:45, F2 comprises SEQ ID NO:46, F3 comprises SEQ ID NO:47, F4 comprises SEQ ID NO:48, F5 comprises SEQ ID NO:49 and F6 comprises SEQ ID NO:50.
96. The protein of claim 92, wherein the zinc finger domain comprises a sequence having at least 75% sequence identity to SEQ ID NO:3.
97. The protein of claim 96, wherein the zinc finger domain comprises the sequence of SEQ ID NO:3.
98. The protein of claim 92, wherein the protein further comprises a transcriptional repressor.
99. The protein of claim 98, wherein the transcriptional repressor comprises a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof.
100. The protein of claim 99, wherein the transcriptional repressor comprises a KRAB domain.
101. The protein of claim 99, wherein the transcriptional repressor comprises a KRAB domain and meCP2.
102. The protein of claim 92, further comprising a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof.
103. The protein of claim 92, comprising a sequence having at least 75% sequence identity to SEQ ID NO:12.
104. The protein of claim 103, comprising the sequence of SEQ ID NO:12.
105. A protein comprising a zinc finger domain capable of binding a sequence within a long terminal repeat (LTR) of Human T-cell lymphotropic virus type 1 (HTLV-1), wherein the sequence has at least 75% sequence identity to SEQ ID NO:29.
106. The protein of claim 105, wherein the sequence within the HTLV-1 LTR comprises the sequence of SEQ ID NO:29.
107. The protein of claim 105, wherein the sequence within the HTLV-1 LTR is operably linked to a nucleic acid encoding HTLV-1 bZIP factor (HBZ).
108. The protein of claim 105, wherein the zinc finger domain comprises six zinc finger recognition helix regions designated F1 to F6, wherein F1 comprises SEQ ID NO:63, F2 comprises SEQ ID NO:64, F3 comprises SEQ ID NO:65, F4 comprises SEQ ID NO:66, F5 comprises SEQ ID NO:67 and F6 comprises SEQ ID NO:68.
109. The protein of claim 105, wherein the zinc finger domain comprises a sequence having at least 75% sequence identity to SEQ ID NO:6.
110. The protein of claim 109, wherein the zinc finger domain comprises the sequence of SEQ ID NO:6.
111. The protein of claim 105, wherein the protein further comprises a transcriptional repressor.
112. The protein of claim 111, wherein the transcriptional repressor comprises a Kruppel associated box (KRAB) domain, methyl CpG binding protein 2 (meCP2), a DNA methyltransferase (DNMT) domain, or combinations thereof.
113. The protein of claim 112, wherein the transcriptional repressor comprises a KRAB domain.
114. The protein of claim 112, wherein the transcriptional repressor comprises a KRAB domain and meCP2.
115. The protein of claim 105, further comprising a nuclear localization signal, a Tat domain, a Myc tag, or combinations thereof.
116. The protein of claim 105, comprising a sequence having at least 75% sequence identity to SEQ ID NO:15.
117. The protein of claim 116, comprising the sequence of SEQ ID NO:15.
118. A nucleic acid encoding the protein of claim 1.
119. A vector comprising the nucleic acid of claim 118.
120. An extracellular vesicle (EV) comprising a nucleic acid encoding the protein of claim 1.
121. The EV of claim 120, wherein the EV further comprises an EV membrane-associated protein and an oncogenic T-cell targeting protein.
122. The EV of claim 121, wherein the EV membrane-associated protein is CD63 or PTGFRN.
123. The EV of claim 121, wherein the oncogenic T-cell targeting protein is an anti-CCR4 antibody or fragment thereof.
124. The EV of claim 121, wherein the oncogenic T-cell targeting protein is fused to an extracellular portion of the EV membrane-associated protein.
125. A pharmaceutical composition comprising the protein of claim 1, the nucleic acid of claim 118, the vector of claim 119, or the EV of claim 120.
126. A cell comprising the protein of claim 1, the nucleic acid of claim 118, the vector of claim 119, or the EV of claim 120.
127. The cell of claim 126, wherein the cell is an oncogenic T-cell.
128. The cell of claim 127, wherein the oncogenic T-cell is an adult T-cell leukemia cell or an adult T-cell lymphoma cell.
129. A method of treating a human T-cell lymphotropic virus type 1 (HTLV-1) associated disease in a subject in need thereof, comprising administering to the subject an effective amount of the protein of claim 1, the nucleic acid of claim 118, the vector of claim 119, or the EV of claim 120.
130. The method of claim 129, wherein the HTLV-1 associated disease is adult T-cell leukemia, adult T-cell lymphoma, HTLV-1 associated myelopathy, tropical spastic paraparesis, or HTLV-1 infection.
131. The method of claim 130, wherein the HTLV-1 associated disease is adult T-cell leukemia.
132. The method of claim 130, wherein the HTLV-1 associated disease is adult T-cell lymphoma.