Designer recombinases for precise removal of HTLV-1 provirus

By designing recombinases with specific amino acid substitutions, the problem of the difficulty in removing HTLV-1 virus was solved, achieving efficient and precise removal of HTLV-1 virus and showing significant antiviral activity.

JP2026508335APending Publication Date: 2026-03-10PROVIREX GENOME EDITING THERAPIES GMBH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-27
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively remove the proviral DNA of HTLV-1 virus, and the use of nuclease therapy may lead to the emergence of drug-resistant viral clones.

Method used

A recombinase with specific amino acid substitutions was designed to efficiently recognize and recombine the asymmetric target sequence of HTLV-1 virus. The engineered recombinase was used to specifically remove the proviral DNA of HTLV-1 virus in human cells.

Benefits of technology

It achieved efficient and precise removal of HTLV-1 virus, avoiding the generation of drug-resistant viral clones, and showed significant antiviral activity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026508335000001_ABST
    Figure 2026508335000001_ABST
Patent Text Reader

Abstract

The present invention relates to nucleic acids encoding a recombinase capable of recombining the asymmetric target sequence of SEQ ID NO: 1 within the long terminal repeats of multiple proviral DNAs. HTLV-1 strains contain recombinases with specific amino acid substitutions compared to the sequence of the cre recombinase. The present invention provides expression vectors containing the nucleic acids, cells containing the same, the recombinase in protein form, and pharmaceutical compositions containing them. Methods for providing such recombinases are also provided. The present invention also provides methods for treating subjects infected with HTLV-1, optionally with adult T-cell leukemia (ATL) and / or HTLV-1-associated myelopathy (HAM).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a nucleic acid encoding a recombinase capable of recombining the asymmetric target sequence of SEQ ID NO: 1 within the long terminal repeats of the proviral DNA of multiple HTLV-1 strains, where the recombinase has specific amino acid substitutions compared to the sequence of the cre recombinase. The present invention provides expression vectors containing the nucleic acid, cells containing the same, the recombinase in protein form, and pharmaceutical compositions containing them. Methods for providing such recombinases are also provided. The present invention also provides methods for treating a subject infected with HTLV-1, where the subject may, optionally, have adult T-cell leukemia (ATL) and / or HTLV-1-associated myelopathy (HAM). [Background technology]

[0002] Retroviral infections remain a challenge for modern medicine, primarily due to the lack of effective therapies to treat them. Permanent integration of proviral DNA into the host cell genome results in lifelong infection. For this reason, of the more than 37 million people worldwide infected with human immunodeficiency virus (HIV), those who are eligible for treatment rely on lifelong drug treatments that suppress the virus but do not cure the infection. A promising approach to solving this problem is the use of programmable DNA nucleases to disrupt the virus. While some nuclease strategies aim to target viral entry receptors in host cells, others focus on directly disrupting viral genetic material. The use of designer nucleases to treat HIV-1 infection has become possible because the inserted proviral DNA serves as a target substrate for genomic scissors such as zinc finger nucleases (ZNFs), transcription activator-like effector nucleases (TALENs), or CRISPR / Cas9. However, nucleases cause double-strand breaks that are processed in an unpredictable manner, depending on the cellular DNA repair machinery. Indeed, the use of nucleases targeting viral sequences can lead to the rapid formation of resistant clones, undermining the utility of this approach (Wang et al., 2016a, 2016b, Yoder & Bundschuh 2016, De Silva Feelixge et al., 2016). Another class of genome editing tools, site-specific recombinases (SSRs), circumvent this problem because they do not rely on host DNA repair mechanisms and edit DNA in a seamless manner (Meinke et al., 2016). However, adapting SSRs to incorporate custom-selected sequences presents challenges. Nevertheless, directed evolution methods have proven useful for redirecting recombinase specificity to recognize new target sequences (Buchholz & Stewart, 2001; Santoro & Schultz, 2002; Eroshenko & Church, 2013).Indeed, designer recombinases have been developed as powerful tools for elimination of HIV-1 proviruses with high efficiency and precision (Hauber et al., 2013; Karpinski et al., 2016; Sarkar et al., 2007; WO 2008 / 08391, WO 2011 / 147590, WO 2016 / 034553). Importantly, these designer recombinases have demonstrated antiretroviral activity in humanized mouse models without the emergence of resistant clones, indicating that they may have significant therapeutic potential (Hauber et al., 2013; Karpinski et al., 2016).

[0003] Targeting infectious diseases with genome editing tools is not limited to the HIV virus; similar approaches can be used to treat other human retroviral infections. Human T-cell leukemia virus type 1 (HTLV-1) is a human pathogenic retrovirus that causes adult T-cell leukemia (ATL), an aggressive tumorigenesis in approximately 5% of carriers (Martin et al., 2018). Furthermore, a small percentage of infected individuals develop an inflammatory neurological disorder called HTLV-1-associated myelopathy / tropical spastic paraplegia (HAM / TSP). Approximately 10 million people worldwide are infected with HTLV-1, with the majority of cases concentrated in endemic regions such as Japan, Central Australia, South America, the Caribbean, and sub-Saharan Africa. The lack of effective treatments for ATL and HAM / TSP highlights the need for novel therapeutic approaches. Virus transmission is primarily via close cell-cell contact or cell-to-cell spread via cellular protrusions. Due to this infection mechanism and viral amplification through clonal proliferation, HTLV-1 has a remarkably stable genome, with little sequence variation between infected individuals and within cells (Furukawa et al., 2001; Wattel et al., 1995). The HTLV-1 proviral genome is approximately 9 kb and contains identical 5' and 3' long terminal repeats (Boxus & Willems, 2009). The stability of the viral genetic sequence and the fact that the provirus integrates into the genome of infected cells make HTLV-1 a suitable candidate for treatment with gene editing tools. Indeed, different groups have attempted to disrupt the virus using nucleases. Tanaka et al. used zinc finger nucleases to disrupt the LTR promoter, while Nakagawa et al. used CRISPR / Cas9 to target the Hbz gene. This resulted in reduced cell proliferation in three different ATL cell lines (Tanaka et al., 2013; Nakagawa et al., 2018). However, the use of nucleases may lead to the emergence of resistant HTLV-1 clones, similar to HIV-1. Summary of the Invention

[0004] In light of this, the inventors have addressed the problem of providing an advantageous enzyme capable of removing the HTLV-1 provirus from infected cells, which, for example, overcomes at least one of these drawbacks. This problem is solved by the subject matter of the claims.

[0005] In particular, the present invention provides nucleic acids encoding a tailored recombinase capable of recombining the asymmetric target sequence of SEQ ID NO: 1 within the long terminal repeat of the proviral DNA of multiple HTLV-1 strains, wherein the amino acid sequence of the tailored recombinase has at least 70% sequence identity with the sequence set forth in SEQ ID NO: 13, and further wherein the tailored recombinase contains at least an amino acid substitution selected from the group consisting of P15L, S38P, A84T, Q90K, S108G, K122R, I166V, A175S, N245Y, E266G, and T268A compared to SEQ ID NO: 11.

[0006] We identified a conserved loxP-like sequence (loxHTLV, SEQ ID NO: 1) present in the long terminal repeats of most virus isolates and thus could be advantageously used for proviral elimination. After 181 cycles of substrate-linked molecular evolution (SLiDE), we isolated designer recombinases that recognized loxHTLV and could efficiently recombine loxHTLV sequences with high specificity in bacterial and human cells. Expression of these recombinases in human T cells demonstrated antiviral activity against HTLV-1 infection. Furthermore, expression of these recombinases in chronically infected SP cells resulted in elimination of HTLV-1 proviral DNA. Our data indicate that recombinase-mediated elimination of HTLV-1 proviruses is a promising strategy for therapeutically reversing HTLV-1 infection.

[0007] Because the prototype recombinase on which the recombinases of the present invention were originally based is cre, the sequences of the claimed recombinases are provided relative to the sequence of cre in SEQ ID NO: 11. This applies when the substituted amino acid is defined (e.g., P15L means that P at position 15 of SEQ ID NO: L is substituted with L), or when only the position and the amino acid present at that position are specified (e.g., 15L also means that P at position 15 of SEQ ID NO: L is substituted with L). As defined herein, further amino acid substitutions can be made compared to the cre sequence, but they may be the same as in cre.

[0008] SEQ ID NO: 13 is the amino acid sequence of a particularly preferred hyperactive recombinase of the present invention, also called RecHTLV or G9 (used synonymously).

[0009] P15L, S38P, A84T, Q90K, S108G, K122R, I166V, A175S, N245Y, E266G, and T268A are present in all recombinases of the present invention but are absent in cre. S38P, A84T, Q90K, and E266G are substitutions not identified in the prior art and are believed to be important for recognizing the HTLV-1 target sequence. Other substitutions are also found in prior art recombinases Brec-1 and Tre, which have been confirmed to target HIV-1 asymmetric target sequences. These amino acid substitutions are considered essential for the recombinases of the present invention. The present invention provides nucleic acids, such as expression vectors, encoding all of the recombinases of the present invention.

[0010] As used herein, the term "recombinase" refers to a protein involved in recombination. Recombinases recognize and bind to two specific DNA sequences, called "recombination sites" or "target sites," and mediate recombination between these two target sites. Therefore, the term "recombinase" refers to any protein component of any recombination system that mediates DNA rearrangement at a specific DNA location. Naturally occurring recombinases recognize symmetric target sites, or "half sites," consisting of two identical sequences of approximately 9-20 base pairs. These half-site sequences are separated by a 5-12 base pair spacer sequence. Recombinases of the tyrosine integrase family are characterized by having tyrosine as the active site nucleophile used for DNA cleavage. On the other hand, recombinases of the serine integrase family use serine instead of tyrosine. The recombinases of the present invention have been modified or engineered to integrate asymmetric target sequences. Therefore, they are also referred to as designer recombinases or engineered recombinases.

[0011] Desirably, the nucleic acids of the invention encode a modified recombinase having at least 80%, preferably at least 90%, sequence identity to SEQ ID NO: 13 (i.e., G9 / RecHTLV). The sequence identity may be at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%. In a preferred embodiment, the recombinase comprises SEQ ID NO: 13 (G9 / RecHTLV).

[0012] Other recombinases of the invention are E5_E9 (also referred to as E5), A3_A9_C6, A11_E1, B11, B12, C2, D7, D11_F7_F9_G8, F6, F11, F12, G7, G10, and H2 (SEQ ID NOs: 14 and 26-37, 39).

[0013] Preferably, the tuned recombinase of the present invention further comprises SEQ ID NO: 38. SEQ ID NO: 38 is the consensus sequence of all recombinase clones of the present invention. SEQ ID NO: 38 has 343 amino acids and 101 variable positions X. Thus, the consensus sequence provides a guide to those skilled in the art as to which domains and positions of the recombinase are variable and which positions should be carefully mutated. X can be any naturally occurring amino acid. The amino acid at the mutated position is preferably a conservative substitution compared to the sequence of any of the clones described herein. That is, a specific amino acid is substituted with a different amino acid having similar biochemical properties (e.g., charge, hydrophobicity, size). For example, aliphatic amino acids, hydroxyl- or sulfur-containing amino acids, cyclic amino acids, aromatic amino acids, basic or acidic amino acids, or their amines can be substituted within these groups.

[0014] The recombinase preferably has an amino acid present at any of positions G9, E5_E9 (also referred to as E5), A3_A9_C6, A11_E1; B11, B12, C2, D7, D11_F7_F9_G8, F6, F11, F12, G7, G10, H2 (SEQ ID NOs: 13, 14 and 26-37, 39).

[0015] Optionally, in addition to the above substitutions, the modified recombinase further comprises at least one, preferably 2 to 101 (e.g., 5-90, 10-80, 20-70, 30-60, or 40-50) amino acid(s) defined below at the indicated positions (these positions are compared to or refer to SEQ ID NO: 11): 3K, 3N or 3D, preferably 3K; 5L, 5Q or 5P, preferably 5L; 7L or 7I, preferably 7L; 8H, 8P or 8Y, preferably 8H; 9Q or 9H, preferably 9Q; 10S or 10D or 10N, preferably 10S; 12S or 12F or 12T, preferably 12S; 14L or 14S, preferably 14L; 16A or 16V, preferably 16A; 17D or 17G or 17N, preferably 17D; 18A or 18V, preferably 18V; 19T or 19A or 19M, preferably 19T; 23A or 23T, preferably 23A; 25K, 25Q or 25R, preferably 25K; 26N or 26S, preferably 26N; 28M or 28T, preferably 28M; 29D or 29V or 29I, preferably 29D; 30M or 30T or 30V, preferably 30M; 31F or 31L, preferably 31F; 34R or 34H or 34C, preferably 34R 35Q or 35H, preferably 35Q; 39E or 39A, preferably 39E; 43K or 43E or 43R, preferably 43E; 45L or 45F, preferably 45L; 51S or 51A or 51T, preferably 51S; 54A or 54T, preferably 54A; 57E or 57K, preferably 57E; 58S or 58L, preferably 58S; 59N or 59D, preferably 59N; 62K or 62R, preferably 62K; 63W or 63R, preferably 63W; 68P or 68H, preferably 68P; 77H or 77Y or 77D, preferably 77H; 86N or 86S, preferably 86N; 88V or 88I, preferably 88V; 93A or 93T, preferably 83A; 94E or 94D, preferably 94E; 97T or 97M, preferably 97T; 110S or 110N, preferably 110S; 111N or 111S, preferably 111N; 114T or 114S or 114P, preferably 114T; 125V or 125I, preferably 125V; 132K or 132R, preferably 132K; 136A or 136P, preferably 136A; 140T or 140A, preferably 140T; 144R, 144Q or 144L, preferably 144R; 147S, 147T, 147P or 147A, preferably 147S; 149M or 149L, preferably 149M; 151N, 151T, 151D or 151S, preferably 151N; 155C or 155R, preferably 155C; 156L, 156K or 156Q, preferably 156L; 160N or 160D, preferably 160N; 183K or 183R, preferably 183K; 185I or 185V, preferably 185I; 196H or 196R, preferably 196H; 206T or 206A, preferably 206T; 207A or 207T, preferably 207A; 219K, 219R or 219X, preferably 219K; 221V or 221I, preferably 221V; 225I, 225V or 225F, preferably 225I; 226S or 226T, preferably 226S; 227V or 227A, preferably 227A; 229G or 229C, preferably 229G; 230V or 230I, preferably 230V; 231A or 231T, preferably 231A; 232D or 232N, preferably 232D; 235N or 235D, preferably 235N; 241R or 241Q, preferably 241R; 244R or 244K, preferably 244R; 247V or 247I, preferably 247V; 248V or 248A, preferably 248V; 249V or 249A, preferably 249V; 250S or 250P, preferably 250S; 253T or 253A, preferably 253T; 254S, 254K, 254N or 254C, preferably 254S; 255R or 255Q, preferably 255R; 258T or 258A, preferably 258T; 259H, 259Y or 259G, preferably 259H; 261M or 261L, preferably 261M; 262Q or 262R, preferably 262Q; 263G or 263K, preferably 263G; 264I or 264V, preferably 264I; 267A or 267T, preferably 257A; 272I or 272V, preferably 272I; 276K, 276R or 276N, preferably 276K; 277D or 277G, preferably 277D; 278D or 278G, preferably 278D; 280G or 280S, preferably 280G; 281Q, 281R or 281L, preferably 281Q; 299L or 299M, preferably 299L; 304V or 304A, preferably 304V; 305S or 305P, preferably 305S; 306I, 306M or 306L, preferably 306I; 307A, 307P or 307V, preferably 307A; 317R or 317T, preferably 317R; 319N or 319D, preferably 319N; 320S, 320T or 320G, preferably 320S; 332T or 332A, preferably 332T; 341D or 341G or 341E, preferably 341D; 342G or 342D, preferably 342G; and / or 343D or 343Y, preferably 343D.

[0016] Amino acids may be selected independently of one another. Amino acids present in the most active clone, G9, are considered preferred. A recombinase may contain at least one, and preferably 2-101 (e.g., 5-90, 10-80, 20-70, 30-60, or 40-50) preferred positions. Furthermore, it is generally desirable to have combinations of positions present together in an active recombinase.

[0017] Amino acid analysis of the different recombinases of the present invention shows that the preferred amino acid positions are 94E or 94D, with 94E being particularly preferred. These substitutions are not known in the art.

[0018] New combinations were identified at amino acid positions 247 and 248. Preferably, the engineered recombinase of the invention comprises: i) 247V and 248V and 244R, or ii) 247I and 248A Contains any of the following.

[0019] The recombinases of the present invention include 247V, 248V, and 244R, and exhibit particularly high activity. These recombinases preferably include 249V.

[0020] Preferably, the recombinase comprises a substitution at position N317, where cre has an N, i.e., the amino acid at position 317 is not N. It has been found that the position can be, for example, N317R or N317T. Preferably, N317R is desired.

[0021] The following mutations are present in both the preferred recombinases E5 and G9 / RecHTLV: P12S, P15L, V23A, S38P, Y77H, A84T, K86N, Q90K, G93A, Q94E, M97T, S108G, K122R, I166V, A175S, K244R, N245Y, A248V, A249V, R259H, L261M, E262Q, E266G, T268A, N317R or N317T, and I320S. These mutations are present in more than 75% of the sequenced clones. Therefore, they are considered to represent the most important mutations. These substitutions make the enzyme particularly suitable for recombination with the target sequence of SEQ ID NO: 1 or target sequences with high sequence identity (e.g., at least 80%, at least 90%, or at least 95%) to SEQ ID NO: 1.

[0022] The mutations shown in Figure 2C, namely V7L, P12S, P15L, V23A, S38P, Y77H, A84T, K86N, Q90K, G93A, Q94E, M97T, S108G, K122R, I166V, A175S, K244R, N245Y, A248V, A249V, R259H, L261M, E262Q, E266G, T268A, N317T, and I320S, are present in more than 75% of the sequenced clones. With the exception of N317T (G9 has N317R), these are also present in the highly active recombinase mutant G9 / RecHTLV (but not necessarily in E5). The present invention provides recombinases (and nucleic acids encoding them) containing these amino acid substitutions compared to cre.

[0023] The nucleic acid of the invention preferably encodes a recombinase containing at least 33 amino acid substitutions selected from the group consisting of N3K, N10D, P12S, P15L, V16A, V23A, S38P, K57E, L58S, Y77H, A84T, K86N, I88V, Q90K, G93A, Q94E, M97T, S108G, K122R, I166V, A175S, K244R, N245Y, A248V, A249V, Q255R, R259H, L261M, E262Q, E266G, T268A, P307A, N317R, and I320S, compared to SEQ ID NO: 11. These amino acid substitutions are present in both E5 and G9 / RecHTLV.

[0024] Preferably, the encoded regulated recombinase has at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 12. SEQ ID NO: 12 is the consensus sequence of two advantageous regulated recombinases provided by the present invention, E5 and G9 / RecHTLV. Thus, the recombinase may comprise the amino acid sequence set forth in SEQ ID NO: 12.

[0025] 10. Any of the nucleic acids according to claim 9, wherein the adjusted recombinase has any of the following sequences compared to SEQ ID NO: 11: N3K, V7L, N10D, P12S, P15L, V16A, A18V, V23A, S38P, K43E, K57E, L58S, Y77H, A84T, K86N, I88V, Q90K, G93A, Q94E, M97T, S108G, S114T, K122R, Q144R, The RecHTLV vector contains at least 34, and preferably all, amino acid substitutions selected from the group consisting of Q156L, I166V, A175S, T206A, S226T, V227A, K244R, N245Y, A248V, A249V, P250S, Q255R, R259H, L261M, E262Q, E266G, A267T, T268A, M299L, P307A, N317R, and I320S. All of these substitutions are contained in RecHTLV.

[0026] In particularly preferred embodiments, the encoded recombinase comprises the amino acid sequence of SEQ ID NO: 13 (G9 / RecHTLV) or 14 (E5), or A3_A9_C6, A11_E1; B11, B12, C2, D7, D11_F7_F9_G8, F6, F11, F12, G7, G10, and H2 (SEQ ID NOs: 26-37, 39), or any combination of said sequences. The above sequence combinations have the consensus sequence SEQ ID NO: 38 and are composed of at least a portion of the above-defined sequences (e.g., SEQ ID NO: 13) and another portion of the above sequences (e.g., SEQ ID NO: 14), for example, the N-terminal portion of one of the above sequences (e.g., SEQ ID NO: 13) and the C-terminal portion of another of the above sequences, or vice versa. The engineered recombinase can be a combination of three or more (e.g., four or five) portions from two or more (e.g., three, four, or five) of the above sequences.

[0027] The engineered recombinase may have high specificity for recombining the particular SEQ ID NO: 1. It may also recombine other sites, such as loxP.

[0028] However, as shown below, the recombinases of the present invention have the advantage that they have no detectable activity against human sequences that would allow for the removal or modification of human sequences from or within the genome. When applying the recombinases of the present invention to human therapeutics, there is an advantage in that the risk of cross-reaction and recombination with human sequences in host cells is minimal. This is one of the factors that contribute to the tolerability of the recombinase in human cells. However, the issue here is not only the short-term effects of recombinase expression, but also safety aspects such as the potential carcinogenic effects of non-specific recombination, even at low efficacy. Complete elimination of residual activity of the regulated recombinase contributes to the safety and reliability of the resulting recombinase in a therapeutic environment.

[0029] In the context of the present invention, a nucleic acid or protein comprising a sequence may consist of that sequence.

[0030] It may contain additional sequences, i.e., it may be a fusion protein. Such additional sequences may be, for example, a signal sequence that provides expression / localization in a specific cellular compartment, such as a nuclear localization signal as set forth in SEQ ID NO: 25, or a nucleic acid sequence encoding such a signal.

[0031] Proteins used in pharmaceutical compositions are preferably expressed as fusion proteins with a protein transduction domain, such as the tat protein transduction domain, which allows for protein transduction into target cells.

[0032] Preferably, the recombinase of the present invention is produced as a fusion protein with a nuclear localization sequence and a protein transduction domain such as tat, and a nucleic acid encoding the recombinase of the present invention can encode such a fusion protein. For example, the following protein transduction domains can be used in fusion proteins with the recombinase of the present invention, and preferably further contain a nuclear localization signal:

[0033] -The basic domain of the HIV-1 Tat transactivator (Fawell et al, 1994) (Fawell S, Seery J, Daikh Y, Moore C, Chen LL, Pepinsky B, Barsoum J., Intracellular delivery of heterologous proteins by Tat. Proc Natl Acad Sci U S A. 1994 Jan 18;91(2):664-8) - The homeodomain of Drosophila Antennapedia (Antp) (Derossi et al, 1994) (Derossi D, Joliot AH, Chassaing G, Prochiantz A., "The third helix of the Antennapedia homeodomain crosses biological membranes", J Biol Chem. 1994 Apr 8; 269(14): 10444-50) -HSV VP22 transcription factor (Elliott & O'Hare, 1997) (Elliott G, O'Hare P., Intercellular transport and protein delivery by herpesvirus structural proteins. Cell. 1997 Jan 24;88(2):223-33) -The cell-permeability transfer motif (TLM) of the PreS2 surface antigen of hepatitis B virus (HBV) (Oess & Hildt, 2000) (Oess S, Hildt E., A novel cell-permeability motif derived from the PreS2 domain of the hepatitis B virus surface antigen. Gene Ther. 2000 May;7(9):750-8)

[0034] If the protein needs to be purified, a tag that facilitates protein purification, such as a His tag, can also be added.

[0035] Fusion proteins can be prepared by methods well known in the art. For example, the expression vector into which the nucleic acid encoding the regulated recombinase is cloned may already contain a nucleotide sequence encoding a second polypeptide or protein. By cloning the nucleic acid encoding the regulated recombinase in frame with the sequence of the second polypeptide or protein, both sequences are expressed as a fusion protein.

[0036] The codon usage of the nucleic acid encoding the above-mentioned recombinase of the present invention can be selected by those skilled in the art.For example, when expression in human cells is intended, particularly for therapeutic purposes, the codon usage suitable for expression in human cells can be selected.The codon usage can also be based on the codon usage of, for example, Cre recombinase.

[0037] Recombinases or nucleic acids encoding the recombinases can be obtained or obtained by the methods described herein, and based on the sequences described herein, these sequences can be optionally combined or further varied to test their activity in recombining asymmetric target sites, such as sequences of SEQ ID NO: 1.

[0038] The present invention provides a nucleic acid encoding a regulated recombinase capable of recombining an asymmetric target sequence, such as SEQ ID NO: 1, within the LTR of the proviral DNA of the majority of HTLV-1 strains, wherein the regulated recombinase comprises the amino acid sequence defined above.

[0039] In the methods of the present invention, the nucleic acid encoding the regulated recombinase active against the asymmetric target sequence in the LTR of retroviral DNA is preferably cloned into an expression vector. Thus, the present invention also relates to an expression vector comprising a nucleic acid encoding the regulated recombinase defined herein.

[0040] An expression vector is a genetic construct for expressing a protein encoded by a nucleic acid within the vector. Such expression vectors may be either self-replicating extrachromosomal vectors or vectors that integrate into a host genome. Generally, these expression vectors contain transcriptional and translational regulatory nucleic acids operably linked to the nucleic acid encoding the regulated recombinase of the present invention.

[0041] As used herein, the term "nucleic acid" refers to a polymeric compound composed of covalently linked subunits called nucleotides. Nucleic acids include polyribonucleic acid (RNA), e.g., mRNA, and polydeoxyribonucleic acid (DNA), which may be either single-stranded or double-stranded. DNA includes cDNA, genomic DNA, synthetic DNA, and semi-synthetic DNA.

[0042] The term "regulatory sequence" refers to DNA sequences necessary for the expression of an operably linked coding sequence in a particular host organism. Regulatory sequences suitable for prokaryotes include a promoter, optionally an operator sequence, and a ribosome binding site. Eukaryotic cells are known to utilize promoters, polyadenylation signals, and enhancers.

[0043] A nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. For example, a promoter or enhancer is operably linked to a coding sequence if it affects the transcription of the coding sequence. Or, a ribosome binding site is operably linked to a coding sequence if it is positioned so as to promote translation. Linking is accomplished by ligation at convenient restriction sites. If such sites do not exist, synthetic oligonucleotide adapters or linkers are used according to conventional techniques. Nucleic acids involved in regulating transcription and translation are generally appropriate for the host cell used to express the regulated recombinase. Many types of appropriate expression vectors and appropriate regulatory sequences are known in the art for a variety of host cells.

[0044] The expression vector used in the present invention can be a retroviral vector, lentiviral vector, spumavirus vector, adenovirus vector, or adeno-associated virus vector. It can be a transposon transposable by Sleeping Beauty, such as SB100X or its mutants. Another transposon can be transposable by Frog Prince. However, in a preferred embodiment, the expression vector is selected from lentiviral vectors derived from HIV-1, SIV, FIV, or EIAV. Lentiviral vectors are described, for example, in Schambach et al. (2006) or European Patent Application No. 1 1000751.5.

[0045] In a preferred embodiment of the invention, the expression vector comprises a cellular, bacterial, viral or hybrid promoter.

[0046] Generally, for purposes of the present invention, a promoter can be a constitutive promoter or an inducible promoter. Furthermore, a promoter can be either a naturally occurring promoter, such as a bacterial, cellular, or viral promoter, or a hybrid promoter. Hybrid promoters combine elements of multiple promoters and are known in the art and are useful in the present invention. Furthermore, a promoter used in the present invention can be a derivative of a naturally occurring promoter. As used herein, the term "derivative" of a naturally occurring promoter refers to a combination of cis-acting elements obtained from promoters or sequences of different origins, or to a derivative obtained by deleting or mutating cis-acting elements within a particular naturally occurring promoter (e.g., as taught in Edelmann et al., 2000; Hartenbach & Fussenegger, 2006; Alper et al., 2005).

[0047] In an embodiment of the invention, the constitutive promoter or derivative thereof is selected from or derived from the promoter of cytomegalovirus, Ruth's sarcoma virus, murine leukemia virus-associated retrovirus, phosphoglycerol kinase gene, murine spleen focus-forming virus, or human elongation factor 1 alpha.

[0048] In a further embodiment of the invention, the inducible promoter or derivative thereof is selected from or derived from LTRs or derivatives thereof derived from lentiviruses, spumaviruses and deltaretroviruses.

[0049] As used herein, the term "LTR" refers to the 5' and 3' long terminal repeats of the provirus, which have promoter function (for a review, see the standard text "Retroviruses" (Coffin et al., 1997)).

[0050] Preferably, the inducible promoter or derivative thereof is selected from or derived from LTRs or derivatives thereof from HIV-1, HIV-2, MVV, EIAV, CAEV, SIV, FIV, BIV, HTLV-I and HTLV-II.

[0051] The present invention further provides methods for preparing a regulated recombinase, comprising preparing an expression vector encoding the regulated recombinase of the present invention, and further comprising expressing the regulated recombinase (or a fusion polypeptide comprising the amino acid sequence of the regulated recombinase) from the expression vector in a suitable host cell.

[0052] Preferably, the final recombinase is tested in mammalian cells to ensure it functions in a mammalian cellular environment. Furthermore, to obtain good expression in mammalian cells, the recombinase can be optimized for expression in these cells (e.g., by optimizing codon usage using methods known in the art). See, for example, (Shimshek et al., 2002). Additionally, a signal sequence, such as an NLS sequence (Macara, 2001), necessary for directing the protein to the nucleus of a mammalian cell can be added to the engineered recombinase nucleic acid. Expression of the engineered recombinase-encoding nucleic acid can be achieved, for example, using bacterial, insect, or mammalian expression systems. However, other expression systems known in the art can also be used. Methods for introducing exogenous nucleic acids into mammalian, insect, or bacterial hosts, as well as other hosts, are widely known in the art and may vary depending on the host cell used. Techniques include dextran-mediated transfection, calcium phosphate precipitation, polybrane-mediated transfection, protoplast fusion, electroporation, viral infection, encapsulation of polynucleotides in liposomes, and direct microinjection of DNA into the nucleus.

[0053] Thus, the present invention provides regulated recombinases, i.e., proteins, encoded by the nucleic acids of the present invention. As used herein, the term "protein" includes proteins, polypeptides, and peptides. The nucleic acid sequences of the present invention can be translated into protein sequences, as will be understood by those skilled in the art. The recombinase may optionally be a fusion protein. In one embodiment, the regulated recombinase protein is prepared as a fusion polypeptide to increase expression. In another embodiment, the regulated recombinase protein is produced as a fusion polypeptide so that the polypeptide can be introduced into living cells. Generally, purified proteins cannot cross cell membranes and enter cells due to their size. However, fusing specific peptide sequences to proteins can result in these fusion proteins being internalized within cells, where the protein can perform its function. Site-specific recombinases, particularly Cre recombinase, have been successfully introduced into cells using this approach (Peitz et al., 2002). Cell-permeable recombinases are further described by (Nolden et al., 2006) and (Lin et al., 2004). This strategy can be used to deliver a regulated recombinase to cells to remove the provirus from infected cells. Therefore, the second polypeptide of the fusion polypeptide can contain a signal peptide. The signal peptide can be a protein transduction domain such as the TAT peptide, or a peptide from the third helix of the Antennapedia homeodomain (Derossi et al., 1996, 1994; Viveus et al., 1997, 2003; Richard et al., 2005), or an NLS (nuclear localization sequence) for transporting the fusion polypeptide to the nucleus of eukaryotic cells (Macara, 2001). For example, the fusion protein of the present invention can contain a nuclear localization signal (NLS) such as SEQ ID NO: 25, tat, and the recombinase sequence of the present invention, RecHTLV or E5.

[0054] In yet another embodiment, the present invention provides a transformed cell containing a nucleic acid of the present invention. Transformed cells used to express the recombinase prepared from the expression vector include prokaryotic cells, such as bacterial or yeast cells, or eukaryotic cells, preferably insect or mammalian cells, most preferably human cells. The host cell may be a hematopoietic cell, such as an adult hematopoietic stem cell or a T cell, e.g., a CD4+ cell. The cell may be obtained from a subject infected with a retrovirus, and the cell may be returned to the subject after transformation and, if desired, cultured and / or expanded.

[0055] The transformed cell is preferably a human cell, for example, a human stem cell, particularly in the context of treating a human subject. The cell is preferably a hematopoietic stem cell, for example, a CD34+ stem cell. The stem cell may be an adult stem cell, an embryonic stem cell, or an induced pluripotent stem cell. In the state of the art, the term "stem cell" refers to a cell that (a) has the ability to self-renew, and (b) has the ability to form at least one, and often multiple, specific cell types by asymmetric division (Donovan & Gearhart, 2001).

[0056] Stem cells are preferably infected or transfected with the expression vector of the present invention. In a preferred embodiment, the stem cells are hematopoietic stem cells that express the regulated recombinase, the above-mentioned fusion polypeptide, or the above-mentioned expression vector. Hematopoietic stem cells (HSCs) are CD34+ cells derived from bone marrow and can be purified by conventional leukapheresis from, for example, peripheral blood of donors mobilized with G-CSF (e.g., HIV-infected patients) (Scherr & Eder, 2002). Cells genetically modified in vitro can then be formulated for reinfusion into patients.

[0057] In yet another embodiment, the expression vectors of the invention are used to transform T cells, such as primary CD4+ cells (blood cells) from patients infected with HTLV-1. Cell lines such as Jurkat T cells can also be used.

[0058] Alternatively, the regulated recombinase of the present invention can be formulated for delivery by virus-like particles (VLPs). VLPs can be used to package recombinase mRNA, recombinase protein (e.g., fusion proteins), or DNA, such as a DNA plasmid expressing the recombinase, or a construct containing a promoter-recombinase cDNA-polyA site. Thus, the nucleic acid of the present invention can further include a packaging signal.

[0059] The present inventors have for the first time identified the target sequence of SEQ ID NO: 1 in HTLV-1, which can be advantageously used to eliminate HTLV-1 proviruses containing this sequence, which applies to the majority of strains.

[0060] The present invention provides pharmaceutical compositions comprising the nucleic acids of the present invention, the regulated recombinases of the present invention, and / or the transformed cells of the present invention. The pharmaceutical compositions may optionally contain pharmaceutically acceptable excipients, additives, and / or solvents. Suitable solvents or carriers, excipients, and additives for use in pharmaceutical compositions are known in the art. The pharmaceutical compositions are preferably in the form of a solution suitable for intravenous administration (infusion).

[0061] According to the present invention, the nucleic acid of the present invention, i.e., the expression vector, recombinase protein, fusion protein, or cells (e.g., stem cells) obtained by the method of the present invention, comprising the nucleic acid encoding the regulated recombinase of the present invention, is formulated as a pharmaceutical composition. The composition can be used to prevent and / or treat HTLV-1 infection and reduce the viral load in subjects infected with HTLV-1. A further object of the present invention is a pharmaceutical composition obtained by the method of the present invention described herein.

[0062] The pharmaceutical compositions of the invention, which comprise an expression vector encoding a regulated recombinase (or the regulated recombinase as a protein or fusion polypeptide, or as stem cells containing the expression vector), can reduce the viral load in subjects infected with HTLV-1, eradicating the viral genetic reservoir within the host cell and thereby preventing further progression of the viral life cycle.

[0063] The pharmaceutical composition administered to a subject infected with the HTLV-1 retrovirus is selected from the group consisting of humans, primates, monkeys, cows, horses, goats, sheep and domestic cats, although the subject is preferably a human.

[0064] Generally, an effective amount of the expression vector, regulated recombinase, or transformed cell of the present invention should be administered to a subject. Administration may be, for example, intravenous or intramuscular.

[0065] For example, a pharmaceutical composition can include a transformed cell comprising a nucleic acid of the invention, which nucleic acid is capable of expressing a regulated recombinase of the invention. The transformed cell can be, for example, a human hematopoietic stem cell.

[0066] In some embodiments, the pharmaceutical composition is prepared to be administered simultaneously with an antiviral agent, or is formulated to be administered following systemic immune activation therapy or specific activation of proviral gene expression. Systemic immune activation (activation of immune cells, including resting cells) is usually achieved by administering immunotoxins, cytokines (e.g., IL-2), or T cell activating antibodies (e.g., OKT3). Systemic immune activation therapy or specific activation of proviral gene expression, or similar therapeutic strategies, greatly benefit from the simultaneous removal of proviral DNA, thereby reducing the pool of infected cells in patients.

[0067] The pharmaceutical compositions of the present invention are generally preferably used for treating or preventing HTLV-1 infection in a subject, and may also be used, if desired, for the treatment of adult T-cell leukemia (ATL) and / or HTLV-1-associated myelopathy (HAM). Prevention is understood to mean reducing the likelihood of infection or reinfection.

[0068] Optionally, the pharmaceutical composition is formulated for administration to a subject when the proviral DNA contained in a sample obtained from the subject contains the asymmetric target sequence of SEQ ID NO:1.

[0069] The present invention provides a method for treating and / or preventing HTLV-1 infection, comprising administering an effective amount of the pharmaceutical composition of the present invention to a subject. In one embodiment, sequences are analyzed in a sample obtained from a subject infected with HTLV-1, and the subject is administered an expression vector encoding at least one regulated recombinase, at least one regulated recombinase, or at least one cell (e.g., stem cell) transformed with the expression vector. This is the case when the subject's proviral DNA contains the asymmetric target sequence of SEQ ID NO: 1, from which the recombinase is selected. The sample obtained from the subject may be a blood sample containing infected CD4+ cells.

[0070] Prevention of a disease is understood to mean the reduction of the onset of the disease.

[0071] Also disclosed is a method for producing a nucleic acid or expression vector encoding a tailored recombinase capable of recombining asymmetric target sequences within the long terminal repeat (LTR) of proviral DNA of multiple HTLV-1 strains, comprising the steps of providing a target sequence present in the LTR of proviral DNA of multiple HTLV-1 strains, the target sequence having at least 30% homology (i.e., sequence identity) to the left half-site sequence and right half-site sequence in the LTR sequence of at least one known recombinase target site, where the homologous sequences are separated by a 5-12 nucleotide spacer, and the target sequence is preferably SEQ ID NO: 1, through the repetition of the following steps until at least one recombinase has activity on the asymmetric target sequence within the LTR of HTLV-1 DNA: i) directed evolution of molecules using as substrates, for at least one recombinase that recognizes a known homologous target site, a target sequence based on the sequence of an asymmetric target sequence but modified to include limited mutations from the known target sequence; wherein in each round, the target sequence can be mutated so that it differs by one, two, or three nucleotides from the target sequence on which the recombinase is known to act; and ii) shuffling the recombinase library to obtain a recombinase library capable of recombining target sequences with higher homology to the asymmetric target sequence; repeated steps of molecular directed evolution and library shuffling to negatively select for recombination at known target sites; expressing the library in human cells, particularly HeLa cells or human T cells, culturing the human cells expressing the selected recombinases for at least one week, preferably at least two weeks, and selecting the modified recombinase(s) by isolating the nucleic acid of the recombinase from the cultured cells expressing the selectable marker; and Optionally, cloning the nucleic acid encoding the recombinase into a suitable expression vector. A method of manufacture is also described, including:

[0072] The present invention uses SEQ ID NO: 1 to provide a target sequence in the HTLV-1 proviral LTR suitable for proviral elimination in a high percentage of HTLV-1 strains. Recombinases are tailored to recognize asymmetric target sites that differ from the original symmetric target site, which may be present in multiple retroviral strains, by dividing the substrate into new subsets with minor differences from the original target and gradually tailoring the recombinase to recognize these subsets (see, e.g., WO 2008 / 083931 and WO 2011 / 147590). A combinatorial approach allows for the selection of functional molecules that recognize asymmetric target sites within a given sequence. Thus, by passing through substrate intermediates during molecular-directed evolution, it is possible to generate enzymes with novel, distant asymmetric target specificities. This approach is also employed in the present invention. The present invention targets HTLV-1, specifically SEQ ID NO: 1, using the methods described in WO 2008 / 083931 and WO 2011 / 147590.

[0073] By "at least one recombinase," we mean that the methods of the invention can direct one or more (single) recombinases, each having independent activity, to recombine asymmetric target sequences. It is not intended to encompass multiple different recombinases capable of recombining asymmetric target sequences. Indeed, the methods used in the invention do not result in the selection of a recombinase that must be combined with other, different recombinases to recombine asymmetric target sequences, because only one recombinase is expressed in each cell.

[0074] The regulated recombinase obtained by the method of the present invention has the ability to recombine asymmetric target sequences within the LTR of HTLV-1 proviral DNA. The proviral DNA targeted by the recombinase can be inserted into the genome of a host cell. The regulated recombinase of the present invention can recombine asymmetric target sequences within the LTR of HTLV-1 that have not (yet) been inserted into the genome of a host cell, i.e., sequences that exist as non-inserted preintegration complexes (PICs). Therefore, both HTLV-1 that has not yet been inserted into the genome of a host cell and HTLV-1 that has already been inserted can be inactivated by the regulated recombinase of the present invention.

[0075] It is noted that in this invention and related art, the terms "target sequence," "target site," and "recombination site" are used interchangeably.

[0076] As described herein, a suitable asymmetric target sequence for HTLV-1 has been determined, shown as SEQ ID NO: 1 (see Figure 1A, loxHTLV). This allows the generation of therapeutic recombinases for retroviral genomes. These recombinases target recognition sites present in as many retroviral strains as possible.

[0077] SEQ ID NO: 1 is present in 86% of the HTLV-1 strains included in the HTLV-1 molecular epidemiology database.

[0078] The present invention also discloses a method for preparing a regulated recombinase, which method comprises inserting a nucleic acid encoding the recombinase into an expression vector and expressing it in a suitable host cell, which may be an adult stem cell, where the recombinase may be expressed as a fusion polypeptide comprising the amino acid sequence of the regulated recombinase.

[0079] The present invention also discloses a method for preparing a transformed cell, which method comprises preparing an expression vector comprising a nucleic acid sequence encoding a recombinase of the present invention and the further step of introducing the expression vector into a cell in vitro, which optionally is an adult stem cell. [Brief explanation of the drawings]

[0080] [Figure 1] Figure 1. A) Alignment of the loxHTLV (RecHTLV) target sequence (SEQ ID NO: 1) with loxP(Cre), loxLTR(Tre), and loxBTR(Brec1) (SEQ ID NOs: 9, 21, and 22, respectively). Mismatches with the loxP sequence are indicated in lowercase, and asymmetries within the sequence are underlined. Spacers are italicized. B) Scheme of the RecHTLV evolution process. Recombinase library activity from the first and last cycles at each subsite is shown as an agarose gel image. The upper band indicates no recombination (two triangles), while the lower band indicates recombination (one triangle). The number of evolutionary cycles performed at each subsite is indicated between the two agarose gels. The lower triangle indicates the evolutionary decrease in L-arabinose (L-ara) levels at each target site. C) Alignment of the loxHTLV (RecHTLV) target sequence (SEQ ID NO: 1) to evolutionary subsites (1, 1A, 1B, 2, 2A, 2B, 2C) (SEQ ID NOs: 2-8). Discrepancies compared to loxHTLV are indicated in lowercase and spacers are italicized. [Figure 2]Figure 2A) Overview of the clone selection process. The top panel shows the configuration of the construct used to express the library: expression of the recombinase fused to a nuclear localization signal (NLS-Rec) is regulated by a tetracycline-inducible promoter (TRE3G) bound by Tet-On 3G (Tet-On Advanced transactivator protein) upon doxycycline treatment; dual expression of the Tet-On 3G and GFP cassette is driven by the elongation factor 1α (EF1a) promoter and separated by a self-cleaving 2A peptide (P2A). The bottom panel shows the procedure used to select candidate clones: the recombinase library was cloned into a lentiviral vector. After lentivirus production, HeLa cells were infected with recombinase-containing lentiviral particles. After 7 days of expression, the recombinase was extracted by PCR and cloned into the pEVO bacterial plasmid. This allowed for sequencing and activity evaluation against loxHTLV. The clones were then tested in cell culture to evaluate activity against the loxHTLV target site. B) Top: Agarose gel for plasmid activity assay showing the recombination efficiency of Cre, Brec1, and RecHTLV against their respective target sites in bacteria. Test digests were performed, yielding a small fragment for the recombined plasmid (one triangle) and a large fragment for the unrecombined plasmid (two triangles). Recombinases were tested at 0 μg / mL, 1 μg / mL, 10 μg / mL, 50 μg / mL, 100 μg / mL, and 200 μg / mL of L-arabinose, as described above. Quantification of recombination is shown below the gel as a fraction of 1 (1 indicates fully recombined, 0 indicates no recombination). Bottom: Line graph showing quantification of recombination at different levels of arabinose for Cre, Brec1, and RecHTLV. Part of the figure was created using BioRender.com. C) Quantification of recombination activity based on flow cytometry analysis and quantification of the frequency of mCherry-positive cells relative to the total number of transfected (GFP+) cells. RecHTLV clone G9 has been shown to exhibit the highest recombination activity.D) Graph showing the frequency of mutant sequences in the library compared to Cre. Each bar represents an amino acid position. The indicated mutations, V7L, P12S, P15L, V23A, S38P, Y77H, A84T, K86N, Q90K, G93A, Q94E, M97T, S108G, K122R, I166V, A175S, K244R, N245Y, A248V, A249V, R259H, L261M, E262Q, E266G, T268A, N317T, and I320S, are present in more than 75% of the libraries and are also present in RecHTLV. E) Alignment of different recombinases Cre, Tre, BREC1, libraries A3_A9_C6, A11_E1; B11, B12, C2, D7, D11_F7_F9_G8, E5_E9 (also called E5), F6, F11, F12, G7, G9 (RecHTLV), G10 and H2 (SEQ ID NOs: 11, 15, 16, 24, 14, 13 and 26–37, 39). [Figure 3]Figure 3. A) Schematic diagram of the reporter and expression constructs used to assess RecHTLV activity in HeLa cells. The reporter construct, pCAGGS-lox-pA-lox-mCherry, consists of three polyA (pA) sequences flanking two loxHTLV sites, preventing mCherry expression driven by the CAG promoter. mCherry expression is only possible upon recombination of the loxHTLV sites. An IRES (internal ribosome entry site) bicistronic expression construct allows simultaneous expression of a recombinase fused to a nuclear localization signal (NLS-Rec) and a GFP cassette driven by a CMV promoter. B) Representative images showing a fluorescence-based recombination reporter assay. HeLa cells were cotransfected with the indicated recombinase and reporter constructs. RecHTLV expression was confirmed only when mCherry expression was observed. Scale bar represents 200 µm. First, cells were gated to remove doublets, and then single cells were gated based on transfection efficiency (GFP+). mCherry+ cells were quantified from the transfected population. C) Left: Representative dot plots obtained by flow cytometry of HeLa cells transfected with a loxHTLV reporter construct and a Cre or RecHTLV expression construct. On the right, quantification of recombination efficiency, measured as the percentage of mCherry+ cells and the frequency of transfected (GFP+) cells, is shown. Bars represent the mean ± standard deviation of results from five independent experiments (***p<0.001, unpaired two-tailed t-test). [Figure 4]Figure 4. A) Schematic of the reporter cell line and expression construct used to assess RecHTLV activity in HeLa cells in its genomic context. The reporter construct consists of a puromycin cassette flanked by two loxHTLV sites, which prevent expression of mCherry driven by the SFFV promoter. Expression of mCherry is only possible if recombination of the loxHTLV sites occurs. In the dual-cistronic recombinase expression construct, an IRES (internal ribosome entry site) allows simultaneous expression of the recombinase fused to a nuclear localization signal (NLS-Rec) and a GFP cassette driven by a CMV promoter. B) Left: Representative dot plots obtained by flow cytometry of loxHTLV reporter cells transfected with Cre or RecHTLV expression constructs. On the right, quantification of recombination efficiency, measured as the ratio of mCherry+ cells to all transfected cells (GFP+), is shown. Bars represent the mean ± standard deviation of three independent experiments (****p<0.0001, ns: not significant, unpaired two-tailed t-test). C) PCR performed on genomic DNA from transfected loxHTLV HeLa reporter cells. Left: Schematic of PCR with primer F and primer R located in the outer flanking regions of loxHTLV. The expected amplification product is 1.2 kb for the unmodified reporter and 560 bp for the modified reporter. Right: Agarose gel of PCR products from genomic DNA extracted from different samples: Lane 1: HeLa cells, Lanes 2-4: reporter HeLa cells transfected with the indicated constructs, Lane 5: water control. The upper band on the gel represents the unmodified reporter (two triangles), and the lower band represents the modified reporter (one triangle). [Figure 5]Figure 5. A) Top: Alignment of the loxHTLV target sequence to in silico predicted off-target sites 1, 2, and 3 (SEQ ID NOs: 17-20). Off-target site 3 is asymmetric and therefore indicated as 3.1 and 3.2 in the alignment. Mismatches with the loxHTLV sequence are indicated in red, and the spacer is gray. Bottom: Agarose gel showing the recombination activity of RecHTLV against the on-target sequence (loxHTLV) and three predicted off-targets in bacteria under 50 μg / mL L-arabinose conditions. Off-target sites 1 and 2 were inserted twice as excision substrates in the pEVO vector, while off-target site 3 consists of sequences 3.1 and 3.2 inserted as excision substrates. Quantification of recombination efficiency is shown below each lane. The unrecombined band at the top is indicated by two triangles, and the recombined band is indicated by a single triangle. B) Schematic diagram of the plasmid used to validate peaks selected from ChIP-seq experiments in bacteria. From the selected peak, a 92-bp fragment around the peak's apex was cloned twice into the pEVO vector as an excision substrate (shown as peak 7, chr12). C) Bacterial plasmid-based RecHTLV activity assay of 10 peaks selected from ChIP-seq binding sites. All peaks were tested with 200 μg / mL L-arabinose. Quantitative recombination is shown below each lane. D) Peak 7 was analyzed to search for RecHTLV recombinase recognition sequences. The 92-bp sequence (SEQ ID NO: 10) tested in the plasmid assay was divided into the following subgroups of interest: "left" (positions 1-46), "center" (24-70), "right" (47-92), and "lox" (40-73P). An agarose gel from a bacterial test digestion assay shows the recombination efficiency of different portions of peak 7. Only the lox-peak7 sequence (SEQ ID NO: 23) exhibits recombination under 200 μg / mL L-arabinose. Recombination quantification results are shown below each lane. The upper unrecombined band is indicated by two triangles, and the recombinant band is indicated by a single triangle. E) Alignment of the loxHTLV sequence to peak7-lox. Mismatches with the loxHTLV sequence are shown in red, and the spacer is shown in gray. [Figure 6] Figure 6. A) Construction of the recombinase expression vector and HTLV-1 provirus expression plasmids pCS-HTLV-X1MT and pCMV-HT1M. pCS-HTLV-X1MT contains two full-length LTRs, whereas pCMV-HT1M contains only one truncated 5' LTR and therefore lacks one loxHTLV target site. In the dual-cistronic recombinase expression construct, the self-cleaving 2A peptide (P2A) enables coexpression of the recombinase fused to the 3xFLAG peptide and nuclear localization signal (NLS-Rec) with a GFP cassette driven by the elongation factor 1α promoter (EF1a). B) Detection of Gag p55 / p19 proteins and FLAG-tagged recombinase 72 hours after transient transfection of 293T cells with proviral expression plasmids (pCS-HTLV-X1MT, pCMV-HT1M) along with the recombinase expression constructs Puro (control), Cre, and RecHTLV. Hsp90 was used as a control. Numbers indicate densitometric values ​​of Gag (p55 + p19) proteins normalized to Hsp90. Red arrows indicate the expected molecular weights of p55 and p19 products. The blot has been truncated for technical reasons. C) Densitometric analysis of Gag proteins (normalized to Hsp90). Bars indicate the mean ± standard deviation of three to four independent experiments. (*p<0.05, ns: not significant, unpaired two-tailed t-test using the logarithm of the normalized value). [Figure 7]Figure 7. A) Schematic of the experimental procedure: Chronically HTLV-1-infected C91-PL donor cells were cocultured for 5 days with Jurkat recipient T cells pre-stained with CMAC-Blue, which constitutively express the RecHTLV recombinase. Cre was used as a control. Infected Jurkat T cells were detected by intracellular FACS staining of the Tax protein. B) Representative dot plots obtained by flow cytometry show the gating strategy for the coculture of Jurkat recipient cells (CMAC+) and C91-PL donor cells (CMAC-). The two populations were distinguished based on CMAC staining, and Tax expression was further detected in the different cell populations. C) Quantification of the frequency of Tax+ cells in Jurkat (CMAC+) recipient cells. Left: integrase inhibitor raltegravir compared to the solvent control DMSO; Right: Jurkat cells expressing Cre or RecHTLV. For statistical analysis, the percentage of C91-PL Tax+ cells in the coculture was considered. Error bars indicate 95% confidence intervals (n=8, **p<0.01; ***p<0.001, t-test following linear mixed model). Portions of the figure were generated using BioRender.com. [Figure 8]Figure 8. A) Schematic diagram of the lentiviral expression constructs SIN-LV-RecHTLV and SIN-LV-GFP used. The EF1a promoter drives expression of the HTLV-1-specific recombinase RecHTLV fused to a nuclear localization signal (NLS-Rec) or GFP, followed by a P2A peptide, enabling expression of a puromycin resistance gene. B) Experimental workflow: HTLV-1-infected SP cell lines were transduced with lentiviral vectors expressing RecHTLV or GFP. After 9–10 days, genomic DNA was extracted, PCR was performed, and Sanger sequencing was performed to confirm correct excision of the HTLV-1 provirus. C) Structure of the HTLV-1 provirus integrated into the genome of SP cells. The black arrow indicates the position of the primer binding to the host cell genomic DNA. Recombination by RecHTLV at the loxHTLV site results in excision of the proviral genome, leaving a "genomic scar" consisting of one LTR sequence. D) PCR was performed on genomic DNA of recombinase-treated SP cells using the primers shown in panel B for proviruses on chromosomes 6 and 20. Red arrows indicate PCR products indicating excision of proviral DNA. MOI indicates the multiplicity of infection of the recombinase or GFP construct, and d indicates the number of days since transduction. Portions of the figure were generated using BioRender.com. [Example]

[0081] Example result Target site selection Target sites for Cre-like recombinases typically consist of an 8-base-pair spacer sequence flanked by a 13-base-pair inverted repeat and may contain asymmetric positions (Sarkar et al., 2007; Karpinski et al., 2016; Meinke et al., 2016). To identify suitable recombinase target sites, we applied the SeLOX algorithm (Surendranath et al., 2010) to a collection of 949 LTR sequences registered in the HTLV-1 Molecular Epidemiology Database (Araujo et al., 2012). From this data, we identified the most conserved 34-mers and restricted the results to sequences present in at least 80% of all registered LTRs. After removing sequences containing single-nucleotide repeats of five or more identical bases, we obtained a set of eight candidate target sites. To avoid unwanted recombination at off-target sites in humans, we examined all occurrences of sequences similar to the candidate target sites (reference assembly GRCh38, December 2013). A maximum of two mismatches per half-site were allowed, and mismatches in the 8-base-pair spacer region were ignored. Consequently, five candidates with potential off-target risk were eliminated. To select the optimal target site from the remaining three, we compared the sequences of these half-sites with target sites from previously evolved recombinase libraries (Buchholz & Stewart, 2001; Sarkar et al., 2007; Karpinski et al., 2016; Lansing et al., 2020, 2022). We selected the candidate with the fewest mismatches among existing target sites and named it loxHTLV (Figure 1A). The final target site was found at nucleotide positions 317 to 350 and 8,597 to 8,630 (GenBank accession number AB513134.1). loxHTLV is highly conserved and was detected without mismatches in 86% of sequences in the HTLV-1 molecular epidemiology database. Because loxHTLV is located in the HTLV-1 long terminal repeat (LTR) sequences and flanks the retroviral genome, recombination at these sites results in the excision of the entire proviral coding sequence from infected cells.

[0082] Directed evolution of HTLV-1-targeted recombinase To develop designer recombinases capable of recombining loxHTLV sites, we performed substrate-linked directed evolution (SLiDE) (Materials and Methods) (Buchholz & Stewart, 2001; Karpinski et al., 2016; Lansing et al., 2020, 2022). SLiDE is performed in Escherichia coli (E. coli) by linking the recombinase excision activity to a plasmid encoding it. In each evolution cycle, recombinases that successfully excised the target site were selected, and mutations were introduced into the library by PCR amplification. The amplified recombinases were then reinserted into the evolution plasmid (pEVO) to allow for successive adaptations.

[0083] To obtain a library that could efficiently recombine loxHTLV sequences at low recombinase expression levels, a total of 181 evolutionary cycles were performed using seven subsites, indicating that clones with high activity were selected during the SLiDE process (Figures 1B and 1C).

[0084] Selection of RecHTLV recombinase clones To enrich for recombinases that do not have deleterious effects when expressed in mammalian cells, the final library was inserted into a lentiviral vector that allows recombinase expression from a tetracycline-inducible promoter (Figure 2A). After infecting HeLa cells at a low infection rate to ensure that most cells expressed a single recombinase clone, doxycycline was added to the medium and the library was expressed for 7 days to confirm that recombinase expression was well tolerated in the cells (Figure 2A) (Lansing et al., 2022). The coding sequences of tolerated recombinases were PCR amplified from genomic DNA and inserted into a pEVO vector for evaluation of recombinase activity in E. coli. Recombinase expression was induced by the addition of 10 μg / mL L-arabinose, and 96 clones were tested for excision activity using a PCR assay. Of the 96 clones, 44 showed recombinase activity, as indicated by the appearance of a specific recombinase band. Sanger sequencing analysis of the clones revealed that several clones shared the same sequence (data not shown). Therefore, 29 unique active clones were selected for further evaluation in mammalian cell culture. The selected recombinases were inserted into mammalian expression vectors and separately cotransfected into HeLa cells with a fluorescent reporter plasmid that enables expression of mCherry upon recombination of the loxHTLV target site within the hoxLTR. Activity against the loxHTLV site in mammalian cells was evaluated (Figure 2C). Clone G9 (hereafter referred to as RecHTLV) showed the highest activity in this assay and was therefore selected for further evaluation as a candidate HTLV-1 recombinase (Figure 2C). Other clones A3_A9_C6, A11_E1; B11, B12, C2, D7, D11_F7_F9_G8, E5_E9 (also referred to as E5), F6, F11, F12, G7, G10, and H2 also showed high activity. Clones A11_E1, B11, B12, C2, D7, D11_F7_F9_G8, F6, F11, F12, and G9 showed the highest activity (mCherry expression above 60%). F6 and especially G9 resulted in mCherry expression above 70%.

[0085] To confirm the success of the initial screening for tolerance in HeLa cells, we inserted the RecHTLV coding sequence into the tetracycline-inducible lentiviral system described above. In this system, recombinase expression is driven by a tet-inducible promoter, and we generated HeLa-derived cell lines (Figure 2A). These vectors also express GFP, allowing us to track transformed cells over time. Continuous addition of 100 ng / mL doxycycline to the culture medium resulted in a slight decrease in the percentage of GFP+ cells in the RecHTLV-expressing cell lines, similar to the GFP-only control (data not shown). These results confirmed the success of the initial screening and demonstrated that RecHTLV does not affect cell growth in this system. Therefore, RecHTLV is a candidate for a highly tolerated HTLV-1 recombinase.

[0086] Mutations acquired during evolution To gain insight into the mutations acquired during evolution, the final evolved library was deep sequenced using PacBio's long-read sequencing technology. During the evolutionary process, we observed that the majority of clones in the library acquired dominant mutations at 35 residues compared to Cre (L5Q, V7L, N10S, P12S, P15L, V23A, K25Q, D29V, M30T, R34H, S38P, K57E, Y77H, A84T, K86N, Q90K, G93A, Q94E, M97T, S108G, K122R, I166V, A175S, K244R, N245Y, A248V, A249V, R259H, L261M, E262Q, E266G, T268A, P307A, N317T, and I320S).

[0087] As mentioned above, many of these mutations are in residues directly bound to or in close proximity to the target site DNA, indicating their putative function in modulating selectivity for new target sites (Karpinski et al., 2016; Sarkar et al., 2007; Buchholz & Stewart, 2001). Next, we sequenced the selected RecHTLV clone G9 using Sanger sequencing and found that it differs from Cre in 46 residues (N3K, V7L, N10D, P12S, P15L, V16A, A18V, V23A, S38P, K43E, K57E, L58S, Y77H, A84T, K86N, I88V, Q90K, G93A, Q94E, M97T, S108G, S11). We observed differences in the following: 4T, K122R, Q144R, Q156L, I166V, A175S, T206A, S226T, V227A, K244R, N245Y, A248V, A249V, P250S, Q255R, R259H, L261M, E262Q, E266G, A267T, T268A, M299L, P307A, N317R, I320S). During the evolution of E5, 42 mutations were acquired compared to cre: N3K, V7I, H8Y, N10D, P12S, P15L, V16A, V23A, S38P, K57E, L58S, Y77H, A84T, K86N, I88V, Q90K, G93A, Q94E, M97T, S108G, S110N, K122 R, S147A, C155R, I166V, A175S, K183R, K244R, N245Y, A248V, A249V, Q255R, R259H, L261M, E262Q, E266G, T268A, P307A, N317R, N319D, I320S, and D343Y (see alignment in Figure 2F). These are the most common mutations observed to date in Cre-derived recombinases and may be due to the significant differences in loxHTLV compared to loxP.

[0088] The mutations shown in Figure 2D, i.e., V7L, P12S, P15L, V23A, S38P, Y77H, A84T, K86N, Q90K, G93A, Q94E, M97T, S108G, K122R, I166V, A175S, K244R, N245Y, A248V, A249V, R259H, L261M, E262Q, E266G, T268A, N317T, and I320S, were present in more than 75% of the sequenced clones and, with the exception of N317T, were also present in RecHTLV when RecHTLV had N317R. The following mutations are present in both E5 and G9 / RecHTLV and are found in over 75% of the clones sequenced: P12S, P15L, V23A, S38P, Y77H, A84T, K86N, Q90K, G93A, Q94E, M97T, S108G, K122R, I166V, A175S, K244R, N245Y, A248V, A249V, R259H, L261M, E262Q, E266G, T268A, N317R or N317T, and I320S. These mutations are therefore considered to be the most important.

[0089] RecHTLV / G9 and E5 contain most of the mutations that emerged through library evolution (see alignments in Figure 2D and Figure 2F), including mutations previously identified in other Cre-based designer recombinases (Y77H, S108G, I166V, A175S, and I320S), suggesting that these residues may play a general role in protein stability (Abi-Ghanem et al., 2015; Karpinski et al., 2016). We also observed that both E5 and G9 / RecHTLV exhibit several mutations in the first 20 amino acids (positions 3, 7, 10, 12, 15, 16, and 18). The role of the N-terminus of Cre is controversial, as wild-type Cre can efficiently recombine loxP even when the first 20 amino acids are removed. However, recent reports have shown that these amino acid mutations contribute to protein stability in evolved Cre-type recombinases (Guilleun-Pingarroun et al., 2022; Rongrong et al., 2005; Warren et al., 2008). Therefore, amino acid mutations in the first 20 residues may also contribute to protein stability of RecHTLV or E5. In one embodiment, one or more, and up to all, of the first 20 amino acids at the N-terminus of a recombinase of the present invention (e.g., SEQ ID NO: 12, 13, or 14) are deleted.

[0090] One of the mutational hotspots for recombinases (e.g., RecHTLV) is located in helix D, particularly residues 84–94, which have been shown to be important for recognizing nucleotide positions near the loxP spacer in Cre. Interestingly, some of these mutations were previously identified in Brec1 recombinase (Karpinski et al., 2016). Nevertheless, most of the acquired mutations are unique to RecHTLV, which may confer specificity to nucleotide mutations in loxHTLV that differ from the loxBTR target site. Another mutational hotspot resides in helix J, which directly contacts the DNA major groove and has been shown to be involved in the recognition of nucleotide positions 8, 9, 10–25, 26, and 27 of the target site (Guo et al., 1997; Ruefer & Sauer, 2002; Kim et al., 2001). The loxHTLV target sequence differs significantly at these nucleotide positions when compared to loxP or other lox-like targets, and unique mutations are consistently observed in this protein region, including residues 261 and 267. These residues have never been mutated in other designer recombinases investigated to date (Lansing et al., 2022, 2020; Sarkar et al., 2007; Karpinski et al., 2016) (Figure 2F).

[0091] Collectively, these data provide new insights into how evolved recombinases recognize their target sites and may contribute to efforts to more rationally design Cre-based recombinases toward novel target sites (Schmitt et al., 2022; Soni et al., 2020; Abi-Ghanem et al., 2013).

[0092] RecHTLV specifically recombines with loxHTLV in Escherichia coli For more detailed characterization of RecHTLV, we tested its recombination activity at different expression levels in E. coli. The pEVO vector allows us to test the activity of recombinase clones in bacteria at different expression levels by adjusting the concentration of L-arabinose added to the culture medium. We loaded plasmid DNA isolated after digestion with diagnostic restriction enzymes onto an agarose gel at L-arabinose concentrations of 0 μg / mL, 1 μg / mL, 10 μg / mL, 50 μg / mL, 100 μg / mL, and 200 μg / mL. As expected, the intensity of the band specific to the recombinant form of the plasmid increased, indicating that the loxHTLV target site was recombined in an L-arabinose dose-dependent manner (Figure 2B). We compared the recombination activity of loxBTR by Brec1, a recombinase with potent antiretroviral activity in HIV-1-infected cells, with that of Cre at loxP (Karpinski et al., 2016) (Figure 2B, target sequence in Figure 1A). At low levels of expression (0 and 1 μg / mL L-arabinose), Cre efficiently recombined loxP, whereas Brec1 and RecHTLV showed less than 10% recombination activity at their respective target sites at this induction level. However, when expressed at high levels, RecHTLV efficiently recombined at loxHTLV sites (>80% recombination) and showed activity comparable to Brec1 at loxBTR. Overall, RecHTLV showed slightly lower activity at loxP compared to Cre, but similar recombination activity at loxBTR.

[0093] RecHTLV efficiently recombines with loxHTLV in HeLa cells To test the activity of RecHTLV in mammalian cells, we constructed a mammalian reporter vector containing three SV40 poly(A) cassettes flanked by two loxHTLV sites that inhibit expression of the downstream mCherry gene. Upon excision of the poly(A) cassette by recombination of the loxHTLV sites, mCherry expression was driven by the CAG promoter (Figure 3A). To express the recombinase (RecHTLV or Cre as the target), we used a plasmid containing an IRES sequence that drives expression from the cytomegalovirus (CMV) promoter and allows for the linked expression of GFP. This allowed us to control the transfection efficiency of the plasmids (Figure 3A). When both plasmids were co-transfected into HeLa cells, mCherry-positive cells were observed only when RecHTLV was co-transfected with the reporter; no mCherry-positive cells were detected in Cre-transfected cells (Figure 3B, C).

[0094] RecHTLV recombines with loxHTLV in the genomic environment To test whether RecHTLV can recombine loxHTLV sequences within a genomic context, we generated a HeLa reporter cell line containing two loxHTLV target sites flanking a puromycin cassette followed by an out-of-frame mCherry cassette. Upon recombination, the puromycin cassette is excised, allowing expression of mCherry driven by the spleen focus-forming virus (SFFV) promoter (Figure 4A). To express the recombinase, we transiently transfected the cells with a pIRES vector driving GFP expression, allowing us to track transfection efficiency. Forty-eight hours after transfection, we detected more than 40% of mCherry-positive cells transfected with RecHTLV (Figure 4B). This demonstrates that RecHTLV can recombine loxHTLV sequences within a genomic context. Importantly, expression of Cre recombinase in this reporter cell line did not result in a significant number of mCherry-positive cells (Figure 4B). To confirm correct excision of the puromycin cassette, we extracted genomic DNA from transfected cells. PCR assays using primers flanking the loxHTLV sites produced a band pattern consistent with correct excision via the loxHTLV sites (Figure 4C). Furthermore, sequencing of the PCR bands confirmed the precise nucleotide sequence predicted from recombinase-mediated excision. We conclude that RecHTLV is active in human cells and efficiently excises sequences flanked by two loxHTLV sites from the genome.

[0095] Experimental detection of potential off-target sites To investigate the specificity of RecHTLV, we bioinformatically screened the human genome based on sequence similarity to loxHTLV to identify potential off-target sites. The three most closely related sites, with half-site sequences containing five and six mismatches to the loxHTLV sequence, respectively (Figure 5A), were experimentally tested in bacteria. No detectable recombination was observed at these three putative off-target sites, even at high recombinase expression levels, indicating that RecHTLV is rather specific and does not recombine with the human target site with the closest sequence similarity to loxHTLV (Figure 5A).

[0096] To experimentally identify putative RecHTLV off-targets in human cells using an unbiased approach, we used chromatin immunoprecipitation followed by sequencing analysis (ChIP-seq) (Lansing et al., 2022). To this end, we fused RecHTLV to GFP and stably expressed the fusion protein in HeLa cells. We used established anti-GFP antibodies for immunoprecipitation of the recombinase to enrich for genomic sequences bound by RecHTLV (Ding et al., 2015, 2021; Lansing et al., 2022; Poser et al., 2008; Hein et al., 2015). After pulldown, DNA fragments were deep-sequenced using Illumina NGS sequencing. Based on the stacking of reads between control and test samples, we identified 112 putative binding sites within the genome and selected 10 peaks for experimental validation. To investigate whether the sequence bound by RecHTLV could recombine, we inserted a 92-base pair sequence found near the peak apex twice into a pEVO vector containing RecHTLV as an excision substrate (Figure 5B). After induction of recombinase expression, nine of the ten constructs showed no specific recombination band (Figure 5C). This confirmed the high specificity of RecHTLV and demonstrated that sequences bound by designer recombinases are not necessarily recombination substrates (Lansing et al., 2022). However, of the sites tested, peak 7 (Figure 5C) showed a weak signal of the band expected for recombination products. To identify the exact 34-base pair sequence recombined by RecHTLV, we divided the 92-base pair sequence (SEQ ID NO: 10) into three regions: left, center, and right. These sequences were then inserted twice into pEVO as excision substrates. We also tested the sequence with the highest similarity to loxHTLV (lox-peak7) in a similar manner (Figures 5D and E). We observed that only the lox-peak7 sequence displayed a recombination-specific band, thereby pinpointing the 34 base pairs recognized by RecHTLV in the genomic context (Fig. 5D).Despite being a potential off-target, the lox-peak7 sequence occurs only once in the human genome (Nurk et al., 2022). Therefore, deleterious recombination events involving this sequence are highly unlikely, and because the spacer sequence differs from that of loxHTLV, recombination between these two sites is unlikely. This result demonstrates the high specificity of RecHTLV and the suitability of ChIP-sequencing for experimentally identifying putative genomic off-target sites for engineered recombinases.

[0097] RecHTLV recombines with loxHTLV in the full-length HTLV-1 LTR sequence To test whether RecHTLV can recombine loxHTLV sequences into the full-length HTLV-1 LTR, we used the plasmid pHpX, which contains two full-length HTLV-1 LTR sequences flanking the Tax protein coding sequence (Nerenberg et al., 1987). For recombinase expression, a lentiviral plasmid was constructed in which the recombinase coding sequence was under the control of the elongation factor-1α (EF1α) promoter. Both plasmids were cotransfected into HEK293T or Jurkat T cell lines. In both cell lines, Tax protein expression was significantly reduced only when transfected with the RecHTLV-expressing plasmid, but not in the empty vector control (Puro, lacking the recombinase cassette) or Cre-transfected cells. This indicates excision of the Tax expression cassette from the plasmid by RecHTLV (data not shown). To clarify whether recombinase functions in full-length HTLV-1 proviral constructs, we performed transient assays using a plasmid (X1MT) containing the entire HTLV-1 genome (Derse et al., 1997). Cotransfection with a recombinase-expressing plasmid (Figure 6A) revealed that Gag expression (detected with Gag p19 antibody) was significantly reduced in the presence of RecHTLV compared with empty vector (Puro) or Cre controls (Figure 6B). As an additional control, we performed the same assay using the HTLV-1 packaging plasmid pCMV-HT1 M (Derse et al., 2001). In this case, viral protein expression is driven by the CMV promoter, and the 5' LTR is partially deleted, thereby lacking the loxHTLV site (Figure 6A). In contrast to the full-length proviral sequence, we observed that coexpression of RecHTLV with a 5' LTR-deleted provirus did not reduce Gag expression. This suggests that the reduction of Gag expression by RecHTLV functions via recombination of the loxHTLV target site (Fig. 6C ).We conclude that RecHTLV is capable of recombining loxHTLV in the full-length HTLV-1 genome.

[0098] Expression of RecHTLV reduces HTLV-1 infection in Jurkat T cells Next, we tested whether expression of RecHTLV reduces HTLV-1 infection in Jurkat T cells. To this end, we generated a Jurkat cell line that persistently expresses RecHTLV under the control of the EF1a promoter, as well as a control cell line expressing Cre. Because HTLV-1 infection depends on close cell-cell contact and completion rarely occurs from free viral particles (Fan et al., 1992; Pique & Jones, 2012; Alais et al., 2015), we cocultured a chronically infected C91-PL cell line capable of producing infectious viral particles with a Jurkat T cell line expressing the recombinase (Alais et al., 2017, 2015; Popovic et al., 1983). To distinguish Jurkat and C91-PL cells, we stained the Jurkat cell line with CMAC blue dye before coculture (Figure 7A and B) (Donhauser et al., 2018). As an indicator of productive infection, intracellular Tax protein expression was measured using flow cytometry 5 days after coculture. As a positive control, we applied raltegravir, a compound shown to inhibit HTLV-1 integrase activity and prevent de novo infection of cells (Seegulam & Ratner, 2011). As expected, raltegravir reduced Tax levels in Jurkat T cells compared with the DMSO solvent control (Figure 7C). This is likely due to the drug inhibiting HTLV-1 integration into the Jurkat cell genome. Importantly, a similar decrease in Tax expression was observed in Jurkat T cells expressing RecHTLV, but not in Cre-expressing cells (Figure 7B, C). These results indicate that expression of RecHTLV reduces stable HTLV-1 integration in Jurkat T cells.

[0099] RecHTLV expression excises integrated HTLV-1 provirus from infected cells Finally, we tested whether RecHTLV could excise HTLV-1 proviruses from cell lines isolated from ATL patients. We used a cell line derived from a chronically infected SP patient that had been shown to harbor four proviruses (two of which had full-length long terminal repeats) with known HTLV-1 integration sites (Meissner et al., 2017). First, we confirmed that SP cells harbored these integrated proviruses on chromosomes 6 and 20, both of which contained loxHTLV sequences within their respective long terminal repeats (data not shown). Next, we transduced the SP cell line with a lentivirus constitutively expressing RecHTLV recombinase (SIN LV RecHTLV) or a lentivirus expressing GFP (SIN LV GFP) as a negative control (Figure 8A). On days 9 or 10 postinfection, genomic DNA was extracted from infected cells and PCR reactions were performed using primers flanking the integration sites (Figures 8B and 8C). Consistent with proviral excision, genomic scars were detected at both loci, and PCR bands of the expected size were detected only in cells expressing RecHTLV recombinase (Figure 8D). Sequencing of the DNA fragments confirmed accurate and precise removal of the provirus from these loci (data not shown). Thus, RecHTLV has the ability to excise the HTLV-1 provirus from patient-derived cells.

[0100] Consideration To date, HTLV-1 remains an infectious disease with no cure. However, the proviral integration into the host genome and the stability of viral sequences within infected cells make HTLV-1 a suitable target for gene editing tools. In this study, we developed a recombinase-based approach to remove the HTLV-1 provirus from infected cells. RecHTLV recombinase (also known as G9) and the related clone E5 (another clone identified by this method) can target and efficiently excise sequences present within the HTLV-1 LTR in bacteria and human cells. Sequencing of genomic DNA isolated from cells treated with RecHTLV revealed that editing was successful and accurate. Notably, the amino acid sequence of RecHTLV contains 46 mutations compared to Cre, which is likely due to mismatches between the loxHTLV target sequence and loxP. This demonstrates the flexibility of designer recombinases, allowing targeting of sequences with only slight similarity to the original loxP sequence.

[0101] To investigate potential side effects of RecHTLV, we performed in silico and experimental off-target analyses. We failed to detect RecHTLV activity at predicted potential off-target sequences identified based on sequence similarity to loxHTLV target sites. Furthermore, using an experimentally unbiased approach, we demonstrated that ChIP-seq can be used to identify nonspecific binding sites within the genome, and this sensitivity allowed us to detect off-target sites with less than 10% recombination activity. Fortunately, this potential off-target sequence, lox-peak7, occurs only once within the genome, mitigating the threat of unwanted editing. Furthermore, recombination events between lox-peak7 and loxHTLV are highly unlikely because recombination requires identical spacer sequences, whereas the spacer sequences of lox-peak7 and loxHTLV are different (Meinke et al., 2016).

[0102] Previous attempts at gene editing in HTLV-1 have focused on disrupting viral protein expression, either by using zinc-finger nucleases to target the LTR and disrupt its promoter function, or by using CRISPR / Cas9 directly on the viral Hbz gene (Tanaka et al., 2013; Nakagawa et al., 2018). While the results were impressive, the editing ability of nucleases relies on the repair mechanisms of the treated cells. Therefore, the resulting sequences are unpredictable, and escape mutants are likely to emerge, as has been described in the literature when HIV-1 is targeted with nucleases (Wang et al., 2016a, 2016b; Yoder & Bundschuh, 2016; De Silva Feelixge et al., 2016). Tanaka et al. used a method similar to ours to target the LTR and eliminate the provirus in infected cells. Nevertheless, they observed that most of the edited cells contained indels within the LTRs, but not larger deletions (Tanaka et al., 2013). Therefore, the appearance of indels after treatment could be detrimental, as the edited virus may still be functional and therefore resistant to further treatment. In contrast, designer recombinases offer predictability with nucleotide precision, making them an excellent tool for excising proviruses from infected cells. Taken together, our data indicate that designer recombinases are promising tools for reversing HTLV-1 infection in human cells.

[0103] material and method Plasmid The evolved plasmid pEVO-loxHTLV and all PEVO-loxHTLV subsite plasmids, as well as PEVO-lox containing loxHTLV off-target sites, were generated using the Cold Fusion Cloning Kit (System Biosciences). Inserts were PCR amplified using pEVO as a template with primers containing the new lox sites as overhangs. The backbone was prepared by digesting the pEVO vector with BglII restriction enzyme (NEB). Cold fusion reactions were performed according to the manufacturer's protocol.

[0104] The pCAGGS-lox-pA-lox-mCherry reporter plasmid used in this study to study the activity of RecHTLV in HeLa cells was constructed from the plasmid described in Lansing et al. (2020 and 2022) and was originally derived from pCAG-loxPSTOPloxP-ZsGreen, a gift from Pawel Pelczar (Addgene plasmid #51269; http: / / n2 t.net / addgene:51269 ; RRID:Addgene_51269) (Hermannet et al., 2014). The target site was introduced by PCR using overhanging primers and inserted into the pCAGGS reporter using SalI and EcoRI restriction enzyme digestion cloning.

[0105] The expression vector pIRES-NLS-eGFP was generated from the pIRESneo-Cre vector described in Lansing et al. (2020). The recombinase was inserted into the pIRES expression plasmid by restriction enzyme digestion with XbaI and BsrGI.

[0106] The SVFF-loxHTLV-puro-mCherry reporter lentiviral plasmid (pLenti-loxHTLV-reporter) was generated from a bicistronic tagged BFP / eGFP reporter obtained from D. Sueruen (Sueruen et al., 2018) and inserted into the plentiCRISPR V2 lentiviral backbone, a gift from Feng Zhang (Addgene plasmid #52961; http: / / n2 t.net / addgene:52961 ;RRID:Addgene_52961) (Sanjana et al., 2014). Briefly, eGFP was replaced with an mCherry cassette using AgeI and XhoI (NEB). A puromycin cassette was inserted by Golden Gate cloning using BsmBI, flanked by loxHTLV sites, by PCR amplification with primers:

[0107] The pLentiX inducible system was constructed as described by Lansing et al. (2022), and the DNA fragment of TRE3 G-EF1 a-Tet-ON®3G-p2 A-eGFP was inserted into the lentiviral backbone of plentiSAMv2, a gift from Feng Zhang, using NheI and KpnI restriction enzymes (NEB) (Addgene plasmid #75112; http: / / n2 t.net / addgene:75112 ; RRID: Addgene_75112) (Joung et al., 2017). The recombinase was inserted into pLentiX using restriction enzyme digestion and ligation with BsrGI-HF and XbaI (NEB). For ChIP-seq experiments, pLentiX was further modified to fuse the recombinase to eGFP. Briefly, the GFP cassette was replaced with a puromycin cassette using HpaI restriction enzyme digestion. Additionally, the GFP cassette was inserted downstream of the XbaI restriction site. The recombinase was PCR amplified using a reverse primer that removed the stop codon from the recombinase coding sequence and cloned using BsrGI and XbaI.

[0108] pLenti-EF1 a-p2 A-Puro was generated from the pLentiX vector by replacing the TRE3 G promoter with a minimal EF1 a promoter using NheI and AscI. The Tet-On-3 G coding sequence and the additional EF1 a promoter were then removed using XbaI and AgeI restriction enzymes. The resulting plasmid, pLenti-EF1 a-p2 A-Puro, allows for the cloning of recombinases using BsrGI and XbaI restriction enzyme digestion.

[0109] The coding sequence of RecHTLV was codon-optimized for mammalian expression and synthesized as a fragment by Twist Bioscience (South San Francisco, CA). The codon-optimized RecHTLV was then cloned into the vector used for virus experiments (pLenti-EF1 a-p2 A-Puro) by digestion with BsrGI and XbaI restriction enzymes.

[0110] In addition, the following plasmids were used in transient transfection experiments: pcDNA3.1 (Life Technologies); the Tax-1 wild-type expression vector pHpX-Tax (Nerenberg et al., 1987); the full-length HTLV-1 proviral clone pCS-HTLV-X1MT (Dull et al., 1998); and the HTLV-1 packaging vector pCMV-HT1 M, which has a deletion in the 5' LTR and expresses all HTLV-1 gene products under the control of the cytomegalovirus (CMV) promoter (Derse et al., 2001).

[0111] For the delivery of RecHTLV in SP cells, an HIV-1-derived replication-incompetent lentiviral vector (SIN LV RecHTLV) was constructed. This vector provides high safety through its split packaging system and self-inactivating (SIN) vector design (Dull et al., 1998). The RecHTLV coding sequence and a P2A peptide puromycin resistance gene expression cassette were placed under the control of the EF1a promoter. To this end, the following sequences with homologous overhangs were amplified by PCR: the coding RecHTLV sequence (insert 1) from pLentiX-EF1a-flag-NLS-G9co-P2A-puro; the P2A puromycin resistance gene cassette (insert 2) from plasmid 1555; and the lentiviral plasmid backbone from plasmid 1539. For the lentiviral GFP control vector, insert 1 was replaced with GFP. The GFP coding sequence was amplified from plasmid 1455. The lentiviral vector backbone was also amplified. This construction was performed using the NEBuilder® HiFi DNA Assembly Cloning Kit (New England BioLabs Inc.) according to the manufacturer's instructions.

[0112] Recombinase library construction and evolution strategy The initial library was generated by shuffling a previously generated Cre-derived library. Evolution of substrate-binding proteins was performed by cloning a library of recombinases into the pEVO vector with corresponding target sites (subsites to final target site) as previously described (Buchholz & Stewart, 2001; Sarkar et al., 2007; Karpinski et al., 2016; Lansing et al., 2020, 2022).

[0113] XL1 Blue cells were transformed with the pEVO library and cultured in 100 mL of LB medium containing chloramphenicol (25 μg / mL). Library expression was induced overnight by adding arabinose to the system. After 14–16 h, 10 mL of culture was used to extract plasmids using the GeneJET Plasmid Miniprep Kit. 500 ng of plasmid DNA was digested with NdeI and AvrII (NEB) restriction enzymes, and 25 ng of the digested DNA was used as a PCR template to amplify active recombinases within the library. PCR was performed using MyTaq polymerase (Bioline), which lacks proofreading activity and therefore introduces mutations. The PCR product was digested with XbaI and BsrGI, and the recombinase band (approximately 1 kb) was extracted from an agarose gel. The new insert was cloned into the pEVO backbone, thereby initiating a new evolutionary cycle.

[0114] Shuffling was performed every fifth cycle of evolution. To this end, 25 ng of NdeI- and AvrII-digested minipreps containing the recombinase library were PCR-amplified using Herculase II Fusion polymerase (Agilent). The PCR products were purified and sonicated using a Covaris M220 to obtain approximately 200-300 bp fragments (50 W peak incident power, 200 cycles per second, 150 s run time, 20% duty cycle). The fragments were reassembled by PCR using the sonicated fragments as templates. This PCR was followed by a second reassembly PCR to amplify the full-length recombinase using Herculase II Fusion polymerase. The PCR was purified using the ISOLATE II PCR and Gel Kit (Bioline) according to the manufacturer's protocol, and the eluate was digested with XbaI and BsrGI-HF (NEB). After gel purification, the digested product was used as an insert for ligation into the pEVO plasmid to continue evolution.

[0115] Test digestion for library / recombinase activity assessment To assess the activity of the library during evolution, 500 ng of miniprep DNA plasmid from the induced culture was digested with XbaI and BsrGI-HF restriction enzymes (NEB), and 250 ng of the digest was scanned on a 0.8% agarose gel stained with RedSafe (Intron Biotechnology). A similar digestion was performed using 5 mL of plasmid DNA (500 ng) extracted from the induced culture to test individual clones.

[0116] Quantification of recombination from agarose gels Recombination efficiency was quantified by the ratio of unrecombined to recombinant bands obtained from test digests of the pEVO plasmid. Gel images were acquired using an Infinity VX2-3026 transilluminator and Infinity Capt software (Vilber). Band intensities were calculated using GelAnalyzer 19.1 (www.gelanalyzer.com) software, developed by Dr. Istvan Lazar Jr. and Dr. Istvan Lazar Sr., CSc.

[0117] Deep sequencing of recombinase libraries To prepare the loxHTLV library for deep sequencing, 500 ng of pEVO plasmid obtained from the last cycle of evolution was digested with the restriction enzymes NdeI and AvrII (NEB) to select for active mutants within the library. The digest was further desalted using an MF membrane and transformed into XL1-Blue E. coli competent cells (Agilent). The bacterial culture was grown for 14–16 h in 100 mL of LB medium containing 25 μg / mL chloramphenicol. Plasmid DNA was extracted using a GeneJet Plasmid Miniprep Kit (ThermoFisher) according to the manufacturer's protocol. Five μg of this plasmid DNA was digested with BsrGI-HF and XbaI (NEB) restriction enzymes, and the fragment containing the 1041 bp recombinase library was enriched by double-binding the pEVO backbone to custom SPRI beads (Ramawatar & Schwessinger, 2018). The supernatant was purified with Ampure XP beads (Beckman Coulter), and the resulting DNA was quantified using the Qubit HS Assay Kit on a Qubit 2.0 Fluorometer (ThermoFisher) and sequenced with the Sequel System 6.0 using the PacBio HiFi method at the CRTD Deep Sequencing Facility. Circular consensus sequence data was generated using PacBio's ccs v3.4.1 and filtered to retain only sequences between 1034 and 1200 bp in length and meet a minimum predicted accuracy of 99.997% (Phred score approximately 25). Data was converted to FASTA format using SAMtools 1.11, and the recombinase coding sequence was translated into amino acids using the protein2DNA:bestfit alignment model in exonerate v2.3.0. The sequences were further filtered and confirmed to all begin with a methionine and be 344 amino acids in length using grep, sed, and awk. Further processing and analysis of sequence data was performed using dplyr, the Sequence tool ( https: / / github.com / ltschmitt / SequenceTools ) and was performed in R (v4.1.0) using the ggplot2 package.

[0118] cell culture HeLa Kyoto (MPI-CBG, Dresden) cell line was cultured in Dulbecco's modified Eagle's medium (DMEM, Gibco) containing 10% fetal bovine serum and 1% penicillin / streptomycin (Pen / Strep) in a humidified incubator at 37°C with 5% CO. HEK-293 T cells were maintained in DMEM (GIBCO, Life Technologies, Darmstadt, Germany) containing 10% fetal bovine serum (FCS; Capricorn Scientific, Ebsdorfergrund, Germany), L-glutamine (0.35 g / L), and penicillin / streptomycin (0.12 g / L each). The CD4+ T cell line Jurkat (Schneider et al., 1977) was cultured in RPMI 1640 (45%; GIBCO, Life Technologies) and Panserin 401 medium (45%; PAN-Biotech, Aidenbach, Germany) supplemented with 10% FCS, L-glutamine, and penicillin / streptomycin. The CD4+ T cell line C91-PL (Ho et al., 1984), transformed in vitro with HTLV-1, was cultured in RPMI 1640 containing 10% FCS, L-glutamine, and penicillin / streptomycin. To confer puromycin resistance, C91-PL cells were transduced with the pLenti-EF1 a-p2 A-Puro empty vector as previously described (Millen et al., 2019) and then cultured in C91-PL medium containing 1 μg / ml puromycin.

[0119] HTLV-1-infected cells (SP) (NIH AIDS Reagent Program, #3059) were cultured in RPMI 1640 (LONZA) medium containing 2.0 mM L-glutamine (PAN Biotech), 100 U / ml penicillin-100 mg / ml streptomycin (Merck), 10% FBS (PAN Biotech), and 250 U / ml human IL-2 (Biomol) at 37°C in a humidified incubator with 5% CO.

[0120] Screening for tolerant recombinases in HeLa cells The final library was cloned into the pLentiX vector, and reporter loxHTLV cells were infected with lentiviral particles containing the recombinase. Cells were transduced at a low MOI to ensure that only one recombinase was inserted into each cell. Library expression was induced with 50 ng / mL doxycycline for 7 days. GFP+ cells were isolated by flow cytometry and grown for 5 days, after which GFP+mCherry+ cells were again isolated to recover active recombinase from the library. Approximately 3500 cells were sorted twice.

[0121] PCR to screen active clones Genomic DNA from pooled sorted cells was extracted using a QIAamp DNA Blood Kit according to the manufacturer's protocol. The recombinase was then recovered by PCR using high-fidelity Herculase II Phusion DNA polymerase (Agilent) (primers 21-22). The PCR product was cloned into pEVO, and the bacterial transformation was plated onto chloramphenicol plates. The next day, 96 colonies were transferred to a deep-well 96-well plate and grown overnight in chloramphenicol LB containing 10 μg / mL L-arabinose. Colony PCR was performed using 1 μL of the overnight culture as template and MyTaq polymerase (Bioline) (primers 82-84). The PCR products were loaded onto a 2% agarose gel for analysis.

[0122] Cell culture plasmid transfection assay Plasmid assays in cell culture were performed to test the on-target activity of recombinases as follows: 15,000 HeLa cells were seeded per well in a 96-well plate. The following day, 100 ng of reporter plasmid and 150 ng of pIRES expression plasmid were transfected using Lipofectamine 2000 (0.3 μL per well). 48 hours after transfection, cells were analyzed using MascQuant VYB (Miltenyi). After single-cell gating, recombination efficiency was calculated by dividing the percentage of double-positive cells (mCherry+GFP+) by the frequency of all GFP+ cells.

[0123] Tests with genomic reporter cells were performed as follows: 100,000 cells were seeded per well in a 24-well plate and transfected with 100 ng of pIRES expression plasmid using 1 μL of Lipofectamine per well. Forty-eight hours after transfection, half of the cells were analyzed for mCherry and GFP expression using a MacsQuant VYF flow cytometer (Miltenyi), and genomic DNA was extracted from the other half of the cells using a QIAamp DNA Blood Mini Kit (Qiagen) according to the manufacturer's protocol. To detect reporter recombination by PCR, 100 ng of genomic DNA was used as a template (primers 25-26). For other transfections using HEK293 T cells, 5 × 10 cells were seeded in a 6-well plate 24 hours before transfection. 5 293 T cells were seeded at 1000 x g for 1 hour. Transfection was performed using GeneJuice® transfection reagent (Merck Millipore, Darmstadt, Germany) according to the manufacturer's protocol, using a total of 2 μg of DNA at a 1:1 ratio of plasmid to 2 μg.

[0124] Jurkat T cells were transfected as previously described (Mann et al., 2014) with minor modifications. Briefly, 5 × 10 6 The cells were electroporated with a total of 50 μg of DNA (30 μg of pLenti-EF1 a-P2 A-Puro and 20 μg of pHpX-Tax) at 290 V and 1500 μF using a Gene Pulser X Electroporation System (BioRad).

[0125] Lentivirus production The lentiviral transfer plasmid was co-transfected into HEK293 T cells using a standard polyethyleneimine (PEI) transfection method with a lentiviral gag / pol packaging plasmid (psPAX2, Addgene #12260) and the envelope plasmid VSV-G (pMD2.G, Addgene #12259) at a molar ratio of 3:1:1. Forty-eight hours after transfection, the supernatant containing the viral particles was collected and filtered through a 0.45 mm pore size PVDF membrane filter (Millipore). The supernatant was either used directly to infect the desired target cells or stored at -80°C for later use.

[0126] VSV-g pseudotyped lentiviral particles were generated for transduction of SP cells by transiently cotransfecting HEK293 T cells with the lentiviral plasmid and the respective packaging plasmid (Dull et al., 1998). TransitLT-1 was used as the transfection reagent, following the manufacturer's protocol (Mirus Bio LLC). Seventy-two hours after transfection, the viral supernatant was collected and passed through a 0.2 μm pore size filter to remove viral aggregates. The viral supernatant was concentrated by ultracentrifugation using a 20% sucrose cushion and resuspended in RPMI medium. For titer determination (IU / ml), HEK293 T cells were transduced with various volumes of vector. Genomic DNA was isolated 72 hours later and analyzed by ddPCR using primers 73, 74, and 75 and the housekeeping gene ribonuclease P / MRP subunit P30 (RPP30; PrimePCR). (商標) ddPCR (商標) Copy number assay: titers were determined using primers specific for RPP30 (Human, Bio-Rad). Supernatants were stored at -80°C for later use.

[0127] Transduction and infection of cell lines HeLa Kyoto or Jurkat cells were transduced with lentiviral particles generated from the corresponding lentiviral vectors. For constructs containing the puromycin cassette, cells were selected with 2 μg / mL puromycin for 7 days at 72 hours post-transduction.

[0128] SP and HEK293 T cells were transduced with various amounts of vector in the presence of 5 μg / mL protamine sulfate (Merck) and spun at 450 g for 10 minutes at room temperature. After spinoculation, cells were cultured at 37°C and 5% CO2 until further use. SP cells were selected with 0.25 to 1 μg / mL puromycin for 7 days.

[0129] Toxicity assay A HeLa-derived cell line carrying a GFP-expressing transgene and a tetracycline-inducible recombinase system was transduced with approximately 46% RecHTLV and approximately 89% blank control and cultured in tetracycline-free medium (Capricorn Scientific, FBS-TET-12 A). Cells were seeded at 15,000 cells / well in flat-bottom 96-well plates and continuously induced with 100 ng / mL doxycycline for 3, 6, or 9 days. The percentage of cells expressing GFP was measured by flow cytometry using a MacsQuant VYB.

[0130] ChIP-seq and qPCR validation ChIP-seq experiments were performed using RecHTLV, RecHTLV-H289 L (G9), and additional clones (E5 and E5-H289 L). The recombinase was fused to eGFP and cloned into a modified tetracycline-inducible pLentiX vector. HeLa cells were infected with the lentivirus, selected with 2 μg / mL puromycin 72 hours posttransduction, and maintained for 7 days. Cells were cultured in 10 cm dishes, and expression of the GFP-fused recombinase was induced with 100 ng / mL doxycycline for 24 hours. Cells were examined microscopically for nuclear GFP expression to confirm the construct was functional. Cells were crosslinked with 1% formaldehyde for 10 minutes, and chromatin was extracted and sheared using the Covaris truChIP® Chromatin Shearing Kit according to the manufacturer's protocol for large cell numbers. Chromatin was sheared using a Covaris M220 sonicator. 1% of the sheared chromatin was isolated as the input sample for further qPCR validation, and the remaining sample was used for immunoprecipitation. Recombinase-protein fusions were immunoprecipitated using a goat GFP antibody (MPI-CBG antibody facility) and Protein G Sepharose beads (Protein G Sepharose, 4 Fast Flow, GE Healthcare). The complexes were washed once with low-salt immune complex wash buffer, once with strong-salt wash buffer, once with LiCl wash buffer, and once with TE buffer, and then eluted with elution buffer (1% SDS, 0.1 M NaHCO3). The eluted and input samples were treated with RNAse, reverse-crosslinked, and further purified using the ISOLATE II PCR and Gel Kit (Bioline). The eluted DNA was sent to the CRTD deep sequencing facility for library preparation and Illumina sequencing.Illumina libraries were prepared using 17-75 ng of ChIP DNA with the NEBNext Ultra™ DNA Library Prep Kit, followed by 15 PCR amplification cycles and size selection using AMPure XP beads. Paired-end sequencing was performed on an Illumina HiSeq2000. Sequencing reads were aligned to the GRCh38.p12 human genome assembly using the STAR aligner (Dobin & Gingeras, 2015) with ChIP-Seq analysis parameters. Peak calling was performed using Genrich (https: / / github.com / jsh58 / Genrich) based on the ENCODE blacklist (v2) (Amemiya et al., 2019). All genomic interval manipulations and comparisons were performed using BEDTools (Quinlan & Hall, 2010), and visualizations of pile-up reads were generated using the UCSC Genome Browser (Kent et al., 2002; Raney et al., 2014) directly from BAM files sorted with the samtools command line tool (Li et al., 2009).

[0131] Recombination assay of ChIP-seq selected peaks in bacteria In the plasmid-based assay, 10 peaks were selected in bacteria for further recombination testing. 92 bp around the top of the peak was cloned twice into a modified pEVO vector using the BglII and PciI restriction sites. RecHTLV recombinase was cloned into pEVO using the BsrGI and XbaI restriction sites. After transformation in XL1blue E. coli, recombinase expression was induced with 200 μg / mL L-arabinose. Plasmid extractions were performed, and recombination was analyzed using test digests (see above).

[0132] Western blot At the indicated time points (48 or 72 hours posttransfection), transfected cells were resuspended in lysis buffer (150 mM NaCl, 10 mM Tris / HCl (pH 7.0), 10 mM EDTA, 1% Triton™ X-100, 2 mM DTT, and the protease inhibitors leupeptin, aprotinin (20 μg / ml each), and 1 mM phenylmethylsulfonyl fluoride (PMSF)) and subjected to repeated freeze-thaw cycles between -196°C (liquid nitrogen) and 30°C, followed by three 20-second sonications. Equal amounts of protein (40 μg) were heated at 95°C for 5 minutes in sodium dodecyl sulfate (SDS) loading dye containing 10 mM Tris / HCl (pH 6.8), 10% glycerol, 2% SDS, 0.1% bromophenol blue, and 5% β-mercaptoethanol. SDS-PAGE and immunoblotting were performed using Immobilon (登録商標) -FL PVDF (Merck Millipore, Billerica, MA, United States) or nitrocellulose transfer membrane (Whatmann (登録商標) , Protran (登録商標) , Whatmann GmbH, Dassel, Germany) according to standard protocols. Proteins were detected using the following primary antibodies: mouse anti-Tax (derived from hybridoma cell line 168B17-46-34, provided by B. Langton of the AIDS Research and Reference Reagent Program, NIAID, NIH) (Langton et al., 1988), mouse monoclonal anti-HTLV-1 gag p19 (TP-7, ZeptoMetrix Corporation), mouse monoclonal anti-FLAG (M2, Sigma), mouse monoclonal anti-Hsp90α / β (F-8, Santa Cruz Biotechnology), and mouse monoclonal anti-α-tubulin (T9026, Sigma). Secondary antibodies were anti-mouse Alexa Fluor 800 (Alexa Fluor 800). (登録商標)The antibodies were either anti-mouse (Life Technologies) or anti-horseradish peroxidase (HRP; GE Healthcare, Little Chalfont, United Kingdom) conjugated with anti-mouse antibody. Fluorescent or chemiluminescent signals were detected using an advanced fluorescence imager camera (ChemoStar, Intas Science Imaging GmbH, Gouttingen, Germany). Densitometric analysis was performed using an Advanced Image Data Analyzer (AIDA Version 4.22.034, Raytest Isotopenmessgeraete GmbH, Straubenhardt, Germany) to compare the expression levels of Tax or Gag.

[0133] Co-culture assay of chronically infected C91-PL-Puro donor cells with pre-stained Jurkat-pLenti-EF1 a acceptor cells 1*10 6 C91-PL-Puro cells were cultured at 1 x 10 6Jurkat-pLenti-EF1a T cells were cocultured with Jurkat-pLenti-EF1a T cells in a 1:1 mixture of C91-PL and Jurkat medium containing puromycin for 5 days to maintain recombinase expression in the Jurkat-pLenti-EF1a T cells. The cocultures were split into two at a 1:2 ratio 24 and 72 hours after the start of the experiment. Jurkat-pLenti-EF1a T cells were pre-stained with the viability dye CellTracker™ Blue 7-amino-4-chloromethylcoumarin (CMAC; Thermo Fisher Scientific) as previously described (Donhauser et al., 2018). As a control, pre-stained Jurkat-pLenti-EF1a T cells were pre-treated with 10 μM of the integrase inhibitor raltegravir 24 hours before the start of coculture. Raltegravir was replenished every time new medium was added to the coculture, maintaining a final concentration of 10 μM. Similarly, Jurkat-pLenti-EF1a T cells were treated with DMSO as a solvent control. Cells were stained with mouse anti-Tax antibody (Langton et al., 1988) and mouse AlexaFluor® 647-conjugated secondary antibody (Life Technologies). Briefly, cells were permeabilized with 0.3% Triton™ X-100, dispensed into FACS buffer (PBS / 1% FCS / 2 mM EDTA), and incubated at room temperature for 10 minutes. Primary and secondary antibodies were diluted in FACS buffer and stained at 37°C for 45 minutes, with three wash steps between. Cocultures were analyzed using a BD LSR II flow cytometer (BD Biosciences, San Jose, CA, USA). Cells were differentiated by CMAC staining (Jurkat-pLenti-EF1a: CMAC positive; C91-PL-Puro: CMAC negative) and by size (FSC / SSC). The percentage of Tax-positive cells and CMAC-positive cells (Jurkat-pLenti-EF1 a T cells) within the Tax-positive C91-PL donor was investigated to measure productive infection.For statistical analysis, coefficients were estimated by linear mixed models followed by t-tests using the Satterthwaite method provided by the lmerTest package in R ( Kuznetsova et al., 2017 ).

[0134] Verification of HTLV1 insertion sites in SP cells To confirm the insertion sites on chromosomes 6 and 20 and demonstrate the integrity of the proviral LTRs in SP cells as published in (Meissner et al., 2017), sequence regions were amplified and confirmed by Sanger sequencing using the same primers.

[0135] Detection of genomic scars in SP cells Genomic DNA from cells transfected with RecHTLV or the GFP control vector was isolated 9 or 10 days after transfection and subjected to PCR to detect the genomic scar. PCR fragments were visualized by agarose gel electrophoresis and confirmed by Sanger sequencing using the same primers.

[0136] Reference list Abi-Ghanem J, Chusainow J, Karimova M, Spiegel C, Hofmann-Sieber H, Hauber J, Buchholz F & Pisabarro MT (2013). Nucleic Acids Res 41: 2394-2403 Abi-Ghanem J, Samsonov SA & Pisabarro MT (2015). J Comput Aided Mol Des 29: 271-282 Alais S, Dutartre H & Mahieux R (2017). Methods Mol Biol 1582: 47-55 Alais S, Mahieux R & Dutartre H (2015). J Virol 89: 10580-10590 Alper H, Fischer C, Nevoigt E & Stephanopoulos G (2005). Proc Natl Acad Sci U S A 102: 12678-12683 Amemiya HM, Kundaje A & Boyle AP (2019). Sci Rep 9: 9354 Araujo THA, Souza-Brito LI, Libin P, Deforche K, Edwards D, de Albuquerque-Junior AE, Vandamme A-M, Galvao-Castro B & Alcantara LCJ (2012). PLoS One 7: e42123 Boxus M & Willems L (2009). Br J Cancer 101: 1497-1501 Buchholz F & Stewart AF (2001). Nat Biotechnol 19: 1047-1052 Coffin JM, Hughes SH & Varmus HE (1997) The Interactions of Retroviruses and their Hosts. In Retroviruses, Coffin JM Hughes SH & Varmus HE (eds) Cold Spring Harbor (NY): Cold Spring Harbor Laboratory Press De Silva Feelixge HS, Stone D, Pietz HL, Roychoudhury P, Greninger AL, Schiffer JT, Aubert M & Jerome KR (2016). Antiviral Res 126: 90-98 Derossi D, Calvet S, Trembleau A, Brunissen A, Chassaing G & Prochiantz A (1996). J Biol Chem 271: 18188-18193 Derossi D, Joliot AH, Chassaing G & Prochiantz A (1994). J Biol Chem 269:10444–10450 Derse D, Hill SA, Lloyd PA, Chung Hk null & Morse BA (2001). J Virol 75:8461–8468 Derse D, Mikovitz J & Ruscetti F (1997). Virology 237:123–128 Ding L, Paszkowski-Rogacz M, Mircetic J, Chakraborty D & Buchholz F (2021). Life Sci Alliance 4: e202000792 Ding L, Paszkowski-Rogacz M, Winzi M, Chakraborty D, Theis M, Singh S, Ciotta G, Poser I, Roguev A, Chu WK, et al (2015). Cell Systems 1:141–151 Dobin A & Gingeras TR (2015). Curr Protoc Bioinformatics 51:11.14.1–11.14.19 Donhauser N, Heym S & Thoma-Kress AK (2018). Front Microbiol 9:400 Donovan PJ & Gearhart J (2001). Nature 414:92–97 Dull T, Zufferey R, Kelly M, Mandel RJ, Nguyen M, Trono D & Naldini L (1998). J Virol 72:8463–8471 Edelmann GM, Meech R, Owens GC, Jones FS (2000). PNAS 97 , 3038 – 3043 Elliott G & O'Hare P (1997). Cell 88:223–233 Eroshenko N & Church GM (2013). Nat Commun 4:2509 Fan N, Gavalchin J, Paul B, Wells KH, Lane MJ & Poiesz BJ (1992). J Clin Microbiol 30:905–910 Fawell S, Seery J, Daikh Y, Moore C, Chen LL, Pepinsky B & Barsoum J (1994). Proc Natl Acad Sci USA 91:664–668 Furukawa Y, Kubota R, Tara M, Izumo S & Osame M (2001). Blood 97:987–993 Guillen-Pingarron C, Guillem-Gloria PM, Soni A, Ruiz-Gomez G, Augsburg M, Buchholz F, Anselmi M & Pisabarro MT (2022). Comput Struct Biotechnol J 20:989–1001 Guo F, Gopaul DN & van Duyne GD (1997). Nature 389:40–46 Hartenbach S & Fussenegger M (2006). Biotechnol Bioeng 95:547–559 Hauber I, Hofmann-Sieber H, Chemnitz J, Dubrau D, Chusainow J, Stucka R, Hartjen P, Schambach A, Ziegler P, Hackmann K, et al (2013). PLoS Pathog 9: e1003587 Hein MY, Hubner NC, Poser I, Cox J, Nagaraj N, Toyoda Y, Gak IA, Weisswange I, Mansfeld J, Buchholz F, et al (2015). Cell 163:712–723 Hermann M, Stillhard P, Wildner H, Seruggia D, Kapp V, Sanchez-Iranzo H, Mercader N, Montoliu L, Zeilhofer HU & Pelczar P (2014). Nucleic Acids Res 42:3894–3907 Joung J, Konermann S, Gootenberg JS, Abudayyeh OO, Platt RJ, Brigham MD, Sanjana NE & Zhang F (2017). Nat Protoc 12:828–863 Karpinski J, Hauber I, Chemnitz J, Schafer C, Paszkowski-Rogacz M, Chakraborty D, Beschorner N, Hofmann-Sieber H, Lange UC, Grundhoff A, et al (2016). Nat Biotechnol 34:401–409 Kent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM & Haussler and D (2002). Genome Res 12:996–1006 Kim ST, Kim GW, Lee YS & Park JS (2001). J Cell Biochem 80:321–327 Kuznetsova A, Brockhoff PB & Christensen RHB (2017). J Stat Soft 82 Langton B, Sliwkowski M, Tran K, Knapp S, Keitelmann E, Smith C, Wallingford S, Liu H, Ralston J & Brandis J (1988). Med Virol 8: 295 Lansing F, Mukhametzyanova L, Rojo-Romanos T, Iwasawa K, Kimura M, Paszkowski-Rogacz M, Karpinski J, Grass T, Sonntag J, Schneider PM, et al (2022). Nat Commun 13: 422 Lansing F, Paszkowski-Rogacz M, Schmitt LT, Schneider PM, Rojo Romanos T, Sonntag J & Buchholz F (2020). Nucleic Acids Res 48: 472-485 Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R, & 1000 Genome Project Data Processing Subgroup (2009). Bioinformatics 25: 2078-2079 Lin Q, Jo D, Gebre-Amlak KD & Ruley HE (2004). BMC Biotechnol 4: 25 Macara IG (2001) Transport into and out of the nucleus. Microbiol Mol Biol Rev 65: 570-594, table of contents Mann MC, Strobel S, Fleckenstein B & Kress AK (2014). Virology 464-465: 98-110 Martin F, Tagaya Y & Gallo R (2018). The Lancet 391: 1893-1894 Meinke G, Bohm A, Hauber J, Pisabarro MT & Buchholz F (2016) C. Chem Rev 116:12785–12820 Meissner ME, Mendonca LM, Zhang W & Mansky LM (2017). J Virol 91:e00369–17 Millen S, Gross C, Donhauser N, Mann MC, Peloponese Jr. JM & Thoma-Kress AK (2019). Front Microbiol 10:2439 Nakagawa M, Shaffer AL, Ceribelli M, Zhang M, Wright G, Huang DW, Xiao W, Powell J, Petrus MN, Yang Y, et al (2018). Cancer Cell 34:286–297.e10 Nerenberg M, Hinrichs SH, Reynolds RK, Khoury G & Jay G (1987). Science 237: 1324-1329 Nolden L, Edenhofer F, Haupt S, Koch P, Wunderlich FT, Siemen H & Bruestle O (2006). Nat Methods 3:461–467 Nurk S, Koren S, Rhie A, Rautiainen M, Bzikadze AV, Mikheenko A, Vollger MR, Altemose N, Uralsky L, Gershman A, et al (2022). Science 376:44–53 Oess S & Hildt E (2000). Gene Ther 7:750–758 Peitz M , Pfannkuche K , Rajewsky K & Edenhofer F (2002). Proc Natl Acad Sci USA 99:4489–4494 Pique C & Jones KS (2012). Front Microbe 3 Popovic M , Lange-Wantzin G , Sarin PS , Mann D & Gallo RC ( 1983 ). Proc Natl Acad Sci USA 80:5402–5406 Poser I, Sarov M, Hutchins JRA, Heriche JK, Toyoda Y, Pozniakovsky A, Weigl D, Nitzsche A, Hegemann B, Bird AW, et al (2008). Nat Methods 5:409–415 Quinlan AR & Hall IM (2010). Bioinformatics 26: 841–842 Ramawatar & Schwessinger B (2018) DNA size selection (>3-4kb) and purification of DNA using an improved homemade SPRI beads solution. protocols.io Raney BJ, Dreszer TR, Barber GP, Clawson H, Fujita PA, Wang T, Nguyen N, Paten B, Zweig AS, Karolchik D, et al (2014). Bioinformatics 30:1003–1005 Richard JP , Melikov K , Brooks H , Prevot P , Lebleu B & Chernomordik LV (2005). J Biol Chem 280:15300–15306 Rongrong L, Lixia W & Zhongping L (2005). Acta Biochim Pol 52: 541-544 Ruefer AW & Sauer B (2002). Nucleic Acids Res 30: 2764-2771 Sanjana NE, Shalem O & Zhang F (2014) I. Nat Methods 11: 783-784 Santoro SW & Schultz PG (2002). Proceedings of the National Academy of Sciences 99: 4185-4190 Sarkar I, Hauber I, Hauber J & Buchholz F (2007). Science 316: 1912-1915 Schambach A, Bohne J, Chandra S, Will E, Margison GP, Williams DA, Baum C (2006). Molecular Therapy 13, 391-400 Scherr M & Eder M (2002). Curr Gene Ther 2: 45-55 Schmitt LT, Paszkowski-Rogacz M, Jug F & Buchholz F (2022) Prediction of designer-recombinases for DNA editing with generative deep learning. 2022.04.01.486669 doi:10.1101 / 2022.04.01.486669 [PREPRINT] Schneider U, Schwenk HU & Bornkamm G (1977). Int J Cancer 19: 621-626 Seegulam ME & Ratner L (2011). Antimicrob Agents Chemother 55: 2011-2017 Shimshek DR, Kim J, Huebner MR, Spergel DJ, Buchholz F, Casanova E, Stewart AF, Seeburg PH & Sprengel R (2002). Genesis 32: 19-26 Soni A, Augsburg M, Buchholz F & Pisabarro MT (2020). Sci Rep 10: 13985 Surendranath V, Chusainow J, Hauber J, Buchholz F & Habermann BH (2010). Nucleic Acids Res 38:W293–298 Sueruen D, Schwaeble J, Tomasovic A, Ehling R, Stein S, Kurrle N, von Melchner H & Schnuetgen F (2018). Mol Ther Nucleic Acids 10:1–8 Tanaka A, Takeda S, Kariya R, Matsuda K, Urano E, Okada S & Komano J (2013). Leukemia 27:1621–1627 Vives E, Brodin P & Lebleu B (1997). J Biol Chem 272:16010–16017 Vives E, Richard JP, Rispal C & Lebleu B (2003). Curr Protein Pept Sci 4:125–132 Wang G, Zhao N, Berkhout B & Das AT (2016a). Mol Ther 24:522–526 Wang Z, Pan Q, Gendron P, Zhu W, Guo F, Cen S, Wainberg MA & Liang C (2016b). Cell Rep 15:481–489 Warren D, Laxmikanthan G & Landy A (2008). Proceedings of the National Academy of Sciences 105: 18278-18283 Wattel E, Vartanian JP, Pannetier C & Wain-Hobson S (1995). J Virol 69: 2863-2868 Yoder KE & Bundschuh R (2016). Sci Rep 6: 29530 EP 1 1000751.5 WO 2002 / 44409 WO 2008 / 083931 WO 2011 / 147590 WO 2016 / 034553

Claims

1. 1. A nucleic acid encoding an adjusted recombinase capable of recombining asymmetric target sequences of SEQ ID NO: 1 within the long terminal repeats of proviral DNA of multiple HTLV-1 strains, wherein the amino acid sequence of said adjusted recombinase has at least 70% sequence identity with the sequence set forth in SEQ ID NO: 13, and said adjusted recombinase contains at least amino acid substitutions selected from the group consisting of P15L, S38P, A84T, Q90K, S108G, K122R, I166V, A175S, N245Y, E266G, and T268A compared to SEQ ID NO:

11.

2. 2. The nucleic acid of claim 1, wherein the regulated recombinase has at least 80%, preferably at least 90%, sequence identity to SEQ ID NO:

13.

3. 10. The nucleic acid of any preceding claim, wherein the regulated recombinase further comprises SEQ ID NO:

38.

4. 10. The nucleic acid of any of the preceding claims, wherein the modified recombinase further comprises at least one, preferably 2 to 101, of the following defined amino acids at the indicated positions compared to SEQ ID NO: 11: 3K, 3N or 3D, preferably 3K; 5L, 5Q or 5P, preferably 5L; 7L or 7I, preferably 7L; 8H, 8P or 8Y, preferably 8H; 9Q or 9H, preferably 9Q; 10S or 10D or 10N, preferably 10S; 12S or 12F or 12T, preferably 12S; 14L or 14S, preferably 14L; 16A or 16V, preferably 16A; 17D or 17G or 17N, preferably 17D; 18A or 18V, preferably 18V; 19T or 19A or 19M, preferably 19T; 23A or 23T, preferably 23A; 25K, 25Q or 25R, preferably 25K; 26N or 26S, preferably 26N; 28M or 28T, preferably 28M; 29D or 29V or 29I, preferably 29D; 30M or 30T or 30V, preferably 30M; 31F or 31L, preferably 31F; 34R or 34H or 34C, preferably 34R 35Q or 35H, preferably 35Q; 39E or 39A, preferably 39E; 43K or 43E or 43R, preferably 43E; 45L or 45F, preferably 45L; 51S or 51A or 51T, preferably 51S; 54A or 54T, preferably 54A; 57E or 57K, preferably 57E; 58S or 58L, preferably 58S; 59N or 59D, preferably 59N; 62K or 62R, preferably 62K; 63W or 63R, preferably 63W; 68P or 68H, preferably 68P; 77H or 77Y or 77D, preferably 77H; 86N or 86S, preferably 86N; 88V or 88I, preferably 88V; 93A or 93T, preferably 83A; 94E or 94D, preferably 94E; 97T or 97M, preferably 97T; 110S or 110N, preferably 110S; 111N or 111S, preferably 111N; 114T or 114S or 114P, preferably 114T; 125V or 125I, preferably 125V; 132K or 132R, preferably 132K; 136A or 136P, preferably 136A; 140T or 140A, preferably 140T; 144R, 144Q or 144L, preferably 144R; 147S, 147T, 147P or 147A, preferably 147S; 149M or 149L, preferably 149M; 151N, 151T, 151D or 151S, preferably 151N; 155C or 155R, preferably 155C; 156L, 156K or 156Q, preferably 156L; 160N or 160D, preferably 160N; 183K or 183R, preferably 183K; 185I or 185V, preferably 185I; 196H or 196R, preferably 196H; 206T or 206A, preferably 206T; 207A or 207T, preferably 207A; 219K, 219R or 219X, preferably 219K; 221V or 221I, preferably 221V; 225I, 225V or 225F, preferably 225I; 226S or 226T, preferably 226S; 227V or 227A, preferably 227A; 229G or 229C, preferably 229G; 230V or 230I, preferably 230V; 231A or 231T, preferably 231A; 232D or 232N, preferably 232D; 235N or 235D, preferably 235N; 241R or 241Q, preferably 241R; 244R or 244K, preferably 244R; 247V or 247I, preferably 247V; 248V or 248A, preferably 248V; 249V or 249A, preferably 249V; 250S or 250P, preferably 250S; 253T or 253A, preferably 253T; 254S, 254K, 254N or 254C, preferably 254S; 255R or 255Q, preferably 255R; 258T or 258A, preferably 258T; 259H, 259Y or 259G, preferably 259H; 261M or 261L, preferably 261M; 262Q or 262R, preferably 262Q; 263G or 263K, preferably 263G; 264I or 264V, preferably 264I; 267A or 267T, preferably 257A; 272I or 272V, preferably 272I; 276K, 276R or 276N, preferably 276K; 277D or 277G, preferably 277D; 278D or 278G, preferably 278D; 280G or 280S, preferably 280G; 281Q, 281R or 281L, preferably 281Q; 299L or 299M, preferably 299L; 304V or 304A, preferably 304V; 305S or 305P, preferably 305S; 306I, 306M or 306L, preferably 306I; 307A, 307P or 307V, preferably 307A; 317R or 317T, preferably 317R; 319N or 319D, preferably 319N; 320S, 320T or 320G, preferably 320S; 332T or 332A, preferably 332T; 341D or 341G or 341E, preferably 341D; 342G or 342D, preferably 342G; and / or 343D or 343Y, preferably 343D.

5. 5. The nucleic acid of claim 4, wherein the regulated recombinase comprises 94E or 94D, preferably 94E.

6. The adjusted recombinase i) 247V and 248V and 244R, or ii) 247I and 248A The nucleic acid according to claim 4 or 5, comprising any one of:

7. 7. The nucleic acid of claim 6, wherein the regulated recombinase comprises 247V and 248V and 244R, and preferably also 249V.

8. The adjusted recombinases are N3K, V7L, N10D, P12S, P15L, V16A, A18V, V23A, S38P, K43E, K57E, L58S, Y77H, A84T, K86N, I88V, Q90K, G93A, Q94E, M97T, S108G, S114T, K122R, Q144R, Q156L, I166V, A175S, T206 10. A nucleic acid according to any preceding claim, comprising at least 34, and preferably all, amino acid substitutions selected from the group consisting of A, S226T, V227A, K244R, N245Y, A248V, A249V, P250S, Q255R, R259H, L261M, E262Q, E266G, A267T, T268A, M299L, P307A, N317R and I320S.

9. 10. The nucleic acid of any of the preceding claims, wherein the regulated recombinase comprises the amino acid sequence shown in SEQ ID NO: 13, 14, 26-37 or 39, or any combination of said sequences.

10. 10. The nucleic acid of any of the preceding claims, wherein the regulated recombinase comprises the amino acid sequence shown in SEQ ID NO:

13.

11. A regulated recombinase encoded by a nucleic acid according to any of the preceding claims, optionally expressed as a fusion protein.

12. A transformed cell comprising a nucleic acid according to any one of claims 1 to 10, which is preferably a hematopoietic stem cell.

13. A pharmaceutical composition comprising a nucleic acid described in any one of claims 1 to 10, a regulated recombinase described in claim 11, and / or a transformed cell described in claim 12, optionally including a pharmaceutically acceptable excipient and / or solvent.

14. 14. The pharmaceutical composition of claim 13, wherein the pharmaceutical composition is used for treating or preventing HTLV-1 infection in a subject, and optionally, the composition is used for treating adult T-cell leukemia (ATL) and / or HTLV-1 associated myelopathy (HAM).

15. A method for treating a subject infected with HTLV-1, comprising administering to said subject an effective amount of the pharmaceutical composition of claim 13, said subject optionally having adult T-cell leukemia (ATL) and / or HTLV-1 associated myelopathy (HAM).