Fusion of Site-Specific Recombinases for Efficient and Specific Genome Editing

A fusion protein of recombinase monomers linked by an oligopeptide enhances genome editing specificity and efficiency, addressing the limitations of symmetric target site recognition and off-target effects in site-specific recombinases, offering a promising therapeutic approach for hemophilia A.

JP7702149B2Active Publication Date: 2025-07-03TECHNISCHE UNIVERSITAT DRESDEN
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022533406
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-06
Filing Date
2020-12-03
Publication Date
2025-07-03
Estimated Expiration
2040-12-03

AI Technical Summary

Technical Problem

Current genome editing technologies, such as programmable nucleases, face challenges with unpredictable sequence rearrangements and off-target effects, while site-specific recombinases like Cre recombinase are limited by symmetric target site recognition, making them less effective for therapeutic applications like correcting genetic disorders like hemophilia A.

Method used

A fusion protein of recombinase monomers linked by an oligopeptide is developed to specifically recognize asymmetric target sites, enhancing specificity and efficiency in genome editing, particularly for inverting the int1h sequence causing hemophilia A, by forming a heterodimer that prevents homodimer formation and off-target recombination.

Benefits of technology

The fusion protein achieves high specificity and efficiency in inverting the F8 gene sequence, reducing off-target events and improving therapeutic potential for treating hemophilia A without relying on cellular repair pathways.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702149000023
    Figure 0007702149000023
  • Figure 0007702149000024
    Figure 0007702149000024
  • Figure 0007702149000025
    Figure 0007702149000025
Patent Text Reader

Abstract

The present invention generally relates to the field of genome editing and provides a DNA recombinase that efficiently and specifically recombines a genomic target sequence through the fusion of a recombinase monomer. More specifically, the present invention provides a method for producing a fusion protein for efficient and specific genome editing, comprising a recombinase complex including at least a first recombinase enzyme, a second recombinase enzyme, and at least one linker, wherein the first recombinase enzyme and the second recombinase enzyme specifically recognize a first half-site and a second half-site of an upstream target site and / or a downstream target site of the recombinase; the first recombinase enzyme and the second recombinase enzyme are linked to each other via a linker; and the linker comprises or consists of an oligopeptide. The present invention further relates to a fusion protein produced by this method. The present invention also discloses a designer recombinase that catalyzes the inversion of a DNA sequence present in the int1h region on the human X chromosome. The present invention further relates to nucleic acid molecules encoding said DNA recombinases and fusion proteins, and to the use of said fusion proteins, DNA recombinases and nucleic acid molecules in pharmaceutical compositions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Field of the Invention) The present invention generally relates to the field of genome editing, and provides a DNA recombinase that efficiently and specifically recombines a genomic target sequence through the fusion of recombinase monomers. More specifically, the present invention provides a method for producing a fusion protein for efficient and specific genome editing, comprising a complex of recombinases comprising at least a first recombinase enzyme, a second recombinase enzyme and at least one linker, wherein the first recombinase enzyme and the second recombinase enzyme specifically recognize the first half-site and the second half-site of the upstream target site and / or the downstream target site of the recombinase; the first recombinase enzyme and the second recombinase enzyme are linked to each other via a linker; and the linker comprises or consists of an oligopeptide. The present invention further relates to a fusion protein produced by this method. The present invention also discloses a designer recombinase that catalyzes the inversion of a DNA sequence present in the int1h region on the human X chromosome. The present invention further relates to a nucleic acid molecule encoding the DNA recombinase and the fusion protein, and the use of the fusion protein, DNA recombinase and nucleic acid molecule in a pharmaceutical composition.

Background Art

[0002] (Background of the Invention) Genome manipulation is becoming an increasingly important technique in biomedical research. The main methods in today's gene editing field are the introduction of nuclease-mediated double-strand breaks (DSBs) at loci of interest, which are later corrected by cellular repair pathways. There are four types of programmable nucleases, which can be divided into two groups based on the mode of target DNA sequence recognition. Meganucleases, zinc finger nucleases (ZFNs), and transcription activator-like effecter nucleases (TALENs) use protein-DNA interactions to direct the nuclease to specific loci, while clustered, regularly interspaced, short-palindromic repeat-associated (CRISPR) endonucleases direct it using RNA-DNA interactions (Caroll, 2014; Wang et al., 2017). Programmable nucleases are candidates for therapeutic applications, and some of them are already in clinical trials (Tebas et al., 2014; Qasim et al., 2015; Cyranoski, 2016).

[0003] However, one of the main challenges of programmable nucleases is the risk of unpredictable sequence rearrangements. The introduced DSBs are repaired by cells mainly using non-homologous end joining (NHEJ) or homology-directed repair (HDR). Repair by HDR is accurate because the sequence is copied from a second allele or donor sequence that matches the target, maintaining genome stability. However, HDR is mainly active during DNA replication, and in most cells, NHEJ events outnumber HDR. NHEJ is an error-prone repair mechanism that leads to insertions and deletions (indels) in the repaired DNA fragment. This can result in harmful events due to changes in the gene sequence (Caroll, 2014; Cox et al., 2015; Kosicki et al., 2018).

[0004] Alternative tools widely used in genome engineering include site-specific recombinases (SSRs) of the tyrosine recombinase family. Tyrosine SSRs perform complete recombination reactions without cofactors and thus do not rely on cellular DNA repair pathways, giving them considerable advantages over programmable nucleases (Meinke et al., 2016). This leads to highly specific, predictable, and accurate genome editing events that are attractive for therapeutic applications.

[0005] One of the most commonly used SSRs is the tyrosine recombinase Cre from Escherichia coli bacteriophage P1, which recognizes the 34-bp symmetric loxP site consisting of two 13-bp palindromic sequences flanking an 8-bp spacer (Figure 1).

[0006] The recombination reaction carried out by SSRs is a multi-step process. First, four single recombinase molecules bind to the half-sites of each palindrome of loxP to form an active recombination synapse (Figure 2A). At this stage, one of the recombinase dimers is activated to cleave the DNA. Cleavage occurs when the tyrosine at position 324 of Cre nucleophilically attacks a labile phosphate in the spacer of the loxP site. This leads to the formation of a 3'-phosphotyrosine intermediate and the release of a 5'-hydroxyl group. Subsequently, strand exchange, caused by the attack of the released 5'-hydroxyl group on another 3'-phosphotyrosine bond, and the formation of a Holliday junction intermediate follow. Next, the Holliday junction is isomerized, the structure of the molecule changes, the active pair of Cre molecules becomes inactive, and the previously inactive Cre molecules are activated. Then, the same process of cleavage and exchange is repeated using the second set of strands (Meinke et al., 2016).

[0007] The result of recombination depends on the orientation of the spacer. When the spacer sequences of the two loxP sites are in the same orientation, recombination leads to the excision or integration of the DNA fragment (Figure 2B). When the spacers are in the opposite orientation, recombination results in the inversion of the adjacent DNA sequence (Meinke et al., 2016).

[0008] To use the Cre recombinase system, the loxP sites are typically introduced via genomic manipulation at the desired genomic locus. This approach has been successfully used for conditional mutagenesis in animal models for investigating gene function and modeling genetic events underlying human diseases (Albanese et al., 2002; Justice et al., 2011), but its application in humans is severely limited. To overcome this limitation, techniques have been developed to engineer recombinase enzymes to recombine artificial DNA sequences, for example, to excise the HIV provirus from infected cells (Buchholz and Stewart et al., 2001; Santoro and Schultz et al., 2002; Sarkar et al., 2007; Karpinski et al., 2016). The palindromicity of the target site considerably limits the number of possible target sites in the recombinable genome. Nevertheless, previous studies have shown that a single SSR can be engineered to specifically recombine asymmetric target sites (Sarkar et al., 2007; Karpinski et al., 2016). Another possibility for recombining asymmetric target sites is to construct a heterospecific recombinase system in which two different recombinase monomers bind to their respective half-sites and both enzymes cooperate to recombine the final asymmetric target site (Figure 3) (Saraf-Levy et al., 2006). It has recently been demonstrated that such a system can actually be developed for advanced genome editing in human cells (Lansing et al., 2020). Therefore, these systems have been able to greatly expand the utility of the SSR system.

[0009] However, the single recombinase of the heterodimer was also able to form homodimers that are active on symmetric sites. Thus, since each recombinase of the heterodimer has its own symmetric off-target site in addition to the putative asymmetric heterodimer off-target, using two recombinases instead of one bears the potential risk of having more off-target events. For the use of heterodimers in therapeutic applications, high efficiency and specificity of the system should be ensured. One option is the modification of the heterodimer that prevents the individual recombinases from recombining as homodimers. In this scenario, the activity on symmetric sites may be removed or restricted. Thus, creating a system that enforces the recombinase heterodimer is of high interest.

[0010] As a result, several approaches have been developed to enhance Cre specificity, including the creation of forced heterodimers by "hydrophobic size switches" (Gelato et al., 2008), forced hetero-tetrameric complexes with selective destabilization (Zhang et al., 2015), or destabilization of binding cooperativity (Eroshenko and Church et al., 2013). However, all of these modifications lead to a significant decrease in recombination activity and are likely not compatible with efficient genome modification for therapeutic applications.

[0011] One of the human diseases that may be corrected by SSR is hemophilia A (HA). HA is an X-linked monogenic disorder caused by a deficiency in the F8 gene, which encodes coagulation factor VIII (FVIII). The F8 gene is located on the long arm of the X chromosome (Xq28). This gene is 186 kb in length and consists of 27 exons.

[0012] FVIII is synthesized mainly in liver sinusoidal endothelial cells (LSEC) and other human endothelial cells (Shahani et al., 2014; Turner and Moake et al., 2015). FVIII is an important factor in the blood coagulation cascade. FVIII binds to von Willebrand factor and circulates in the blood, dissociating from it after activation by thrombin. Activated FVIII (FVIIIa) binds to FIXa, and this complex then activates FX, leading to a series of reactions that form a thrombus (Dahlback, 2000).

[0013] HA affects approximately 1 in 5,000 males (Graw et al., 2005). The clinical severity of HA depends on the residual activity level of factor VIII in the blood and is classified into three groups: mild (5 - 40% of normal levels), moderate (1 - 5%), and severe (<1%) (White et al., 2001). Patients with mild HA experience bleeding episodes only in the case of major trauma or surgery. Moderately affected individuals have spontaneous bleeding after minor trauma. Severe HA is characterized by frequent spontaneous bleeding into internal organs, muscles, and joints. Repeated joint bleeding often leads to a disabling condition called hemophilic arthropathy. This complication is characterized by chronic pain, joint impairment, and a dramatic decline in the patient's quality of life (Pandey and Mittal et al., 2001; Fischer et al., 2005; Melchiorre et al., 2017). More than half of HA patients have the severe form of this disease, which is most often caused by chromosomal inversion due to intrachromosomal recombination with homologous regions outside the F8 gene (Graw et al., 2005).

[0014] The second most common genetic change in patients with severe HA (about 1 - 5%) is inversion of exon 1. This occurs by homologous recombination between two 1 kb sequences (int1h) located telomeric to intron 1 and the F8 gene (Figure 4). The sequences are in opposite orientations and are located approximately 140 kb apart. Since homologous recombination leads to inversion and translocation of exon 1, the F8 gene is disrupted and rendered non - functional (Castaldo et al., 2007; Tizzano et al., 2003).

[0015] The standard treatment for severe HA is replacement therapy - administration of plasma - derived blood coagulation factor VIII or recombinant concentrates of blood coagulation factor VIII. However, due to its short half - life (8 - 12 hours), very frequent administrations (2 - 3 times a week or even daily) are required. This treatment has a major impact on the patient's quality of life (Peyvandi et al., 2016). Several modifications in proteins such as PEGylation and fusion with fragments of IgG have been developed to increase the half - life of recombinant FVIII, enabling the acquisition of a treatment with equivalent efficiency but with a half - life 1.5 times longer than that of standard FVIII infusions (Mahlangu et al., 2014; Konkle et al., 2015; Peyvandi et al., 2016). The extension of the half - life leads to a slight decrease in the administration frequency from twice a week to once every five days (Carcao, 2014).

[0016] In addition to frequent intravenous injections, a major challenge of current treatments is that some patients (about 30%) produce inhibitory antibodies against the administered FVIII, thereby compromising the treatment (Gouw et al., 2013). Furthermore, the economic burden of HA should also be considered. The average annual direct cost of standard treatment is approximately 200,000 euros per patient in Europe and exceeds 300,000 euros in Germany (O'Hara et al., 2017). Therefore, prevention alone in HA patients without complications is very costly. Thus, it is important to develop alternative treatment options or even curative therapies.

[0017] Alternative therapies that inhibit the anticoagulation pathway to achieve hemostasis have been developed, and some of them, such as siRNA-degrading antithrombin III mRNA and monoclonal antibodies that inhibit tissue factor pathway inhibitor or protein C, are currently in clinical trial stages I-III. The advantages of these approaches are the potential for use by all patients, regardless of the presence of FVIII inhibitors, and a lower treatment frequency (once a week) and subcutaneous administration mode. However, clinical studies have shown that careful adjustment of the treatment dose is necessary because over-dosing increases the risk of thrombotic effects. (References by Peters and Harris et al., 2018).

[0018] Another approach is to use the FVIIIa-mimicking bispecific antibody emicizumab. This antibody cross-links FIXa to FX. Therefore, it basically mimics the function of FVIII. However, this approach lacks regulation by thrombin cleavage. Therefore, it cannot be inactivated, depends on other steps of the coagulation cascade, and also has a high risk of thrombotic effects (References by Lenting et al., 2017).

[0019] Because this disease is monogenic, HA is an attractive target for gene therapy. Four vectors based on adeno-associated virus (AAV) serotypes that transduce hepatocytes are currently in clinical trials. This treatment has shown improvement in some patients. However, since one of the main side effects is an increase in liver enzymes, the safety of this approach still needs to be investigated (References by Pasi et al., 2017; George et al., 2017; Peters and Harris et al., 2018). Furthermore, the non-integrating nature of AAV limits the application of the treatment method in children due to dilution and loss of vector expression as a result of cell proliferation during organ growth (References by Vandamme et al., 2017).

[0020] The above-mentioned treatment method complements the defective FVIII gene function but does not correct the causative mutations. Reversing the gene to its normal orientation enables stable expression of FVIII under physiological conditions. Importantly, increasing the FVIII level to 1-5% of normal significantly reduces the risk of spontaneous internal bleeding, and administration of FVIII is only required in cases of trauma and surgery, thus significantly improving the lives of severe HA patients.

[0021] Jin-Soo Kim's group used programmable nucleases to correct two inversions that cause severe HA. They successfully inverted exon 1 at 0.2 - 0.4% (Lee et al., 2012) using ZFNs in HEK 293 cells and at 1.4% (Park et al., 2014) using TALENs in human iPSCs. Using CRISPR-Cas9 technology, the most efficient corrections (reversions) of 6.7% of exon 1 and 3.7% of exons 1 - 22 within the F8 gene in iPSCs derived from HA patients were achieved (Park et al., 2015). Although impressive, the results obtained using nucleases have limited therapeutic utility. When programmable nucleases perform corrections by introducing two double-strand breaks in the homologous region, either inversion or deletion of adjacent DNA fragments made by the cellular repair pathway occurs. In addition, indels at the cleavage site are highly likely to occur. Therefore, the inversions performed by ZFNs, TALENs, or CRISPR-Cas9 are more stochastic than controlled events.

[0022] To address this drawback, SSR can be used to correct the gene inversion that causes the disease. SSR is highly specific and can invert DNA sequences in the genomic context without depending on cofactors when the target sites of recombinases are present in the appropriate orientation (Yu and Bradley et al., 2002). Therefore, the goal was to develop an SSR system that can correct the int1h inversion in human cells with high efficiency and specificity. SUMMARY OF THE INVENTION

Problems to be Solved by the Invention

[0023] (Summary of the Invention) Accordingly, an object of the present invention is to provide a DNA recombinase that efficiently and specifically recombines genomic target sequences. In particular, the problem that each recombinase monomer must efficiently and specifically find and recognize its binding site must be solved.

Means for Solving the Problems

[0024] The problem of the present invention is solved by providing a method for producing a fusion protein or DNA recombinase for efficient and specific genome editing, particularly for recombination, most preferably for inversion of a DNA sequence at the genome level, in a cell comprising a complex of recombinases comprising at least a first recombinase enzyme, a second recombinase enzyme and at least one linker, wherein the first recombinase enzyme and the second recombinase enzyme specifically recognize a first half-site and a second half-site of an upstream target site and / or a downstream target site of the recombinase, the first recombinase enzyme and the second recombinase enzyme are bound to each other via a linker, and the linker comprises or consists of an oligopeptide.

[0025] The present invention further relates to a non-fusion heterodimer DNA recombinase without a linker peptide for use in recombination, most preferably for inversion of a DNA sequence at the genome level, in a cell.

[0026] According to the present invention, it is shown, for example, in FIGS. 19, 20 and 23 that recombinase monomers can be covalently bound together to one polypeptide without losing their ability to recombine their respective target sites. This is unexpected considering the details of the known molecules of the recombination reaction (stepwise cleavage, Holliday structure formation and isomerization). The current dogma was that the four recombinase molecules required to carry out the recombination reaction had to be flexible monomers so that they could function as an enzyme complex. It is shown herein that the fusion of SSRs via specific linkers can be applied to designer recombinases evolved to recombine sequences present in the inverted repeats that are causally involved in the genetic modification of the FVIII gene that causes HA in humans. Unexpectedly, the fusion of two heterospecific recombinases with a specific linker prevents the molecule from recombining the symmetric target site and forces the system. Furthermore, the fusion protein is highly active and exhibits significantly improved specificity compared to the unbound enzyme. Directed molecular evolution leads to improved linker sequences with desirable characteristics.

[0027] The present invention also discloses a target sequence (loxF8) in the int1h sequence involved in FVIII inversion required for the production of designer recombinases.

[0028] The problem of the present invention is solved, in particular, by providing a fusion protein for efficient and specific genome editing, comprising a complex of recombinases comprising at least a first recombinase enzyme, a second recombinase enzyme and at least one linker, wherein the first recombinase enzyme and the second recombinase enzyme specifically recognize the first half-site and the second half-site of the upstream target site and / or the downstream target site of the recombinase, the first recombinase enzyme and the second recombinase enzyme are covalently bound to each other via a linker, and the linker comprises or consists of an oligopeptide comprising 4 to 50 amino acids.

Advantages of the Invention

[0029] The present invention further relates to a DNA recombinase that specifically recognizes target sequences upstream and downstream of the loxF8 recombinase target site and catalyzes the inversion of the gene sequence between the target sequences upstream and downstream of the loxF8 recombinase target site.

[0030] The present invention further relates to a nucleic acid molecule encoding the fusion protein or DNA recombinase according to the present invention.

[0031] In a further embodiment, the present invention provides a mammalian cell, insect cell, plant cell or bacterial host cell comprising the nucleic acid molecule encoding the fusion protein or DNA recombinase according to the present invention.

[0032] The fusion protein or DNA recombinase or nucleic acid molecule according to the present invention can be used as a medicament and can therefore be included in a pharmaceutical composition, optionally in combination with one or more pharmaceutically acceptable diluents or carriers.

[0033] The fusion protein or DNA recombinase or pharmaceutical composition according to the present invention is suitable for the treatment of diseases that can be cured by genome editing, particularly for the treatment of hemophilia A.

[0034] In a further embodiment, a method for determining recombination at the genomic level in a host cell culture or patient comprising the fusion protein or DNA recombinase for efficient and specific genome editing according to the present invention is provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] (Brief Description of the Drawings)

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

DETAILED DESCRIPTION OF THE INVENTION

[0036] General Definitions As used herein, the terms “cell,” “cell line,” and “cell culture” are used interchangeably, and all such names include progeny. Thus, the terms “transformant” and “transformed cell” include the primary subject cell and cultures derived therefrom, regardless of the number of transfers. It is also understood that all progeny may not be precisely identical in DNA content due to deliberate or inadvertent mutations. Mutant progeny having the same function or biological activity as screened in the originally transformed cell are included. Where a different meaning is intended, this will be apparent from the context.

[0037] As used herein, the terms "polypeptide," "peptide," and "protein" are interchangeable and are defined to mean a biomolecule composed of amino acids joined by peptide bonds.

[0038] When a peptide or amino acid sequence is referred to herein, each amino acid residue is represented by a one-letter or three-letter name corresponding to the conventional name of the amino acid, according to the following conventional list: [Table 1]

[0039] As used herein, the terms "a," "an," and "the" are defined to mean "one or more" and include the plural unless the context is inappropriate.

[0040] As used herein, the term "subject" refers to an animal, preferably a mammal, most preferably a human, that is the subject of treatment, observation, or experiment.

[0041] As used herein, the term "therapeutically effective amount" means the amount of an active compound or pharmaceutical agent that elicits a biological or pharmaceutical response in a tissue system, animal, or human, as determined by a researcher, veterinarian, physician, or other clinician, including alleviation of symptoms of a disease or disorder being treated.

[0042] As used herein, the term "pharmaceutically acceptable" encompasses use in both human and veterinary medicine: for example, the term "pharmaceutically acceptable" includes compounds that are acceptable as veterinary pharmaceuticals or compounds that are acceptable in human medicine and healthcare.

[0043] Naturally occurring DNA recombinases, particularly site-specific recombinase (SSR) systems (such as tyrosine-type SSRs), generally consist of four identical monomers. Generally, they recognize two identical and symmetric palindromic target sites, each consisting of two approximately 13-nucleotide-long half-sites separated by an asymmetric, 8-nucleotide-long spacer with high frequency (Figure 1). Depending on the number and relative orientation of the target sites and their spacers, the DNA recombinase can perform any of excision, integration, inversion, or substitution of gene content (Figure 2; reviewed in Meinke et al., 2016).

[0044] As used herein, "upstream" refers to the 5' target site of the recombinase in DNA, including a first half-site such as a left half-site and a second half-site such as a right half-site, and the first half-site and the second half-site are separated by a spacer sequence.

[0045] As used herein, "downstream" refers to the 3' target site of the recombinase in DNA, a first half-site such as a left half-site, and a second half-site such as a right half-site, and the first half-site and the second half-site are separated by a spacer sequence.

[0046] (Detailed Description of the Invention) The present invention is a method for producing a DNA recombinase for efficient and specific genome editing in cells, particularly for recombination, and most preferably for inversion of DNA sequences at the genome level, comprising: i. Identifying a nucleic acid sequence that is a potential target site for a DNA recombinase capable of inducing site-specific DNA recombination of a sequence of interest in the genome, wherein the potential target site comprises two asymmetric recombinase target sites; ii. Providing a nucleic acid molecule encoding a first recombinase enzyme, wherein the first recombinase enzyme has been evolved by directed evolution or rational design to specifically recognize the first half-site of the recombinase target site; iii. providing a nucleic acid molecule encoding a second recombinase enzyme, wherein the second recombinase enzyme has been evolved by directed evolution or rational design to specifically recognize a second half-site of a recombinase target site; iv. producing an expression vector by cloning a nucleic acid molecule encoding a first recombinase enzyme and a nucleic acid molecule encoding a second recombinase enzyme into the expression vector; v. transfecting a cell containing a DNA sequence to be inverted with the expression vector of step iv); vi. expressing a DNA recombinase comprising a first recombinase enzyme and a second recombinase enzyme; vii. analyzing whether the DNA recombinase expressed in step vi. can invert a DNA sequence on a human chromosome in the cell; and viii. selecting a DNA recombinase capable of inverting a DNA sequence on a human chromosome in the cell according to step vii A method comprising is provided.

[0047] Surprisingly, it has been found that the specificity of the DNA recombinase obtained by this method can be dramatically increased by binding the recombinase monomer to an oligopeptide, thereby producing a fusion protein.

[0048] Accordingly, in a preferred embodiment, the present invention is a method for producing a fusion protein for efficient and specific genome editing in a cell, particularly for recombination, most preferably for inversion of a DNA sequence at the genome level, comprising i. identifying a nucleic acid sequence that is a potential target site for a DNA recombinase capable of inducing site-specific DNA recombination of a sequence of interest in the genome, wherein the potential target site comprises two asymmetric recombinase target sites; ii. providing a nucleic acid molecule encoding a first recombinase enzyme, wherein the first recombinase enzyme has been evolved by directed evolution or rational design to specifically recognize a first half-site of a recombinase target site; iii. providing a nucleic acid molecule encoding a second recombinase enzyme, wherein the second recombinase enzyme has been evolved by directed evolution or rational design to specifically recognize a second half-site of a recombinase target site; iv. providing a nucleic acid molecule encoding a linker peptide, wherein the nucleic acid encodes a linker oligopeptide comprising 6 to 30 amino acids; v. producing an expression vector by cloning the nucleic acid molecule encoding the first recombinase enzyme, the nucleic acid molecule encoding the second recombinase enzyme, and the nucleic acid molecule encoding the linker peptide into an expression vector; vi. transfecting a cell comprising a DNA sequence to be recombined, preferably inverted, with the expression vector of step v); vii. expressing a fusion protein comprising the first recombinase enzyme, the second recombinase enzyme, and the linker peptide; viii. analyzing whether the fusion protein expressed in step vii. can invert a DNA sequence on a human chromosome in the cell; and ix. selecting, according to step viii, a fusion protein capable of inverting a DNA sequence on a human chromosome in the cell A method comprising the above steps is provided.

[0049] The nucleic acid sequence according to step i., which is a potential target site for a DNA recombinase enzyme capable of inducing site-specific DNA recombination of a sequence of interest in the genome, is, for example, a) A sub-step of screening a genome or a part thereof containing a sequence of interest for two arrays that are potential spacer arrays having a length of at least 5 and at most 12 bp, wherein one of the potential spacer arrays is upstream of the sequence of interest and the other potential spacer array is downstream of the sequence of interest, and the two arrays have a maximum distance of 2 megabases or less, preferably 1.5 megabases or less or 1 megabase or less, more preferably 900 kb, 800 kb, 700 kb, 600 kb or 500 kb, most preferably 400 kb or 300 kb, and a minimum distance of 150 bp. b) A sub-step of identifying potential target sites by determining that both the potential half-sites and the spacer array therebetween form potential target sites, and for each potential spacer array, the neighboring nucleotides on one side that form the potential first half-site, preferably 10 - 20 nucleotides, more preferably 12 - 15 nucleotides, most preferably 13 nucleotides, and the neighboring nucleotides on the other side that form the potential second half-site, preferably 10 - 20 nucleotides, more preferably 12 - 15 nucleotides, most preferably 13 nucleotides. c) A sub-step of further screening the potential target sites identified in step b) to select potential target sequences that do not occur anywhere in the genome of the host to ensure sequence-specific recombination, preferably inversion. It can be identified according to the method described in WO 2018 / 229226 A1 including the above steps.

[0050] Preferably, the sequences of the potential target sites identified for the DNA recombinase are naturally present in the genome.

[0051] The first recombinase enzyme and the second recombinase enzyme according to steps ii. and iii. can be evolved by directed evolution or rational design, preferably, for example, by substrate-binding directed evolution (SLiDE) as described in WO 2018 / 229226 A1. The directed evolution results in obtaining at least one first designer DNA recombinase enzyme that is active on the first target site selected in a), and at least one second designer DNA recombinase enzyme that is active on the second target site, until a) selecting the nucleotide sequence upstream of the nucleotide sequence to be modified as the first target site, and selecting the nucleotide sequence downstream of the nucleotide sequence to be modified as the second target site. The sequences of the target sites are preferably not identical, but each target site includes a first half-site and a second half-site each having 10-20 nucleotides separated by a spacer sequence of 5-12 nucleotides, b) applying molecular directed evolution to at least one library of DNA recombinase enzymes using a vector containing the first target site and the second target site selected in a) as substrates is included.

[0052] The selection of asymmetric target sites provides an opportunity to compare two different evolutionary strategies. A single recombinase can be evolved to recognize both 10-20 bp, more preferably 12-15 bp, and most preferably 13 bp half-sites, or two recombinases can be evolved in parallel for each half-site. By combining two recombinases, a functional heterodimer that can recombine asymmetric sites can be formed.

[0053] In one embodiment of the present invention, it is preferred to evolve a single recombinase to recognize both 10-20 bp, more preferably 12-15 bp, and most preferably 13 bp half-sites.

[0054] In a further embodiment of the invention, it is preferred to evolve two recombinases in parallel for each half-site that gives rise to a heterodimer. Since the heterodimer consists of two recombinases that can form either a heterodimer or two different homodimers, the amount of potential recognition sequences increases. This approach can be disadvantageous as it increases the chance of unwanted recombination at off-target sites. A further goal of the invention was to limit the homodimerization of the monomers in order to reduce the chance of recombination at off-target sites. To achieve this goal, the recombinase monomers were physically fused and the desired heterodimer assembly was forced using a linker peptide. The linker peptide was produced by providing a nucleic acid molecule encoding the linker peptide, said nucleic acid encoding a linker oligopeptide comprising 6 to 30 amino acids according to step iv. of the method of the invention. In order to enhance the activity of the fused recombinase monomers, it is preferred to design a linker library. The linker library is suitably evolved in a number of cycles, for example 10 cycles of SLiDE, to find the most active linker variant. As a result, a final library showing robust activity against the desired recombinase target site, indicating that variants with improved activity have been produced according to steps viii. and ix. of the method of the invention, is produced and the fusion protein is selected.

[0055] In a further embodiment of the method of the invention, to select a fusion protein that is highly specific for the desired recombinase target site, i.e., a fusion protein showing reduced off-target recombination, x. identifying potential off-target sites of the desired fusion protein; xi. analyzing the recombinase activity of the desired fusion protein at the off-target sites identified in step x.; and xii. selecting a fusion protein that does not show recombinase activity at at least one off-target site may be included.

[0056] In a preferred embodiment of the present invention, the C-terminus of the second recombinase enzyme is linked to the N-terminus of the first recombinase enzyme via a linker peptide.

[0057] In the most preferred embodiment of the present invention, in the fusion protein expressed in step vii., the C-terminus of the first recombinase enzyme is linked to the N-terminus of the second recombinase enzyme via a linker peptide.

[0058] The off-target sites of recombinases can be identified, for example, using bioinformatics techniques known to those skilled in the art. Other techniques include ChIP-Seq-based assays for identifying putative off-targets in humans, followed by verification by qPCR and DNA enrichment. These methods are also known to those skilled in the art.

[0059] The recombinase activity of the fusion protein on these potential off-target sites can be experimentally tested, for example, by cloning a genomic sequence as a excision substrate into a bacterial reporter vector as described below in the present specification. Recombination at the off-target site can then be detected, for example, by monitoring the expression of the reporter gene using a PCR-based assay. Such assays can also be performed in human tissue cultures to examine whether the off-target site is altered by the fusion protein in vivo.

[0060] The method of the present invention has the particular advantage that it can evolve a recombinase fusion protein that exhibits specific activity against any recombinase target site. The method of the present invention includes providing a target site-specific recombinase complex, and since the linker for binding the monomers of the recombinase heterodimer is also specifically adapted to the evolved recombinase, the method of the present invention further has the advantage that it can dramatically reduce, preferably completely eliminate, the undesirable off-target activity of the recombinase complex, i.e., the recombinase fusion protein. Thereby, the recombinase complex, i.e., the recombinase fusion protein, is particularly suitable for use in gene therapy.

[0061] In another aspect, the present invention provides a fusion protein for efficient and specific genome editing comprising a complex of recombinases, said complex of recombinases comprising at least a first recombinase enzyme, a second recombinase enzyme and at least one linker, said first recombinase enzyme and said second recombinase enzyme specifically recognizing a first half-site and a second half-site of a recombinase target site; said first recombinase enzyme and said second recombinase enzyme being bound to each other via a linker; said linker comprising or consisting of an oligopeptide comprising 4 to 50 amino acids.

[0062] The recombinase contained in the complex is preferably a DNA recombinase and can be a naturally occurring recombinase (i.e., a recombinase isolated from any type of biological sample), a designer recombinase such as a recombinase evolved by directed molecular evolution or rational design, or any combination thereof. Methods for producing designer recombinases are known in the art. For example, WO 2018 / 229226 Al teaches vectors and methods for producing designer DNA recombinases by directed molecular evolution. W02008083931A1 discloses the directed molecular evolution of a recombinase (Tre 1.0) suitable for use with sequences in the terminal repeat sequences (LTR) of HIV as recognition sites (loxLTR Tre 1.0). Further development of this technique using asymmetric target sites is described in W02011147590 A2 (Tre 3.0) and W02016034553 A1 (Tre 3.1 and Tre / Brec1) and in the publication by Karpinski J et al., 2016 (Brec1). Methods for manipulating naturally occurring recombinases or designer recombinases by rational design are also known in the art (e.g., Abi-Ghanem et al., 2013; Karimova et al., 2016).

[0063] In the most preferred embodiment, the recombinase contained in the complex is produced by the above method of the present invention. Accordingly, the features and embodiments regarding the recombinase target site, fusion protein, recombinase, and linker peptide, which will be described later, are also applicable to the above method of the present invention.

[0064] Compositions of the recombinase complex are within the scope of the present invention and include, for example, · all four recombinase monomers are different; · three recombinase monomers are identical and one monomer is different; · two recombinase monomers are identical and the other two monomers are different; · the complex contains two different homodimers; · The complex contains two different heterodimers; · The complex contains two identical heterodimers; · All four recombinase monomers are identical; · The complex contains two different monomers; or · The complex contains two identical monomers.

[0065] In this context, "different" means that the monomers are not identical in their primary structure, i.e., they show differences in their amino acid sequences; and / or show high specificity for one of the four half-sites of the upstream and downstream target sites of the recombinase, which conveniently results in the surprisingly increased specificity of the fusion protein of the present invention.

[0066] In a preferred embodiment, the recombinase complex is a dimer, more preferably a heterodimer, for the recognition of the first and second target sequences of the upstream or downstream recombinase target sites in DNA, and the monomers of said heterodimer are fused via a linker.

[0067] In a more preferred embodiment, the recombinase complex of the present invention is a tetramer, more preferably a heterotetramer, for recognizing the upstream and downstream target sites of the recombinase in DNA, and at least two monomers of said heterotetramer are fused via a linker.

[0068] In a more preferred embodiment, the complex consists of two heterodimers, most preferably two identical heterodimers, and the monomers of each heterodimer are bonded to each other via a linker.

[0069] More preferably, the monomers of the heterodimer have evolved by directed evolution or rational design to specifically recognize the first or second half-site of the recombinase target site. Thus, the first heterodimer in the complex of the two heterodimers specifically recognizes the first and second half-sites of the upstream recombinase target site in the DNA, and the second heterodimer specifically recognizes the first and second half-sites of the downstream recombinase target site.

[0070] Even more preferably, the monomers of the heterodimer are tyrosine site-specific recombinases.

[0071] Most preferably, the first recombinase monomer of the heterodimer has the following specific amino acids: Q at position 5, A at position 40, T at position 44, V at position 80, R at position 90, L at position 94, and R at position 219; or Q at position 5, A at position 40, T at position 44, V at position 80, R at position 90, L at position 94, R at position 219, G at position 232, S at position 323 and L at position 325; or Q at position 5, A at position 40, T at position 44, V at position 80, R at position 90, L at position 94, R at position 132, G at position 150, R at position 219, V at position 272, S at position 323 and L at position 325 and has an amino acid sequence characterized by.

[0072] Even most preferably, the second recombinase monomer of the heterodimer has the following specific amino acids: H at position 90, S at position 94, K at position 249, G at position 266, V at position 272 and K at position 282; or Q at position 5, T at position 80, H at position 90, S at position 94, K at position 249, G at position 266, V at position 272 and K at position 282 and has an amino acid sequence characterized by.

[0073] These amino acid positions refer to the numbering of the sequence of the wild-type recombinase Cre of SEQ ID NO: 68.

[0074] To avoid ambiguity, when the recombinase complex is represented by a dimer, particularly a homodimer or preferably a heterodimer, one linker for the interaction of the two monomers of the dimer is present in the fusion protein of the present invention. When the recombinase complex is represented by a tetramer, one, two, three or four linkers for the interaction of the four monomers of the tetramer are present in the fusion protein of the present invention.

[0075] Each of the linkers included in the fusion protein of the present invention is suitably a peptide, preferably an oligopeptide. In a preferred embodiment, the linker consists of 4 to 50 amino acids, more preferably 5 to 40 amino acids, and most preferably 6 to 30 amino acids.

[0076] Based on the general knowledge regarding the mechanism of action of recombinases and the steric conditions of the mechanism of action of recombinases, it was surprising to those skilled in the art that the fusion of recombinases did not inhibit recombination. Surprisingly, it is experimentally shown herein that recombination events induced by the binding of a fusion protein consisting of a heterodimer and a linker can be observed at the recombinase target site, which indicates a forced mode of recombination events.

[0077] Therefore, each linker included in the fusion protein of the present invention, in one embodiment,

Chemical formula

[0078] However, the fusion of recombinase with a linker containing (G2S) repeats sometimes resulted in a decrease in recombination activity compared to the unbound recombinase. Further experiments have shown that the linker length plays a significant role in the activity of the recombinase. Good results were obtained using a linker containing 8 (G2S) repeats. Further increase in linker length did not improve the recombination efficiency, but it is undesirable that the specificity decreases as the resulting fusion protein shows activity on the symmetric site of the recombinase target site. Therefore, in a preferred embodiment, the linker contained in the fusion protein of the present invention is most preferably an oligopeptide consisting of 24 amino acids.

[0079] We further investigated whether the amino acid composition of the linker affects the recombination activity of the final target site. Three linker libraries were designed. The purpose of the library was to find linkers with the same high specificity (no recombination) on the symmetric site and improved activity on the final loxF8 recombinase target site. Part or all of the linker sequence containing 8 (G2S) repeats was changed to the degenerate codon RVM encoding 9 of the amino acids commonly used in natural linkers (Chen et al., 2013), namely Ala, Arg, Asn, Asp, Glu, Gly, Lys, Ser, Thr amino acids. As a result, a large number of linker variants were created.

[0080] In a more preferred embodiment, the fusion protein of the present invention thus has the following formulas 1, 2 and 3 representing the above-mentioned tree linker library: X1 - X2 - X3 - X4 - X5 - X6 - (G2S)4 - X7 - X8 - X9 - X 10 - X 11 - X 12 (Formula 1); (G2S)2 - X1 - X2 - X3 - X4 - X5 - X6 - X7 - X8 - X9 - X 10 - X 11 -X 12 - (G2S)2 (Formula 2); and X1 - X2 - X3 - X4 - X5 - X6 - X7 - X8 - X9 - X 10 - X 11 -X 12 - X 13 - X 14 - X 15 - X 16 - X 17 -X 18 - X 19 - X 20 - X 21 - X 22 - X 23 -X 24 (Formula 3); (wherein, G is glycine; S is serine; and each of X1 to X 24 is independently selected from the group consisting of alanine, arginine, asparagine, aspartic acid, glutamine, glycine, lysine, serine and threonine) comprises an oligopeptide having an amino acid sequence selected from or consisting of a linker.

[0081] In a preferred embodiment, the oligopeptide of Formula 1, Formula 2 or Formula 3 does not consist only of glycine and serine residues.

[0082] With respect to the recombinase used in the recombinase complex, the fusion protein according to the present invention preferably comprises a heterodimeric complex, and each recombinase enzyme of the heterodimer is directed to independently recognize the half-site of the upstream target site and the half-site of the downstream target site of the tyrosine site-specific recombinase. It is a tyrosine site-specific recombinase evolved by evolution or produced by other means.

[0083] Suitably, the tyrosine site-specific recombinase is selected from the group consisting of Cre, Dre-, VCre-, SCre-, Vika-, lambda-Int-, Flp-, R-, Kw-, Kd-, B2-, B3-, Nigri- and Panto-recombinases. The recognition target sites of these bacterial and yeast T-SSR systems are discussed in Meinke et al., 2016 and Karimova et al., 2016 and are shown in Table 1 below: Table 1. Recognition target sites of bacterial and yeast T-SSR systems

Table 2

[0084] As described above, most of the severe cases of hemophilia A in patients are caused by inversion of exon 1 due to homologous recombination between two 1 kb sequences in the region located telomeric to intron 1 and the F8 gene, i.e., by inversion due to intrachromosomal recombination with homologous regions outside the F8 gene (Graw et al., 2005) (Figure 4).

[0085] It was a further object of the present invention to provide a fusion protein that specifically corrects DNA inversions that cause severe HA, preferably inversion of exon 1. To solve this problem, two loxF8 target sites were identified in the homologous regions located telomeric to intron 1 and gene F8. Subsequently, two SSRs that recognize different half-sites of loxF8 were evolved by directed molecular evolution (Figure 5A). Each of the recombinases so evolved recognizes one half-site of the final target site and performs recombination cooperatively.

[0086] By altering the DNA binding properties of naturally occurring SSRs, these systems can be used for other purposes to recombine DNA sequences with therapeutic value. Methods for creating designer recombinases are known in the art. For example, WO 2018 / 229226 Al teaches vectors and methods for creating designer DNA recombinases by directed molecular evolution. W02008083931A1 discloses the directed molecular evolution of a recombinase (Tre 1.0) suitable for use of a sequence in the HIV long terminal repeat (LTR) as a recognition site (loxLTR Tre 1.0). Further development of this approach using asymmetric target sites is described in W02011147590 A2 (Tre 3.0) and W02016034553 A1 (Tre 3.1 and Tre / Brec1) and in the publication by Karpinski J et al. 2016 (Brec1). Methods for engineering naturally occurring recombinases or designer recombinases by rational design are also known in the art (e.g., Abi-Ghanem et al., 2013; Karimova et al., 2016).

[0087] Instead of creating a single designer recombinase that targets an asymmetric target site, a dual designer recombinase that recombinantly joins target sites together as a heterodimer can also be created (Lansing et al., 2020). However, a comparative study of a single designer recombinase system and a dual designer recombinase system (heterodimer) that target the same DNA target has not yet been performed. To conduct such a comparative study, two classes of recombinases that utilize the SLiDE technology (Buchholz and Stewart et al., 2001; Lansing et al., 2020) were created (loxF8, Figures 4 and 5) for the purpose of recombining conserved sequences in the int1h1 and int1h2 sequences.

[0088] To obtain a single recombinase that recognizes the target array as a monomer, SLiDE was performed at the asymmetric loxLTR and loxBRT sequences for Tre and at the asymmetric loxF8 site corresponding to the directed evolution of Brec1, respectively (Sarkar et al., 2007; Hauber et al., 2013). After 168 rounds of directed evolution, individual recombinases in the pEVO vector were tested in E. coli using different L-arabinose concentrations. The best clone identified (H7) recombined the loxF8 sequence when recombinase expression was induced with low concentrations of L-arabinose (Figure 6).

[0089] To create a recombinase system that recognizes the same target array as a heterodimer, first, SLiDE was performed on the symmetric loxF8 half-site corresponding to the directed evolution of the Hex recombinase, and the sequence on human chromosome 7 was recombined (Lansing et al., 2020). Then, a library of recombinases with activity on the symmetric loxF8 half-site (after 88 and 89 rounds of directed evolution, respectively) was co-expressed in E. coli from a pEVO plasmid carrying two complete asymmetric loxF8 target sites. After an additional 3 rounds of directed evolution, different L-arabinose concentrations were used to test the dual recombinase heterodimer in the pEVO vector in E. coli. The best heterodimer clone identified (D7) recombined the loxF8 sequence more efficiently than the single H7 clone (Figure 6), indicating that the production of the dual recombinase heterodimer is more effective than the production of a single recombinase targeting an asymmetric site.

[0090] To examine the recombination activity of the monomeric (H7) and heterodimeric (D7) F8 recombinases in mammalian cells, HeLa cells were co-transfected with a recombinant reporter and a recombinase expression plasmid (Figure 7A). Similar to the E. coli assay, the H7 recombinase was less efficient than the heterodimeric D7 recombinase (Figure 7B; recombination efficiency 28% vs. recombination efficiency 46%).

[0091] To compare the ability of the single F8 recombinase (H7) and the heterodimeric F8 recombinase (D7) to induce F8 inversion in human cells, the recombinase expression plasmids were transfected into HeLa cells. Genomic DNA was isolated 48 hours after transfection. A PCR reaction was performed to detect genomic inversion. Again, the monomeric H7 recombinase did not function as well as the heterodimeric D7 recombinase in this assay (Figure 8).

[0092] To compare the specificity of the H7 monomer and the D7 heterodimer, the recombinases were tested on nine human sequences that showed the highest similarity to the loxF8 sequence (SEQ ID NOs: 21-23 and 87-92, Table 2). The nine off-target sites and the loxF8 on-target site were cloned as excision substrates into pEVO vectors carrying the respective recombinases, and the plasmids were grown in E. coli for 24 hours in the presence of 10 μg / ml L-arabinose. DNA extracted from these cultures revealed that the H7 monomer recombined the loxF8 site (SEQ ID NO: 17) in addition to off-target sites having SEQ ID NOs: 21, 22, 87, and 91. In contrast, the D7 heterodimer was more efficient and far more specific in the recombination of the loxF8 sequence, only slightly recombining the off-target site with SEQ ID NO: 22 (Figure 9). Based on these cumulative results, the heterodimer D7 was selected for further experiments.

[0093] The D7 heterodimer showed superior properties compared to the H7 monomer, but still showed some off-target recombination. In therapeutic applications, it is desirable to employ genome editing tools that are highly specific and do not edit the genome at sites other than the intended site. In previous attempts, only limited success has been reported in improving the application properties of heterospecific recombinases, so ways to improve the behavior of double recombinases were investigated. One possibility was to physically link two D7 heterodimers by fusing them with a peptide linker. Recent literature suggested that this would probably not work because the individual enzymes need to move quite a bit during the recombination reaction. However, if successful, this approach could potentially provide a simple way to create a very efficient and specific forced recombination system.

[0094] In a further embodiment of the present invention, the recombinase heterodimer D11 was evolved in the same manner as the D7 heterodimer.

[0095] In a more preferred embodiment of the present invention, the recombinase heterodimer A4-L was evolved in the same manner as the D7 heterodimer. To obtain a fusion heterodimer with the best linker properties, selection was made for activity on the loxF8 site and counterselection was made for activity on both symmetric sites (Figure 10). The combined recombinases could theoretically form different protein complexes with activity on symmetric or asymmetric target sites (Figure 11). All of the first three libraries showed activity on the symmetric site, indicating that the linker composition has an important influence on specificity. Fusion proteins using linkers with amino acid sequences selected from SEQ ID NOs: 8-16 showed satisfactory activity on the loxF8 site.

[0096] Therefore, in a further embodiment, the present invention provides a fusion protein, wherein the linker is

Chemical formula

[0097] A further important criterion for linker selection was that it was active on the loxF8 site at low induction levels and not active on the symmetrical site of loxF8 at high induction levels (Figures 16, 17). The best results in this regard were achieved using a fusion protein containing a linker having the amino acid sequence

Chemical formula

[0098] Accordingly, in a preferred embodiment, the present invention provides a fusion protein, wherein the linker comprises or consists of an oligopeptide having the amino acid sequence of SEQ ID NO: 14.

[0099] More preferably, the fusion protein of the present invention specifically recognizes an upstream recombinase target sequence of a loxF8 target site having a nucleic acid sequence

Chemical formula

Chemical formula

[0100] The ability of the fusion protein to catalyze the inversion of the DNA sequence between the upstream recombinase target sequence of SEQ ID NO: 17 and the downstream recombinase target sequence of SEQ ID NO: 18 of the loxF8 recombinase target site can be a) expressing a fusion protein comprising a first recombinase enzyme, a second recombinase enzyme and a linker peptide; and b) analyzing whether the fusion protein expressed in step a) can invert the DNA sequence on the human chromosome in the cell can be tested by a method comprising.

[0101] Most preferably, the fusion protein of the present invention exhibits high specificity for the loxF8 target site having the target sequences of SEQ ID NOs: 17 and 18 and shows no activity against off-target sites at high induction levels (Figure 20). Preferably, the off-target sites not recognized by the fusion protein of the present invention are selected from the group consisting of SEQ ID NOs: 19-29 and 87-92 shown in Table 2.

[0102] Table 2: Nucleic acid sequences of off-target sites

Table 3

[0103] The upstream (5') recombinase target sequence of the loxF8 target site having the nucleic acid sequence of SEQ ID NO: 17; and the downstream (3') recombinase target sequence of the loxF8 target site having the nucleic acid sequence of SEQ ID NO: 18 were identified as part of the present invention. Thus, in a more preferred embodiment, the present invention relates to a loxF8 recombinase target site comprising the 5' target sequence of SEQ ID NO: 17 and the 3' target sequence of SEQ ID NO: 17.

[0104] To generate a DNA recombinase that recombines the target sequences upstream and downstream of the loxF8 target site, a substrate-binding directed evolution approach was employed (Buchholz and Stewart et al., 2001).

[0105] The DNA recombinase according to the present invention is an enzyme that recombines nucleic acid sequences, particularly DNA sequences, by recognizing two target sites (recognition sites, i.e., one upstream and one downstream recognition site or sequence), causing deletion, insertion, inversion, or substitution of the DNA sequence. Conveniently, the DNA recombinase according to the present invention recognizes the asymmetric recognition sites of the loxF8 sequences set forth in SEQ ID NO: 17 (upstream) and SEQ ID NO: 18 (downstream). Since these recognition sites do not exist anywhere else in the human genome, they can be used for specific DNA recombination. The DNA recombinase according to the present invention conveniently does not require target sites artificially introduced into the genome. Even more conveniently and most preferably, the DNA recombinase according to the present invention causes inversion of the DNA sequence. A further advantage is that the DNA recombinase according to the present invention enables accurate genome editing without initiating the endogenous DNA repair pathway.

[0106] After several cycles of evolution, an enzyme having activity against the target sequences of SEQ ID NO: 17 and SEQ ID NO: 18 was produced.

[0107] This problem was first solved by a recombinase monomer (recombinase H7) having an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 66. Thus, in a preferred embodiment, the present invention provides a recombinase monomer having an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 66.

[0108] In a preferred embodiment, the DNA recombinase of SEQ ID NO: 66 contains one or more mutations compared to the Cre recombinase protein of SEQ ID NO: 68. Exemplary mutations include L5Q, V7L, P12S, P15L, V16A, V23T, M30V, F31L, R34S, H40Q, M44S, S47F, K57E, K62E, Y77H, Q90N, Q94S, S108G, T140A, D143S, Q144R, S147A, C155P, I166V, A175S, A231V, K244R, N245G, A249K, R259C, E262Q, E266A, T268A, I272L, S305P, P307V, E308Q, N317T, N319E, and I320S. More preferably, the recombinase of SEQ ID NO: 66 has one or more, preferably 2, 3, 4, 5, 6, 8, 9, 10 or more mutations selected from the group consisting of L5Q, V7L, P12S, P15L, V16A, V23T, M30V, F31L, R34S, H40Q, M44S, S47F, K57E, K62E, Y77H, Q90N, Q94S, S108G, T140A, D143S, Q144R, S147A, C155P, I166V, A175S, A231V, K244R, N245G, A249K, R259C, E262Q, E266A, T268A, I272L, S305P, P307V, E308Q, N317T, N319E, and I320S compared to the Cre recombinase protein of SEQ ID NO: 68. In the most preferred embodiment, the recombinase of SEQ ID NO: 66 has all of the aforementioned mutations.

[0109] To avoid ambiguity, the amino acid denoted by the one-letter code before the position number represents the amino acid in the wild-type sequence (e.g., SEQ ID NO: 68 of Cre recombinase); and the amino acid denoted by the one-letter code after the position number represents the amino acid in the mutated, i.e., evolved, recombinase of the present invention.

[0110] More preferably, sequences having at least 70%, preferably 80%, more preferably 90% sequence identity with the recombinase of SEQ ID NO: 66 are characterized by the following specific amino acids: Q at position 5, Q at position 40, S at position 44, N at position 90, S at position 94, S at position 143, R at position 144, P at position 155, V at position 231, K at position 249, and L at position 272. These amino acid residues are predicted to be important residues for target site recognition based on deep sequencing and are not found in the sequence of the wild-type Cre recombinase of SEQ ID NO: 68.

[0111] Since the recombination efficiency and specificity of this recombinase monomer were not satisfactory, DNA recombinase heterodimers were also developed after several cycles of evolution.

[0112] As a result, the problem of the present invention is further solved by a DNA recombinase composed of a heterodimer, said heterodimer comprising or consisting of a first recombinase enzyme and a second recombinase enzyme, said first recombinase enzyme having an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 30 (recombinase 28-L), and said second recombinase enzyme having an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 31 (recombinase 28-R).

[0113] Preferably, the DNA recombinase of SEQ ID NO: 30 contains one or more mutations as compared to the Cre recombinase protein of SEQ ID NO: 68. Exemplary mutations include N3D, L5Q, V7L, P12S, P15L, V23A, M30V, H40A, K43N, M44T, L58S, K62E, Y77H, A80V, K86N, Q90R, G93A, Q94L, S108G, A131V, K132R, E150G, N151S, C155R, Q156N, I166V, A175S, V182I, I195V, K219R, D232G, T253S, S257T, R259D, A260V, E262R, E266V, T268A, I272V, Y273H, K276R, A285R, P307A, N317T, N319E, I320S, N323S, and I325L. More preferably, the recombinase of SEQ ID NO: 30 has one or more, preferably 2, 3, 4, 5, 6, 8, 9, 10 or more mutations selected from the group consisting of N3D, L5Q, V7L, P12S, P15L, V23A, M30V, H40A, K43N, M44T, L58S, K62E, Y77H, A80V, K86N, Q90R, G93A, Q94L, S108G, A131V, K132R, E150G, N151S, C155R, Q156N, I166V, A175S, V182I, I195V, K219R, D232G, T253S, S257T, R259D, A260V, E262R, E266V, T268A, I272V, Y273H, K276R, A285R, P307A, N317T, N319E, I320S, N323S, and I325L as compared to the Cre recombinase protein of SEQ ID NO: 68. In the most preferred embodiment, the recombinase of SEQ ID NO: 30 has all of the aforementioned mutations.

[0114] More preferably, the recombinase of SEQ ID NO: 30 exhibits a sequence characterized by the following specific amino acids: Q at position 5, A at position 40, T at position 44, V at position 80, R at position 90, L at position 94, R at position 155, R at position 219, V at position 266, V at position 272, S at position 323, and L at position 325. In a more preferred embodiment, the recombinase of SEQ ID NO: 30 further exhibits a sequence characterized by the following specific amino acids: arginine at position 132, glycine at position 150, glycine at position 232, and arginine at position 276. These amino acid residues are predicted to be important residues for target site recognition based on deep sequencing and are not found in the sequence of the wild-type Cre recombinase of SEQ ID NO: 68.

[0115] Preferably, the DNA recombinase of SEQ ID NO: 31 contains one or more mutations compared to the Cre recombinase protein of SEQ ID NO: 68. Exemplary mutations include L5Q, T6A, V7L, L11P, P12S, A13T, P15L, V16A, S20C, R34S, M44A, L46Q, A53V, N60S, K62R, Y77H, A80T, Q90H, Q94S, R101Q, S108G, K122R, T140A, F142L, S147A, I166V, A175G, K183R, N235D, K244R, N245Y, A249K, R259Y, E262Q, E266G, T268A, I272V, R282K, I306L, N317T, I320S, and G342D. More preferably, the recombinase of SEQ ID NO: 31 has one or more, preferably 2, 3, 4, 5, 6, 8, 9, 10 or more mutations selected from the group consisting of L5Q, T6A, V7L, L11P, P12S, A13T, P15L, V16A, S20C, R34S, M44A, L46Q, A53V, N60S, K62R, Y77H, A80T, Q90H, Q94S, R101Q, S108G, K122R, T140A, F142L, S147A, I166V, A175G, K183R, N235D, K244R, N245Y, A249K, R259Y, E262Q, E266G, T268A, I272V, R282K, I306L, N317T, I320S, G342D compared to the Cre recombinase protein of SEQ ID NO: 68. In the most preferred embodiment, the recombinase of SEQ ID NO: 31 has all of the aforementioned mutations.

[0116] More preferably, sequences having at least 70%, preferably 80%, more preferably 90% sequence identity with the recombinase of SEQ ID NO: 31 are characterized by the following specific amino acids: Q at position 5, A at position 44, T at position 80, H at position 90, S at position 94, K at position 249, G at position 266, V at position 272 and K at position 282. In a further preferred embodiment, the recombinase of SEQ ID NO: 31 further exhibits a sequence characterized by the amino acid arginine at position 183. These amino acid residues are predicted to be important residues for target site recognition based on deep sequencing and are not found in the sequence of the wild-type Cre recombinase of SEQ ID NO: 68.

[0117] Recombinant protein expression using bacteria and other host organisms is a fundamental technology for protein production. An important step in recombinant protein expression is codon optimization, in which the coding sequence of the protein of interest is designed by synonymous substitution aimed at increasing its expression level. In conventional methods, rare codons are replaced by frequent codons according to the genomic codon usage frequency in the host organism. The basis of this method is that endogenous genes consisting of frequent codons have high protein expression levels, so it is considered that recombinant protein expression can also be improved by increasing the codon frequency. Another method is to introduce synonymous substitutions that are computationally predicted to destabilize the mRNA secondary structure. Since a stable mRNA secondary structure may inhibit translation, this method is considered to improve recombinant protein expression by enhancing translation efficiency. The relationship between these sequence features and protein expression levels has been shown by omics analysis of endogenous genes. Direct evidence of their effects in recombinant protein expression has been shown in the art using a relatively small number of genes.

[0118] After performing additional cycles of molecular directed evolution, a DNA recombinase with improved enzyme activity was developed.

[0119] Accordingly, in a further embodiment, the present invention provides a heterodimeric DNA recombinase, wherein the first monomer is a recombinase enzyme (recombinase D7-L) having a sequence with at least 70%, preferably 80%, more preferably 90% sequence identity to the sequence set forth in SEQ ID NO: 32, and the second monomer is a recombinase enzyme (recombinase D7-R) having a sequence with at least 70%, preferably 80%, more preferably 90% sequence identity to the sequence set forth in SEQ ID NO: 33.

[0120] Preferably, the DNA recombinase of SEQ ID NO: 32 contains one or more mutations compared to the Cre recombinase protein of SEQ ID NO: 68. Exemplary mutations include L5Q, V7L, P12S, P15L, V16A, D17N, V23A, M30V, Q35R, H40A, M44T, S51T, Y77H, A80V, K86N, Q90R, G93A, Q94L, S108G, N111S, A131V, K132R, Q144K, E150G, I166V, A175S, V182I, K219R, E222G, D232G, R259D, A260V, E262R, I264V, E266A, T268A, I272V, A275T, R282G, A285T, P307A, N317T, N319E, I320S, N323S, and I325L. In a more preferred embodiment, the recombinase of SEQ ID NO: 32 has one or more, preferably 2, 3, 4, 5, 6, 8, 9, 10 or more mutations selected from the group consisting of L5Q, V7L, P12S, P15L, V16A, D17N, V23A, M30V, Q35R, H40A, M44T, S51T, Y77H, A80V, K86N, Q90R, G93A, Q94L, S108G, N111S, A131V, K132R, Q144K, E150G, I166V, A175S, V182I, K219R, E222G, D232G, R259D, A260V, E262R, I264V, E266A, T268A, I272V, A275T, R282G, A285T, P307A, N317T, N319E, I320S, N323S, and I325L compared to the Cre recombinase protein of SEQ ID NO: 68. In the most preferred embodiment, the recombinase of SEQ ID NO: 32 has all of the aforementioned mutations.

[0121] Even more preferably, a sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the recombinase of SEQ ID NO: 32 is characterized by the following specific amino acids: Q at position 5, A at position 40, T at position 44, V at position 80, R at position 90, L at position 94, K at position 144, R at position 219, V at position 272, G at position 282, S at position 323, and L at position 325. In a further preferred embodiment, the recombinase of SEQ ID NO: 32 further exhibits a sequence characterized by the following specific amino acids: arginine at position 132, glycine at position 150, and glycine at position 232. These amino acid residues are predicted to be important residues for target site recognition based on deep sequencing and are not found in the sequence of the wild-type Cre recombinase of SEQ ID NO: 68.

[0122] Preferably, the DNA recombinase of SEQ ID NO: 33 contains one or more mutations as compared to the Cre recombinase protein of SEQ ID NO: 68. Exemplary mutations include L4I, L5Q, V7P, N10S, P12S, P15L, V16T, E22V, V23T, M28A, R34S, K57E, F64L, A66V, Y77H, A80T, Q90H, Q94S, S102A, S108G, K122R, K132Q, M149V, I166V, A175S, K244R, N245Y, A249K, R259Y, E262Q, E266G, T268A, I272V, D277G, R282K, S305P, N317T, I320S, and G342S. In a more preferred embodiment, the recombinase of SEQ ID NO: 33 has one or more, preferably 2, 3, 4, 5, 6, 8, 9, 10 or more mutations selected from the group consisting of L4I, L5Q, V7P, N10S, P12S, P15L, V16T, E22V, V23T, M28A, R34S, K57E, F64L, A66V, Y77H, A80T, Q90H, Q94S, S102A, S108G, K122R, K132Q, M149V, I166V, A175S, K244R, N245Y, A249K, R259Y, E262Q, E266G, T268A, I272V, D277G, R282K, S305P, N317T, I320S, and G342S as compared to the Cre recombinase protein of SEQ ID NO: 68. In the most preferred embodiment, the recombinase of SEQ ID NO: 33 has all of the aforementioned mutations.

[0123] Even more preferably, a sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the recombinase of SEQ ID NO: 33 is characterized by the following specific amino acids: Q at position 5, T at position 80, H at position 90, S at position 94, K at position 249, G at position 266, V at position 272, and K at position 282. In a further preferred embodiment, the recombinase of SEQ ID NO: 33 further exhibits a sequence characterized by the specific amino acid glutamine at position 132. These amino acid residues are predicted to be important residues for target site recognition based on deep sequencing and are not found in the sequence of the wild-type Cre recombinase of SEQ ID NO: 68.

[0124] Preferably, the DNA recombinase monomer of SEQ ID NO: 32 has a glutamine at the 5th amino acid residue, an asparagine at the 17th amino acid residue, an arginine at the 35th amino acid residue, an alanine at the 40th amino acid residue, a threonine at the 44th amino acid residue, a valine at the 80th amino acid residue, an arginine at the 90th amino acid residue, a leucine at the 94th amino acid residue, a serine at the 111th amino acid residue, an arginine at the 132th amino acid residue, a lysine at the 144th amino acid residue, a glycine at the 150th amino acid residue, an arginine at the 219th amino acid residue, a glycine at the 222th amino acid residue, a valine at the 264th amino acid residue, a valine at the 272th amino acid residue, a threonine at the 275th amino acid residue, a glycine at the 282th amino acid residue, a serine at the 323th amino acid residue, and / or a leucine at the 325th amino acid residue.

[0125] These amino acid residues are specific to the DNA recombinase according to the present invention and are not found in Cre recombinase (SEQ ID NO: 68).

[0126] Preferably, the DNA recombinase monomer of SEQ ID NO: 33 has an isoleucine at the 4th amino acid residue, a glutamine at the 5th amino acid residue, a proline at the 7th amino acid residue, a threonine at the 16th amino acid residue, a valine at the 22nd amino acid residue, a threonine at the 23rd amino acid residue, an alanine at the 28th amino acid residue, a leucine at the 64th amino acid residue, a valine at the 66th amino acid residue, a threonine at the 80th amino acid residue, a histidine at the 90th amino acid residue, a serine at the 94th amino acid residue, an alanine at the 102nd amino acid residue, a glutamine at the 132nd amino acid residue, a valine at the 149th amino acid residue, a lysine at the 249th amino acid residue, a glycine at the 266th amino acid residue, a valine at the 272nd amino acid residue, a glycine at the 277th amino acid residue, a lysine at the 282nd amino acid residue, a proline at the 305th amino acid residue, and / or a serine at the 342nd amino acid residue.

[0127] These amino acid residues are specific to the DNA recombinase according to the present invention and are not found in Cre recombinase (SEQ ID NO: 68).

[0128] In a more preferred embodiment, the present invention provides a DNA recombinase that is a heterodimer, wherein the first monomer is a recombinase enzyme (recombinase A4-L) having a sequence with at least 70%, preferably 80%, more preferably 90% sequence identity to the sequence set forth in SEQ ID NO: 93, and the second monomer is a recombinase enzyme (recombinase A4-R) having a sequence with at least 70%, preferably 80%, more preferably 90% sequence identity to the sequence set forth in SEQ ID NO: 94.

[0129] Preferably, the DNA recombinase of SEQ ID NO: 93 contains one or more mutations compared to the Cre recombinase protein of SEQ ID NO: 68. Exemplary mutations include N3S, L5Q, V7L, P12S, P15L, V23A, K25E, M28I, D29G, M30G, H40A, M44T, N60S, Y77H, A80V, K86N, Q90R, G93A, Q94L, S108G, A131V, Q144R, I166V, I174V, A175S, K211E, K219R, D232G, N257T, R259D, A260V, E262R, E266V, T268A, K276R, P307A, N317T, N319E, I320S, N323S and I325L. In a more preferred embodiment, the recombinase of SEQ ID NO: 93 has one or more, preferably 2, 3, 4, 5, 6, 8, 9, 10 or more mutations selected from the group consisting of N3S, L5Q, V7L, P12S, P15L, V23A, K25E, M28I, D29G, M30G, H40A, M44T, N60S, Y77H, A80V, K86N, Q90R, G93A, Q94L, S108G, A131V, Q144R, I166V, I174V, A175S, K211E, K219R, D232G, N257T, R259D, A260V, E262R, E266V, T268A, K276R, P307A, N317T, N319E, I320S, N323S and I325L compared to the Cre recombinase protein of SEQ ID NO: 68. In the most preferred embodiment, the recombinase of SEQ ID NO: 93 has all of the aforementioned mutations.

[0130] Even more preferably, sequences having at least 70%, preferably 80%, more preferably 90% sequence identity to the recombinase of SEQ ID NO: 93 are characterized by the following specific amino acids: Q at position 5, A at position 40, T at position 44, V at position 80, R at position 90, L at position 94, R at position 144, R at position 219, G at position 232, V at position 266, R at position 276, S at position 323 and L at position 325. These amino acid residues are predicted to be important residues for target site recognition based on deep sequencing and are not found in the sequence of the wild-type Cre recombinase of SEQ ID NO: 68.

[0131] Preferably, the DNA recombinase of SEQ ID NO: 94 contains one or more mutations as compared to the Cre recombinase protein of SEQ ID NO: 68. Exemplary mutations include N3S, L5P, V7L, N10S, P12S, P15L, T19A, D21G, K25T, D29V, R34S, E39V, K57E, E67D, Y77H, Q90H, Q94S, N96D, S102A, S108G, K122R, E123A, K132G, I166V, A175S, K203R, K244R, N245Y, A249K, R259Y, E262Q, E266G, T268A, I272V, A275V, R282K, V304A, I306L, N317T, I320S, and G342S. In a more preferred embodiment, the recombinase of SEQ ID NO: 94 has one or more, preferably 2, 3, 4, 5, 6, 8, 9, 10 or more mutations selected from the group consisting of N3S, L5P, V7L, N10S, P12S, P15L, T19A, D21G, K25T, D29V, R34S, E39V, K57E, E67D, Y77H, Q90H, Q94S, N96D, S102A, S108G, K122R, E123A, K132G, I166V, A175S, K203R, K244R, N245Y, A249K, R259Y, E262Q, E266G, T268A, I272V, A275V, R282K, V304A, I306L, N317T, I320S, and G342S as compared to the Cre recombinase protein of SEQ ID NO: 68. In the most preferred embodiment, the recombinase of SEQ ID NO: 94 has all of the aforementioned mutations.

[0132] Even more preferably, a sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the recombinase of SEQ ID NO: 94 is characterized by the following specific amino acids: P at position 5, H at position 90, S at position 94, G at position 132, R at position 183, K at position 249, G at position 266, V at position 272, and K at position 282. These amino acid residues are predicted to be important residues for target site recognition based on deep sequencing and are not found in the sequence of the wild-type Cre recombinase of SEQ ID NO: 68.

[0133] Preferably, the DNA recombinase monomer of SEQ ID NO: 93 has a serine at the 3rd amino acid residue, a glutamine at the 5th amino acid residue, a glutamic acid at the 25th amino acid residue, an isoleucine at the 28th amino acid residue, a glycine at the 29th amino acid residue, a glycine at the 30th amino acid residue, an alanine at the 40th amino acid residue, a threonine at the 44th amino acid residue, a serine at the 60th amino acid residue, a valine at the 80th amino acid residue, an arginine at the 90th amino acid residue, a leucine at the 94th amino acid residue, an arginine at the 144th amino acid residue, a glutamic acid at the 211th amino acid residue, an arginine at the 219th amino acid residue, a glycine at the 232th amino acid residue, a threonine at the 257th amino acid residue, a valine at the 266th amino acid residue, an arginine at the 276th amino acid residue, a serine at the 323th amino acid residue, and / or a leucine at the 325th amino acid residue.

[0134] These amino acid residues are specific to the DNA recombinase according to the present invention and are not found in Cre recombinase (SEQ ID NO: 68).

[0135] Preferably, the DNA recombinase monomer of SEQ ID NO: 94 has a serine at the 3rd amino acid residue, a proline at the 5th amino acid residue, a glycine at the 21st amino acid residue, a threonine at the 25th amino acid residue, a valine at the 29th amino acid residue, a valine at the 39th amino acid residue, an aspartic acid at the 67th amino acid residue, a histidine at the 90th amino acid residue, a serine at the 94th amino acid residue, an aspartic acid at the 96th amino acid residue, an alanine at the 102nd amino acid residue, an alanine at the 123rd amino acid residue, a glycine at the 132nd amino acid residue, an arginine at the 183rd amino acid residue, a lysine at the 249th amino acid residue, a glycine at the 266th amino acid residue, a valine at the 272nd amino acid residue, a valine at the 275th amino acid residue, a lysine at the 282nd amino acid residue, an alanine at the 304th amino acid residue, a leucine at the 306th amino acid residue, and / or a serine at the 342nd amino acid residue.

[0136] These amino acid residues are specific to the DNA recombinase according to the present invention and are not found in Cre recombinase (SEQ ID NO: 68).

[0137] In a more preferred embodiment, the present invention provides a DNA recombinase that is a heterodimer, wherein the first monomer is a recombinase enzyme (recombinase D11-L) having a sequence with at least 70%, preferably 80%, more preferably 90% sequence identity to the sequence set forth in SEQ ID NO: 99, and the second monomer is a recombinase enzyme (recombinase D11-R) having a sequence with at least 70%, preferably 80%, more preferably 90% sequence identity to the sequence set forth in SEQ ID NO: 100.

[0138] Preferably, the DNA recombinase of SEQ ID NO: 99 contains one or more mutations compared to the Cre recombinase protein of SEQ ID NO: 68. Exemplary mutations include L5Q, V7I, P12T, L14S, P15L, V16A, V23A, M30V, F31L, H40A, M44T, S51T, L58S, Y77H, A80V, K86N, Q90R, G93A, Q94L, S108G, A131V, K132R, F142L, E150G, N151D, Q156K, L164P, I166V, A175S, V182I, K219R, A249V, R259D, A260V, E262R, E266A, A267T, T268A, I272V, K276R, D278G, Y283F, A285T, P307A, N317T, N319G, and I320S. In a more preferred embodiment, the recombinase of SEQ ID NO: 99 has one or more, preferably 2, 3, 4, 5, 6, 8, 9, 10 or more mutations selected from the group consisting of L5Q, V7I, P12T, L14S, P15L, V16A, V23A, M30V, F31L, H40A, M44T, S51T, L58S, Y77H, A80V, K86N, Q90R, G93A, Q94L, S108G, A131V, K132R, F142L, E150G, N151D, Q156K, L164P, I166V, A175S, V182I, K219R, A249V, R259D, A260V, E262R, E266A, A267T, T268A, I272V, K276R, D278G, Y283F, A285T, P307A, N317T, N319G, and I320S compared to the Cre recombinase protein of SEQ ID NO: 68. In the most preferred embodiment, the recombinase of SEQ ID NO: 99 has all of the aforementioned mutations.

[0139] Even more preferably, sequences having at least 70%, preferably 80%, more preferably 90% sequence identity with the recombinase of SEQ ID NO: 99 are characterized by the following specific amino acids: Q at position 5, A at position 40, T at position 44, V at position 80, R at position 90, L at position 94, R at position 132, G at position 150, R at position 219, V at position 272 and R at position 276. These amino acid residues are predicted to be important residues for target site recognition based on deep sequencing and are not found in the sequence of the wild-type Cre recombinase of SEQ ID NO: 68.

[0140] Preferably, the DNA recombinase of SEQ ID NO: 100 contains one or more mutations as compared to the Cre recombinase protein of SEQ ID NO: 68. Exemplary mutations include N3D, L5Q, V7L, N10K, P12S, P15L, V16A, V23A, F31L, R34W, K57E, N60S, Q90H, Q94S, S108G, S147L, A175S, K183R, K211E, N235D, K244R, N245Y, A249K, R259Y, E262Q, E266G, T268A, I272V, R282K, S305P, N317T, N319G and I320S. In a more preferred embodiment, the recombinase of SEQ ID NO: 100 has one or more, preferably 2, 3, 4, 5, 6, 8, 9, 10 or more mutations selected from the group consisting of N3D, L5Q, V7L, N10K, P12S, P15L, V16A, V23A, F31L, R34W, K57E, N60S, Q90H, Q94S, S108G, S147L, A175S, K183R, K211E, N235D, K244R, N245Y, A249K, R259Y, E262Q, E266G, T268A, I272V, R282K, S305P, N317T, N319G and I320S as compared to the Cre recombinase protein of SEQ ID NO: 68. In the most preferred embodiment, the recombinase of SEQ ID NO: 100 has all of the aforementioned mutations.

[0141] Even more preferably, an array having at least 70%, preferably 80%, more preferably 90% sequence identity with the recombinase of SEQ ID NO: 100 is characterized by the following specific amino acids: Q at position 5, H at position 90, S at position 94, R at position 183, K at position 249, G at position 266, V at position 272, and K at position 282. These amino acid residues are predicted to be important residues for target site recognition based on deep sequencing and are not found in the sequence of the wild-type Cre recombinase of SEQ ID NO: 68.

[0142] Preferably, the DNA recombinase monomer of SEQ ID NO: 99 shows a sequence in which the amino acid residue at position 5 is glutamine, the amino acid residue at position 40 is alanine, the amino acid residue at position 44 is threonine, the amino acid residue at position 58 is serine, the amino acid residue at position 80 is valine, the amino acid residue at position 90 is arginine, the amino acid residue at position 94 is leucine, the amino acid residue at position 132 is arginine, the amino acid residue at position 142 is leucine, the amino acid residue at position 150 is glycine, the amino acid residue at position 156 is lysine, the amino acid residue at position 164 is proline, the amino acid residue at position 219 is arginine, and / or the amino acid residue at position 319 is glycine.

[0143] These amino acid residues are specific to the DNA recombinase according to the present invention and are not found in Cre recombinase (SEQ ID NO: 68).

[0144] Preferably, the DNA recombinase monomer of SEQ ID NO: 100 has an aspartic acid at the 3rd amino acid residue, a glutamine at the 5th amino acid residue, a lysine at the 10th amino acid residue, a tryptophan at the 34th amino acid residue, a serine at the 60th amino acid residue, a histidine at the 90th amino acid residue, a serine at the 94th amino acid residue, an arginine at the 183rd amino acid residue, a glutamic acid at the 211th amino acid residue, a lysine at the 249th amino acid residue, a glycine at the 266th amino acid residue, a valine at the 272nd amino acid residue, a lysine at the 282nd amino acid residue, a proline at the 305th amino acid residue, and / or a glycine at the 319th amino acid residue.

[0145] These amino acid residues are specific to the DNA recombinase according to the present invention and are not found in Cre recombinase (SEQ ID NO: 68).

[0146] More preferably, the DNA recombinase according to the present invention is a site-specific recombinase that targets a sequence present in the Int1h repeat sequence of the factor VIII gene that is inverted by the activity of the DNA recombinase according to the present invention.

[0147] The DNA recombinases of the present invention include the polypeptides of SEQ ID NOs: 30 to 33, 93, 94, 99, and 100, and polypeptides having at least 75% similarity (e.g., preferably at least 50%; more preferably at least 70% identity) to the DNA recombinases of SEQ ID NOs: 30 to 33, more preferably at least 85% similarity (e.g., preferably at least 70% identity) to the DNA recombinases of SEQ ID NOs: 30 to 33, 93, 94, 99, and 100, and most preferably at least 95% similarity (e.g., at least 90% identity) to the DNA recombinases of SEQ ID NOs: 30 to 33. Further, they should preferably contain the exact portions of such DNA recombinases that contain sequences of at least 30 amino acids, more preferably at least 50 amino acids. In a preferred embodiment, the term "identity" as used herein refers to the amount of amino acids by one of the amino acid sequences of SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 93, SEQ ID NO: 94, SEQ ID NO: 99, or SEQ ID NO: 100 with respect to the total number of amino acids.

[0148] Fragments or portions of the polypeptides of the present invention can be employed as intermediates for generating the corresponding full-length polypeptides by peptide synthesis. Fragments or portions of the polynucleotides of the present invention can also be used for synthesizing the full-length polynucleotides of the present invention.

[0149] According to the present invention, the activity and specificity of the DNA recombinase are further increased by incorporating them into a fusion protein, and the monomers of the DNA recombinase are linked to each other by a linker as described above.

[0150] Accordingly, in a preferred embodiment, the present invention provides a fusion protein comprising a heterodimer of DNA recombinases, wherein the heterodimer comprises a first recombinase enzyme having an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with a sequence selected from SEQ ID NO: 30, 32, 93 or 99 for recognition of an upstream target sequence and a downstream target sequence of a recombinase target site; and a second recombinase enzyme having an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with a sequence selected from SEQ ID NO: 31, 33, 94 or 100.

[0151] In a more preferred embodiment, the present invention provides a fusion protein comprising a heterodimer of DNA recombinases, wherein the heterodimer comprises a first recombinase enzyme having an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 30 for recognition of an upstream target sequence and a downstream target sequence of a recombinase target site, and a second recombinase enzyme having an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 31.

[0152] In an even more preferred embodiment, the present invention provides a fusion protein comprising a heterodimer of DNA recombinases, wherein the heterodimer comprises a first recombinase enzyme having an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 99 for recognition of an upstream target sequence and a downstream target sequence of a recombinase target site, and a second recombinase enzyme having an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 100.

[0153] In the most preferred embodiment, the present invention provides a fusion protein comprising a heterodimer of DNA recombinases, wherein the heterodimer has an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 32 for recognition of the upstream target sequence and the downstream target sequence of the recombinase target site, a first recombinase enzyme, and an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 33, and a second recombinase enzyme having an amino acid sequence.

[0154] In an even more preferred embodiment, the present invention provides a fusion protein comprising a heterodimer of DNA recombinases, wherein the heterodimer has an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 93 for recognition of the upstream target sequence and the downstream target sequence of the recombinase target site, a first recombinase enzyme, and an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity with the sequence set forth in SEQ ID NO: 94, and a second recombinase enzyme having an amino acid sequence.

[0155] As shown by experiments, the activity of the heterodimer-fusion protein is further affected by the orientation in which the recombinase monomer and the linker are bound to each other. The first recombinase enzyme, the second recombinase enzyme and the linker are i. The C-terminus of the recombinase enzyme having the amino acid sequence of SEQ ID NO: 31 that specifically recognizes the second half-site, such as the right half-site of the recombinase target site, is fused using the N-terminus of the recombinase enzyme having the amino acid sequence of SEQ ID NO: 30 that specifically recognizes the first half-site, such as the left half-site of the recombinase target site, and the linker; more preferably, ii. The C-terminus of a recombinase enzyme having the amino acid sequence of SEQ ID NO: 33 that specifically recognizes a second half-site, such as the right half-site of the recombinase target site, is fused using a linker to the N-terminus of a recombinase enzyme having the amino acid sequence of SEQ ID NO: 32 that specifically recognizes a first half-site, such as the left half-site of the recombinase target site, When joined to each other as such, the best activity of the heterodimer on the loxF8 site was shown.

[0156] The first recombinase enzyme, the second recombinase enzyme and the linker are, i. The C-terminus of a recombinase enzyme having the amino acid sequence of SEQ ID NO: 94 that specifically recognizes a second half-site, such as the right half-site of the recombinase target site, is fused using a linker to the N-terminus of a recombinase enzyme having the amino acid sequence of SEQ ID NO: 93 that specifically recognizes a first half-site, such as the left half-site of the recombinase target site; more preferably, ii. The C-terminus of a recombinase enzyme having the amino acid sequence of SEQ ID NO: 100 that specifically recognizes a second half-site, such as the right half-site of the recombinase target site, is fused using a linker to the N-terminus of a recombinase enzyme having the amino acid sequence of SEQ ID NO: 99 that specifically recognizes a first half-site, such as the left half-site of the recombinase target site, When joined to each other as such, even better activity of the heterodimer on the loxF8 site was shown.

[0157] Therefore, in a preferred embodiment, the full fusion protein of the present invention is, i. The sequence set forth in SEQ ID NO: 71 or SEQ ID NO: 103; or most preferably ii. The sequence set forth in SEQ ID NO: 72 or SEQ ID NO: 97 and has an amino acid sequence having at least 70%, preferably 80%, more preferably 90% sequence identity.

[0158] In a further embodiment, the present invention provides that the first recombinase enzyme, the second recombinase enzyme and the linker are, i. The C-terminus of a recombinase enzyme having the amino acid sequence of SEQ ID NO: 93 that specifically recognizes a first half-site, such as the left half-site of the recombinase target site, is fused using a linker to the N-terminus of a recombinase enzyme having the amino acid sequence of SEQ ID NO: 94 that specifically recognizes a second half-site, such as the right half-site of the recombinase target site; more preferably, ii. The C-terminus of a recombinase enzyme having the amino acid sequence of SEQ ID NO: 99 that specifically recognizes a first half-site, such as the left half-site of the recombinase target site, is fused using a linker to the N-terminus of a recombinase enzyme having the amino acid sequence of SEQ ID NO: 100 that specifically recognizes a second half-site, such as the right half-site of the recombinase target site, to provide a fusion protein that is bound to each other in this way.

[0159] The present invention further relates to a nucleic acid molecule such as a polynucleotide or nucleic acid encoding a DNA recombinase according to the present invention, or a monomer or fusion protein thereof.

[0160] The "polynucleotide" or "nucleic acid" of the present invention may be in the form of RNA or DNA; DNA is to be understood to include cDNA, genomic DNA, recombinant DNA, and synthetic DNA. DNA may be double-stranded or single-stranded, and in the case of single-stranded, it may be the coding strand or the non-coding (antisense) strand. The coding sequence encoding the polypeptide may be identical to the coding sequences of the polypeptides shown in SEQ ID NOs: 30 to 33, preferably SEQ ID NOs: 32 and 33, or may be different coding sequences encoding the same polypeptide as a result of the redundancy or degeneracy of the genetic code or single nucleotide polymorphisms. For example, it may be an RNA transcript containing the full length of the coding sequence of any one of the polypeptides of SEQ ID NOs: 30 to 33, 93, 94, 99, and 100. In a preferred embodiment, the "polynucleotide" according to the present invention is one of SEQ ID NOs: 34 to 37, 67, 95, 96, 101, and 102, wherein the polynucleotide of SEQ ID NO: 34 encodes the DNA recombinase monomer of SEQ ID NO: 30; the polynucleotide of SEQ ID NO: 35 encodes the DNA recombinase monomer of SEQ ID NO: 31, the polynucleotide of SEQ ID NO: 36 encodes the DNA recombinase monomer of SEQ ID NO: 32, and the polynucleotide of SEQ ID NO: 37 encodes the DNA recombinase monomer of SEQ ID NO: 33; the polynucleotide of SEQ ID NO: 67 encodes the DNA recombinase monomer of SEQ ID NO: 66, the polynucleotide of SEQ ID NO: 95 encodes the DNA recombinase monomer of SEQ ID NO: 93, the polynucleotide of SEQ ID NO: 96 encodes the DNA recombinase monomer of SEQ ID NO: 94, the polynucleotide of SEQ ID NO: 101 encodes the DNA recombinase monomer of SEQ ID NO: 99, and the polynucleotide of SEQ ID NO: 102 encodes the DNA recombinase monomer of SEQ ID NO: 100.

[0161] Nucleic acids encoding the polypeptides of SEQ ID NOs: 30-33, 66, 93, 94, 99 and 100, preferably SEQ ID NOs: 32 and 33, or SEQ ID NOs: 93 and 94, may include, but are not limited to, the coding sequence of the polypeptide alone; the coding sequence of the polypeptide and additional coding sequences such as a leader or secretory sequence or a proprotein sequence; and non-coding sequences such as introns or 5' and / or 3' non-coding sequences of the coding sequence of the polypeptide (and optionally additional coding sequences). Nucleic acids encoding the polypeptides of SEQ ID NOs: 30-33, 66, 93, 94, 99 and 100, preferably SEQ ID NOs: 32 and 33, or SEQ ID NOs: 93 and 94, include nucleic acids codon-optimized for expression in human cells. They may further include a nuclear localization sequence.

[0162] Therefore, the term "polynucleotide encoding a polypeptide" or the term "nucleic acid encoding a polypeptide" is to be understood to encompass polynucleotides or nucleic acids that contain only the coding sequence of the DNA recombinase enzyme of the present invention, for example, polypeptides selected from SEQ ID NOs: 30-33, 66, 93, 94, 99 and 100, preferably SEQ ID NOs: 32 and 33, or SEQ ID NOs: 93 and 94, as well as those containing additional coding sequences and / or non-coding sequences. The terms polynucleotide and nucleic acid are used interchangeably.

[0163] The present invention also includes polynucleotides in which the coding sequence of the polypeptide can be fused in the same reading frame to a polynucleotide sequence that aids in the expression and secretion of the polypeptide from the host cell; for example, a leader sequence that functions as a secretion sequence for controlling the transport of the polypeptide from the cell may be so fused. A polypeptide having such a leader sequence is called a preprotein or preproprotein, and the leader sequence can be cleaved by the host cell to form the mature form of the protein. These polynucleotides may have a 5' extension region that encodes a proprotein, which is a mature protein with additional amino acid residues added at the N-terminus. An expression product having such a pro sequence is called a proprotein, which is an inactive form of the mature protein; however, when the pro sequence is cleaved, the active mature protein remains. Additional sequences may also bind to the protein and become part of the mature protein. Thus, for example, the polynucleotides of the present invention may encode a polypeptide, or a protein having a pro sequence, or a protein having both a pro sequence and a pre sequence (such as a leader sequence).

[0164] The polynucleotides of the present invention may also have a coding sequence fused in-frame to a marker sequence that enables the purification of the polypeptides of the present invention. The marker sequence can be an affinity tag or epitope tag such as a polyhistidine tag, streptavidin tag, Xpress tag, FLAG tag, cellulose or chitin binding tag, glutathione-S-transferase tag (GST), hemagglutinin (HA) tag, c-myc tag or V5 tag.

[0165] The HA tag corresponds to an epitope derived from the influenza hemagglutinin protein (Wilson et al., 1984), and the c-myc tag can be an epitope from the human Myc protein (Evans et al., 1985).

[0166] When the nucleic acid of the present invention is mRNA, particularly for use as a medicine, the delivery of mRNA therapeutics has been facilitated by significant progress in maximizing mRNA translation and stability, preventing its immunostimulatory activity, and developing in vivo delivery technologies. The 5' cap and 3' poly(A) tail mainly contribute to the efficient translation of mature eukaryotic mRNA and the extension of its half-life. Incorporating cap analogs such as ARCA (anti-reverse cap analog) and a 120-150 bp poly(A) tail into in vitro transcribed (IVT) mRNA significantly improves the expression of the encoded protein and the stability of mRNA. New types of cap analogs, such as 1,2-dithiodiphosphate-modified caps, have resistance to the RNA decapping complex and can further improve the efficiency of RNA translation. So-called codon optimization, which replaces rare codons within the mRNA protein-coding sequence with synonymous codons that occur frequently, promotes more efficient protein synthesis and limits mRNA destabilization by rare codons, thus preventing accelerated degradation of the transcript. Similarly, manipulating the 3' and 5' untranslated regions (UTRs) containing sequences involved in the recruitment of RNA-binding proteins (RBPs) and miRNAs can increase the level of the protein product. Interestingly, the UTRs can be intentionally modified to encode regulatory elements (e.g., K-turn motifs and miRNA binding sites), providing a means to control RNA expression in a cell-specific manner. Some RNA base modifications, such as N1-methyl-pseudouridine, not only help mask the immunostimulatory activity of mRNA but also have been shown to increase mRNA translation by enhancing translation initiation. In addition to the observed effects on protein translation, base modification and codon optimization affect the secondary structure of mRNA, which in turn affects translation. Each modification of the nucleic acid molecule of the present invention is also contemplated by the present invention.

[0167] The present invention further provides polynucleotides that hybridize to the above-described sequences herein that encode proteins having at least 70%, preferably at least 90%, more preferably at least 95% identity or similarity between the sequences and thus have similar biological activities. Further, as is known in the art, "similarity" exists between two polypeptides when the amino acid sequences contain identical or conserved amino acid substitutions for individual residues in the sequence. Identity and similarity can be measured using sequence analysis software (e.g., ClustalW of PBIL (Pole Bioinformatique Lyonnais) at http: / / npsa-pbil.ibcp.fr). In particular, the present invention provides polynucleotides that hybridize to the polynucleotides described above herein under stringent conditions.

[0168] Suitable stringent conditions can be defined, for example, by the concentration of salt or formamide in the prehybridization and hybridization solutions, or the hybridization temperature, and are well known in the art. In particular, stringency can be increased by decreasing the concentration of salt, increasing the concentration of formamide, and / or increasing the hybridization temperature.

[0169] For example, hybridization under high stringency conditions can be carried out at about 37°C to 42°C, employing about 50% formamide, while hybridization under reduced stringency conditions can be carried out at about 30°C to 35°C, employing from about 35% to 25% formamide. One specific set of conditions for hybridization under high stringency conditions employs 50% formamide, 5x SSPE, 0.3% SDS, and 200 μg / ml sheared and denatured salmon sperm DNA at 42°C. For hybridization under reduced stringency, similar conditions as above can be used with 35% formamide at a reduced temperature of 35°C. The temperature range corresponding to a particular level of stringency can be further narrowed by calculating the purine to pyrimidine ratio of the nucleic acid of interest and adjusting the temperature accordingly. Variations of the above ranges and conditions are well known in the art. Preferably, hybridization should occur only when there is at least 95%, more preferably at least 97% identity between the sequences. In a preferred embodiment, the polynucleotide hybridizing to the polynucleotide described above herein encodes a polypeptide that exhibits a biological function or activity substantially identical to the mature proteins of SEQ ID NOs: 30 - 33, 66, 93, 94, 99, and 100, preferably SEQ ID NOs: 32 and 33, or SEQ ID NOs: 93 and 94.

[0170] As described above, a suitable polynucleotide probe may have at least 14 bases, preferably 30 bases, more preferably at least 50 bases, and hybridizes to a polynucleotide of the present invention having identity therewith as described hereinabove. For example, such a polynucleotide may be employed as a probe for hybridization to a polynucleotide encoding a polypeptide of SEQ ID NOs: 30-33, 66, 93, 94, 99 and 100, such as the polynucleotides of SEQ ID NOs: 34-37, 67, 95, 96, 101 and 102, for example, for the recovery of such a polynucleotide, or as a diagnostic probe, or as a PCR primer. Therefore, the present invention includes a polynucleotide having at least 70% identity, preferably at least 90% identity, more preferably at least 95% identity with a polynucleotide of SEQ ID NOs: 34-37, 67, 95, 96, 101 and 102 encoding a polypeptide of SEQ ID NOs: 30-33, 66, 93, 94, 99 and 100, and fragments thereof having preferably at least 30 bases, more preferably at least 50 bases.

[0171] The terms "homology" or "identity", as used interchangeably herein, refer to sequence similarity between two polynucleotide sequences or between two polypeptide sequences, with identity being a more stringent comparison. The phrases "percent identity or homology" and "identity or homology" refer to the percentage of sequence similarity found in a comparison of two or more polynucleotide sequences or two or more polypeptide sequences. "Sequence similarity" refers to the percent similarity in base pair sequence (determined by any suitable method) between two or more polynucleotide sequences. Two or more sequences can have a similarity of anywhere from 0 to 100%, or any integer value therebetween. Identity or similarity can be determined by comparing positions within each sequence that can be aligned for purposes of comparison. When a position within the sequences being compared is occupied by the same nucleotide base or amino acid, the molecules are identical at that position. The degree of similarity or identity between polynucleotide sequences is a function of the number of identical or matching nucleotides at positions shared by the polynucleotide sequences.

[0172] The degree of identity of polypeptide sequences is a function of the number of identical amino acids at positions shared by the polypeptide sequences. The degree of homology or similarity of polypeptide sequences is a function of the number of amino acids at positions shared by the polypeptide sequences. As used herein, the term "substantially identical" refers to at least 70%, 75%, at least 80%, at least 85%, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identity or homology.

[0173] The degree of sequence identity is determined by selecting one sequence as the query sequence and aligning it with homologous sequences obtained from GenBank using the blastp algorithm (NCBI) and ClustalW, an internet-based tool.

[0174] As is well known in the art, the genetic code is redundant in that certain amino acids are encoded by more than one nucleotide triplet (codon), and the present invention includes polynucleotide sequences that encode the same amino acids using codons different from those specifically exemplified in the sequences herein. Such polynucleotide sequences are referred to herein as "equivalent" polynucleotide sequences. The present invention further includes variants of the polynucleotides described herein that encode fragments, such as some or all of the proteins, analogs, and derivatives of the polypeptides of SEQ ID NOs: 30 - 33, 66, 93, 94, 99, and 100. Variant forms of polynucleotides can be allelic variants of naturally occurring polynucleotides or variants of non-naturally occurring polynucleotides. For example, a variant of a nucleic acid can simply be a difference in the codon sequence of an amino acid due to the degeneracy of the genetic code, or there can be deletion variants, substitution variants, and addition or insertion variants. As is known in the art, an allelic variant is an alternative form of a polynucleotide sequence that can have one or more nucleotide substitutions, deletions, or additions that do not substantially alter the biological function of the encoded polypeptide.

[0175] In a more preferred embodiment, the polynucleotide of the present invention encodes the full-length fusion protein of the present invention, more preferably encodes a heterodimer of a DNA recombinase, and the heterodimer comprises a first recombinase enzyme having an amino acid sequence with at least 70%, preferably 80%, more preferably 90% sequence identity to the sequence set forth in SEQ ID NO: 30, 32, 93 or 99 for recognition of the upstream target sequence and the downstream target sequence of the recombinase target site, and a second recombinase enzyme having an amino acid sequence with at least 70%, preferably 80%, more preferably 90% sequence identity to the sequence set forth in SEQ ID NO: 31, 33, 94 or 100, and the polynucleotide further encodes a linker as described herein.

[0176] In one embodiment, the polynucleotide of the present invention comprises the nucleic acid of SEQ ID NO: 34 and the nucleic acid of SEQ ID NO: 35, and a nucleic acid encoding a linker oligopeptide.

[0177] In a further embodiment, the polynucleotide of the present invention comprises the nucleic acid of SEQ ID NO: 36 and the nucleic acid of SEQ ID NO: 37, and a nucleic acid encoding a linker oligopeptide.

[0178] In a further embodiment, the polynucleotide of the present invention comprises the nucleic acid of SEQ ID NO: 95 and the nucleic acid of SEQ ID NO: 96, and a nucleic acid encoding a linker oligopeptide.

[0179] In a further embodiment, the polynucleotide of the present invention comprises the nucleic acid of SEQ ID NO: 101 and the nucleic acid of SEQ ID NO: 102, and a nucleic acid encoding a linker oligopeptide.

[0180] The nucleic acid encoding the linker oligopeptide is preferably a nucleic acid selected from SEQ ID NOs: 38-54 (see Table 3 of Example 2).

[0181] More preferably, the nucleic acid encoding the linker oligopeptide is preferably a nucleic acid selected from SEQ ID NOs: 55 to 63 (see Table 4 in Example 2).

[0182] Most preferably, the nucleic acid encoding the fusion protein of the present invention is i. a nucleic acid having the sequences of SEQ ID NOs: 34, 35, and 61; a nucleic acid such as SEQ ID NO: 64; or ii. a nucleic acid having the sequences of SEQ ID NOs: 36, 37, and 61; a nucleic acid such as SEQ ID NO: 65; or iii. a nucleic acid having the sequences of SEQ ID NOs: 93, 94, and 61; a nucleic acid such as SEQ ID NO: 98; or iv. a nucleic acid having the sequences of SEQ ID NOs: 99, 100, and 61; a nucleic acid such as SEQ ID NO: 104; or v. a nucleic acid having the sequences of SEQ ID NOs: 95 and 96; or vi. a nucleic acid having the sequences of SEQ ID NOs: 101 and 102; vii. having the sequences of SEQ ID NOs: 34 and 35; or viii. a nucleic acid having the sequences of SEQ ID NOs: 36 and 37; or ix. the nucleic acid of SEQ ID NO: 67 and is a nucleic acid containing the same.

[0183] Most preferably, the nucleic acid encoding the fusion protein of the present invention is a nucleic acid containing or consisting of SEQ ID NO: 64.

[0184] Even more preferably, the nucleic acid encoding the fusion protein of the present invention is a nucleic acid containing or consisting of SEQ ID NO: 65.

[0185] Even more preferably, the nucleic acid encoding the fusion protein of the present invention is a nucleic acid containing or consisting of SEQ ID NO: 98.

[0186] Even more preferably, the nucleic acid encoding the fusion protein of the present invention is a nucleic acid containing or consisting of SEQ ID NO: 104.

[0187] The present invention also relates to a vector containing such a polynucleotide, a host cell genetically engineered with such a vector, and the production of the polypeptides of SEQ ID NOs: 30-33, 66, 93, 94, 99 and 100, preferably SEQ ID NOs: 32 and 33, or SEQ ID NOs: 93 and 94 by recombinant techniques using the foregoing. The host cell is genetically engineered (transduced or transformed or transconjugated or transfected) with such a vector which can be, for example, a cloning vector or an expression vector. The vector can be in the form of, for example, a plasmid, a conjugative plasmid, a virus particle, a phage, etc. The vector or gene can be integrated into the chromosome at a specific site or a non-specific site. Methods of genomic integration of recombinant DNA such as homologous recombination or transposase-mediated integration are well known in the art. The engineered host cell can be cultured in a conventional nutrient medium appropriately modified to activate a promoter, select a transformant, or amplify the gene of the present invention. Culture conditions such as temperature, pH, etc. are those commonly used with the host cell selected for expression and are well known to those skilled in the art. The host cell can be a mammalian cell, an insect cell, a plant cell or a bacterial host cell containing the nucleic acid or recombinant polynucleotide molecule or expression vector described herein.

[0188] The polynucleotide sequence in the expression vector is operably linked to an appropriate expression control sequence(s) (promoter) to direct mRNA synthesis. Representative examples of such promoters include the LTR or SV40 promoter, the E. coli lac, ara, rha or trp, the phage lambda PL promoter and other promoters known to control the expression of genes in prokaryotic or eukaryotic cells or their viruses.

[0189] One skilled in the art can select a vector based on desired properties for the production of the vector in a particular cell, such as a mammalian cell or a bacterial cell, for example.

[0190] Any of a variety of inducible promoters or enhancers can be included in a vector for the expression of the antibodies or nucleic acids that can be regulated of the present invention. Such inducible systems include, for example, the tetracycline inducible system; the metallothionein promoter induced by heavy metals; insect steroid hormones that respond to related steroids such as ecdysone or muristerone; the mouse mammary tumor virus (MMTV) induced by steroids such as glucocorticoids and estrogens; and the heat shock promoter that can be induced by temperature changes; the rat neuron-specific enolase gene promoter; the human β-actin gene promoter; the human platelet-derived growth factor B (PDGF-B) chain gene promoter; the rat sodium channel gene promoter; the human copper-zinc superoxide dismutase gene promoter; and promoters for members of the mammalian POU domain regulatory gene family.

[0191] Regulatory elements containing a promoter or enhancer can be constitutive or regulatable depending on their regulatory nature. A regulatory sequence or regulatory element is operably linked to one of the polynucleotide sequences of the present invention such that the physical and functional relationship between the polynucleotide sequence and the regulatory sequence allows transcription of the polynucleotide sequence. Vectors useful for expression in eukaryotic cells can include, for example, regulatory elements including the CAG promoter, the SV40 early promoter, the cytomegalovirus (CMV) promoter, the mouse mammary tumor virus (MMTV) steroid inducible promoter, Pgtf, the Moloney murine leukemia virus (MMLV) promoter, the thy-1 promoter, and the like.

[0192] Optionally, the vector can include a selectable marker. As used herein, "selectable marker" refers to a genetic element that provides a selectable phenotype to a cell into which the selectable marker has been introduced. Selectable markers are generally genes whose gene products confer resistance to a drug that inhibits cell growth or kills cells. For example, various selectable markers including, for example, the Neo, Hyg, hisD, Gpt, and Ble genes described in Ausubel et al., 1999 and U.S. Patent No. 5,981,830 can be used in the DNA constructs of the present invention. Drugs useful for selecting for the presence of a selectable marker include, for example, G418 for Neo, hygromycin for Hyg, histidinol for hisD, xanthine for Gpt, and bleomycin for Ble. The DNA constructs of the present invention can incorporate a positive selectable marker, a negative selectable marker, or both.

[0193] A variety of mammalian cell culture systems can also be employed to express recombinant proteins. Examples of mammalian expression systems include the COS-7 strain of monkey kidney fibroblasts. Other cell lines capable of expressing compatible vectors include, for example, the C127, 3T3, CHO, HeLa, and BHK cell lines. Mammalian expression vectors generally also include an origin of replication, appropriate promoters and enhancers, and any necessary ribosome binding sites, polyadenylation sites, splice donor and acceptor sites, transcription termination sequences, and 5' flanking non-transcribed sequences. DNA sequences derived from SV40 splicing and polyadenylation sites can be used to provide the necessary non-transcribed genetic elements.

[0194] The polypeptide can be recovered and purified from recombinant cell cultures by methods including ammonium sulfate or ethanol precipitation, acid extraction, anion or cation exchange chromatography, phosphocellulose chromatography, hydrophobic interaction chromatography, affinity chromatography, hydroxylapatite chromatography, and lectin chromatography. When the polypeptide is expressed on the surface of the cell, recovery can be facilitated, but this is not a prerequisite. Also, it may be desirable to recover cleavage products that are cleaved after expression of the longer polypeptide form. Protein refolding steps, such as those known in the art, can be used, if necessary, to complete the conformation of the mature protein. High performance liquid chromatography (HPLC) can be employed in the final purification step.

[0195] According to a further embodiment of the invention, there is provided, for example, a gene therapy vector for use in increasing the systemic or local expression of the fusion protein of the invention in a subject. The gene therapy vector is used for the prevention, alleviation, amelioration, reduction, inhibition, and / or treatment of diseases that can be treated by genome editing, particularly hemophilia A. The gene therapy vector typically comprises an expression cassette comprising a polynucleotide encoding the fusion protein of the invention. In one embodiment, the vector is a viral vector. In a preferred embodiment, the viral vector is derived from a virus selected from the group consisting of adenovirus, retrovirus, lentivirus, herpes virus, and adeno-associated virus (AAV). In a more preferred embodiment, the vector is derived from one or more of adeno-associated virus (AAV) serotypes 1-11, or any subgroup or any engineered form thereof. In another embodiment, the viral vector is encapsulated in an anionic liposome.

[0196] In another embodiment, the vector is a non-viral vector. In a preferred embodiment, the non-viral vector is selected from the group consisting of naked DNA, cationic liposome complexes, cationic polymer complexes, cationic liposome-polymer complexes, and exosomes.

[0197] When the vector is a viral vector, the expression cassette suitably (from the perspective of the transcribed mRNA) includes a first inverted terminal repeat, enhancer, promoter, polynucleotide encoding the fusion protein of the present invention, 3' untranslated region, polyadenylation (polyA) signal, and a second inverted terminal repeat, which are operably linked in the 5' to 3' direction. The promoter is selected from the group consisting of, for example, the cytomegalovirus (CMV) promoter and the chicken beta-actin (CAG) promoter. The polynucleotide preferably includes DNA or cDNA or RNA or mRNA. In a preferred embodiment, the polynucleotide encoding the fusion protein of the present invention includes one or more of the polypeptides of SEQ ID NO: 30, SEQ ID NO: 31, SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 66, SEQ ID NO: 93, SEQ ID NO: 94, SEQ ID NO: 99, and SEQ ID NO: 100. In the most preferred embodiment, the polynucleotide encoding the fusion protein of the present invention has at least about 75%, 80%, 85%, or 90% sequence identity, for example, at least about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity with one or more of SEQ ID NO: 64, SEQ ID NO: 65, SEQ ID NO: 67, SEQ ID NO: 98, and SEQ ID NO: 104.

[0198] The present invention further relates to the fusion protein of the present invention or the nucleic acid molecule, recombinant polynucleotide, or expression vector of the present invention for use as a medicament. In a more preferred embodiment, the present invention further relates to the fusion protein of the present invention or the nucleic acid molecule, recombinant polynucleotide, or expression vector of the present invention for use in the prevention or treatment of diseases that can be treated by genome editing, particularly hemophilia A.

[0199] In a further embodiment, the present invention relates to the use of the fusion protein of the present invention or the nucleic acid molecule, recombinant polynucleotide or expression vector of the present invention for the preparation of a medicament for the prevention or treatment of a disease that can be treated by genome editing, particularly hemophilia A.

[0200] In a further embodiment, the present invention relates to a method for the prevention or treatment of a disease that can be treated by genome editing, particularly hemophilia A, the method comprising administering a therapeutically effective amount of the fusion protein of the present invention or the nucleic acid molecule, recombinant polynucleotide or expression vector of the present invention to a patient in need thereof.

[0201] The fusion protein of the present invention or the nucleic acid molecule, recombinant polynucleotide or expression vector of the present invention is particularly suitable for the treatment of severe hemophilia A.

[0202] The fusion protein of the present invention or the nucleic acid molecule, recombinant polynucleotide or expression vector of the present invention can further be included in a pharmaceutical composition that may further contain one or more pharmaceutically acceptable diluents or carriers.

[0203] Provided is a pharmaceutical composition for use in preventing or treating a disorder that can be treated by genome editing, such as hemophilia A, the composition comprising a nucleic acid sequence of a polynucleotide encoding one or more fusion proteins according to the present invention, or a nucleic acid encoding a therapeutically active amount of the fusion protein of the present invention, or a therapeutically effective amount of a vector comprising a recombinant fusion protein of the present invention (collectively referred to as a "therapeutic active agent").

[0204] It will be understood that the single or total daily dosage of the therapeutically active agents and compositions of the present invention will be determined by the attending physician within the scope of sound medical judgment. The specific therapeutically effective dosage level for any particular patient will depend upon a variety of factors including the disorder being treated and the severity of the disorder; the activity of the specific compound employed; the specific composition employed, the age, weight, general health, sex and diet of the patient; the duration of administration, the route of administration, and the rate of excretion of the specific compound employed; the period of treatment; drugs employed in combination with, or used in conjunction with, the specific nucleic acid or polypeptide employed; and various factors well known in the medical arts. For example, it is within the skill of the art to start the administration of the compounds at a level lower than that required to achieve the desired therapeutic effect and to gradually increase the dosage until the desired effect is achieved. However, the daily dosage of the product can vary widely for an adult. The therapeutically effective amount of a therapeutically active agent such as a vector according to the present invention to be administered, as well as the dosage for treating a pathological condition with a number of viral or non-viral particles and / or pharmaceutical compositions described herein, depends on a number of factors including the age and condition of the patient, the severity of the disturbance or disorder, the method and frequency of administration, and the specific peptide employed.

[0205] The pharmaceutical composition containing a therapeutic agent according to the present invention may be in any form suitable for the selected mode of administration.

[0206] In one embodiment, the pharmaceutical composition of the present invention is administered parenterally.

[0207] As used herein, the phrases "parenteral administration" and "administered parenterally" mean a mode of administration other than enteral and topical administration, usually by injection, and include intradermal, intravenous, intramuscular, intraarterial, intrathecal, intracapsular, intraorbital, intracardiac, intradermal, intraperitoneal, intratendinous, transtracheal, subcutaneous, subcuticular, intraarticular, subcapsular, subarachnoid, intraspinal, intracranial, intrathoracic, epidural and intrasternal injections and infusions.

[0208] The therapeutic agent of the present invention can be administered to animals and humans in unit dosage forms, as a single active agent or in combination with other active agents, as a mixture with conventional pharmaceutical carriers.

[0209] In a further embodiment, the pharmaceutical composition contains a pharmaceutically acceptable vehicle for injectable formulations. These can be, in particular, isotonic sterile physiological saline (such as monosodium phosphate or disodium phosphate, sodium chloride, potassium chloride, calcium chloride or magnesium chloride, or a mixture of such salts), or can be dry, in particular lyophilized compositions, which, upon addition, enable the constitution of an injectable solution, depending on the case, with sterile water or physiological saline.

[0210] Pharmaceutical forms suitable for injectable use include sterile aqueous solutions or dispersions; formulations containing sesame oil, peanut oil or aqueous propylene glycol; and sterile powders for the immediate preparation of sterile injectable solutions or dispersions. In all cases, the form must be sterile and must be liquid. It must be stable under the conditions of manufacture and storage and must be preserved against the contaminating action of microorganisms such as bacteria and fungi.

[0211] Solutions containing therapeutically active agents as free bases or pharmaceutically acceptable salts can be prepared with water suitably mixed with surfactants such as hydroxypropylcellulose. Dispersions can also be prepared with glycerol, liquid polyethylene glycols, and mixtures thereof and oils. Under normal conditions of storage and use, these preparations contain preservatives to prevent the growth of microorganisms.

[0212] Therapeutically active agents can be formulated into compositions in neutral or salt form. Pharmaceutically acceptable salts include acid addition salts (formed with the free amino groups of the protein), for example, inorganic acids such as hydrochloric acid or phosphoric acid, or organic acids such as acetic acid, oxalic acid, tartaric acid, mandelic acid, etc. Salts formed with free carboxyl groups can also be derived from inorganic bases such as sodium, potassium, ammonium, calcium, or ferric hydroxide, and organic bases such as isopropylamine, trimethylamine, histidine, procaine, etc.

[0213] The carrier can also be present as a solvent or dispersion medium, for example, including water, ethanol, polyols (such as glycerol, propylene glycol, and liquid polyethylene glycol, etc.), suitable mixtures thereof, and vegetable oils. Appropriate fluidity can be maintained, for example, by the use of coatings such as lecithin, by maintaining the required particle size in the case of dispersions, and by the use of surfactants. Prevention of microbial action can be brought about by various antibacterial and antifungal agents, for example, parabens, chlorobutanol, phenol, sorbic acid, thimerosal, etc. In many cases, it is preferable to include isotonic agents, for example, sugars or sodium chloride. Prolonged absorption of injectable compositions can be brought about by the use in the composition of agents that delay absorption, for example, aluminum monostearate and gelatin.

[0214] Sterile injectable solutions are prepared, if necessary, by incorporating the required amount of the active polypeptide into a suitable solvent together with some of the other ingredients listed above, followed by filtration sterilization. Generally, dispersions are prepared by incorporating various sterilized active ingredients into a sterile vehicle containing a basic dispersion medium and the necessary other ingredients from those listed above. In the case of sterile powders for the preparation of sterile injectable solutions, the preferred methods of preparation are vacuum drying and freeze-drying techniques that yield a powder of the active ingredient plus any additional desired ingredients from a previously sterile filtered solution.

[0215] Upon formulation, the solution can be administered in a manner compatible with the dosage formulation and in a therapeutically effective amount. The formulation can be easily administered in various dosage forms such as the above-mentioned injection solutions, but drug release capsules and the like can also be employed. It can also be administered in multiple doses. Where appropriate, the therapeutic agents described herein can be formulated with any suitable vehicle for delivery. For example, they can be incorporated into pharmaceutically acceptable suspensions, solutions or emulsions. Suitable media include physiological saline and liposome preparations. More specifically, pharmaceutically acceptable carriers include sterile aqueous carriers of non-aqueous solutions, suspensions, and emulsions. Examples of non-aqueous solvents are propylene glycol, polyethylene glycol, vegetable oils such as olive oil, and organic esters for injection such as ethyl oleate. Aqueous carriers include water including physiological saline and buffer media, alcoholic / aqueous solutions, emulsions or suspensions, but are not limited thereto. Intravenous vehicles include body fluids and nutrient supplements, electrolyte supplements (such as those based on Ringer's dextrose), and the like.

[0216] For example, preservatives and other additives such as antibacterial agents, antioxidants, chelating agents, and inert gases may also be present.

[0217] Colloidal dispersions can also be used for targeted gene delivery. Colloidal dispersions include lipid-based systems such as macromolecular complexes, nanocapsules, microspheres, beads, and oil-in-water emulsions, micelles, mixed micelles, and liposomes.

[0218] An appropriate treatment regimen can be determined by a physician and depends on the age, sex, weight, and disease stage of the subject. As an example, for the delivery of the nucleic acid sequence encoding the fusion protein of the present invention using a viral expression vector, each unit dosage of the fusion protein expression vector can be, for example, in a viral genome concentration in the range of 10 11 ~10 16 and can include a composition of 2.5 μl to 100 μl containing the viral expression vector in a pharmaceutically acceptable fluid.

[0219] The effective dosage and dosing schedule for administering the fusion of the present invention in the form of a recombinant polypeptide depend on the disease or condition to be treated and can be determined by those skilled in the art. Exemplary and non-limiting ranges for the therapeutically effective amount of the fusion protein of the present invention are about 0.1 to 10 mg / kg body weight, 0.1 to 5 mg / kg body weight, etc., for example, 0.1 to 2 mg / kg body weight, 0.1 to 1 mg / kg body weight, etc., for example, about 0.15, about 0.2, about 0.5, about 1, about 1.5 or about 2 mg / kg body weight.

[0220] A physician or veterinarian having ordinary skill in the art can readily determine and prescribe the effective amount of the required pharmaceutical composition. For example, the physician or veterinarian can start with a dosage of the therapeutic agent of the present invention employed in the pharmaceutical composition at a level lower than that required to achieve the desired therapeutic effect and gradually increase the dosage until the desired effect is achieved. Generally, the appropriate daily dosage of the composition of the present invention is the amount of the delivery system that is the lowest dosage effective to produce a therapeutic effect. Such effective dosages generally depend on the factors described above. Administration can be, for example, intravenous, intramuscular, intraperitoneal, or subcutaneous, for example, administered proximal to the target site. Optionally, the effective daily dosage of the pharmaceutical composition can be administered, if desired, in unit dosage form, as two, three, four, five, six or more sub-dosages administered separately at appropriate intervals throughout the day. It is also possible to administer the delivery system of the present invention alone, but it is preferred to administer it as the delivery system in the above-mentioned pharmaceutical composition.

[0221] There is further provided a kit comprising the therapeutic agent described above and herein. In one embodiment, the kit provides the therapeutic agent in one or more unit dosage forms ready for administration to a subject, for example, prepared in a preloaded syringe or ampoule. In another embodiment, the therapeutic agent is provided in lyophilized form.

[0222] In a further embodiment, the present invention is a method for determining genomic-level recombination in a host cell culture, comprising a fusion protein for efficient and specific genomic editing according to the present invention, i. providing a nucleic acid molecule encoding a first recombinase enzyme, wherein the first recombinase enzyme has been evolved by directed evolution or rational design to specifically recognize a first half-site of a recombinase target site; ii. providing a nucleic acid molecule encoding a second recombinase enzyme, wherein the second recombinase enzyme has been evolved by directed evolution or rational design to specifically recognize a second half-site of a recombinase target site; iii. providing a nucleic acid molecule encoding a linker peptide, wherein the nucleic acid encodes a linker oligopeptide comprising 4 to 60 amino acids; iv. preparing an expression vector by cloning a nucleic acid molecule encoding a first recombinase enzyme, a nucleic acid molecule encoding a second recombinase enzyme, and a nucleic acid molecule encoding a linker peptide into an expression vector further comprising a first reporter gene for expressing a first reporter protein; v. transfecting a host cell with the expression vector of step iv) and transfecting the host cell with a reporter plasmid comprising a second reporter gene for expressing a second reporter protein; vi. expressing a fusion protein comprising a first recombinase enzyme, a second recombinase enzyme, and a linker peptide, wherein the fusion protein is fused to the first reporter gene, and expressing a second reporter protein; vii. identifying cells that show dual expression of a first reporter protein and a second reporter protein and indicate successful recombination A method comprising the steps is provided.

[0223] A suitable first reporter gene for step iv) is a gene encoding EGFP. As a result, a suitable first reporter protein is EGFP.

[0224] The appropriate second reporter gene for project v is the gene encoding mCherry. As a result, the appropriate second reporter protein is mCherry.

[0225] This system has the advantage that the transfection efficiency of cells transfected with both the expression plasmid and the reporter plasmid can be measured based on GFP fluorescence. GFP and mCherry double-positive cells reflect the recombination of the reporter in human cells. To calculate the recombination efficiency of the reporter plasmid in human cells, double-positive cells can be normalized to the transfection efficiency.

[0226] The fusion protein described herein was developed to correct a large gene inversion in exon 1 of the F8 gene causing hemophilia A. To examine the inversion efficacy of the heterodimer at the genomic level, an in vitro recombinase assay described in Example 9 was developed. It was found that the inversion efficacy at the genomic level in recombinase-expressing cells was 20.3% for the non-fused heterodimer and 42.7% for the fused heterodimer. Therefore, the fusion of the recombinase heterodimer by the linker of the present invention results in a two-fold higher inversion rate.

[0227] Accordingly, in a further embodiment, the present invention is a method for inverting a DNA sequence at the genomic level in a cell comprising a fusion protein or DNA recombinase for efficient and specific genome editing according to the present invention, i. providing a nucleic acid molecule encoding a first recombinase enzyme, wherein the first recombinase enzyme has been evolved by directed evolution or rational design to specifically recognize a first half-site of a recombinase target site; ii. providing a nucleic acid molecule encoding a second recombinase enzyme, wherein the second recombinase enzyme has been evolved by directed evolution or rational design to specifically recognize a second half-site of a recombinase target site; iii. providing a nucleic acid molecule encoding a linker peptide, wherein the nucleic acid encodes a linker oligopeptide comprising 6 to 30 amino acids; iv. producing an expression vector by cloning a nucleic acid molecule encoding a first recombinase enzyme, a nucleic acid molecule encoding a second recombinase enzyme, and a nucleic acid molecule encoding a linker peptide into the expression vector; v. delivering the expression vector of step iv), an RNA molecule encoding the fusion protein described herein, or the fusion protein described herein to a cell comprising a DNA sequence to be inverted; vi. inverting the DNA sequence on the human chromosome in the cell A method comprising is provided.

[0228] Preferably, the fusion protein of step v. is the fusion protein of the present invention described herein, more preferably the first half-site and the second half-site of the upstream target site and the downstream target site of the recombinase, and most preferably recognizes the upstream target site of SEQ ID NO: 17 and the downstream target site of SEQ ID NO: 18 of the loxF8 recombinase or its reverse complementary sequence.

[0229] In a further embodiment, the method for inversion of a DNA sequence at the genomic level in a cell is v. a) further comprising the step of expressing a fusion protein comprising a first recombinase enzyme, a second recombinase enzyme, and a linker peptide, In step v. of the method, the expression vector of step iv. or the RNA molecule encoding the fusion protein described herein is delivered into the cell, for example, by transfection.

[0230] In one embodiment, the method for inversion of a DNA sequence at the genomic level is performed in a genetically engineered host cell.

[0231] In a preferred embodiment, the method of inverting a DNA sequence at the genomic level is performed in vitro in human cells derived from a patient, more preferably a patient suffering from hemophilia A.

[0232] In a more preferred embodiment, the method of inverting a DNA sequence at the genomic level is performed in vivo in a patient, particularly a patient suffering from hemophilia A.

Examples

[0233] (Examples of the present invention) The following examples are provided only for the purpose of illustrating various embodiments of the present invention and are not intended to limit the present invention in any way.

[0234] Example 1: Substrate-binding directed evolution Recombinases were evolved using the aforementioned substrate-binding protein evolution (SLiDE) (Buchholz and Stewart et al., 2001; Sakara et al., 2007; Karpinski et al., 2016; Lansing et al., 2019). A counterselection strategy was established by varying the selection of active recombinase and inactive recombinase on loxF8 and the symmetric site, respectively.

[0235] Example 2: Comparison of the recombination activities of F8 recombinase monomers and F8 non-fusion recombinase heterodimers To compare whether the heterodimer recombinases consisting of monomers of SEQ ID NOs: 32 and 33 outperform the single F8 monomer recombinase that recognizes the loxF8 sequence, SLiDE using a single recombinase was performed on asymmetric loxF8 sites that are equivalent to the directed evolution of Tre and Brec1 in asymmetric loxLTR and loxBRT sequences, respectively (Sarkar et al., 2007; Hauber et al., 2013). After 168 rounds of directed evolution, the individual recombinases in the pEVO vector were tested in E. coli using different arabinose concentrations. The best clone identified (H7, SEQ ID NO: 66) had a reduced efficiency compared to the D7 heterodimer recombinase (SEQ ID NOs: 32 and 33), but recombined the loxF8 sequence (Figure 6).

[0236] To examine the recombination activities of monomeric (H7) and heterodimeric (D7) F8 recombinases in mammalian cells, HeLa cells were co-transfected with a recombinant reporter and recombinase expression plasmid. Similar to the E. coli assay, the H7 recombinase was less efficient than the heterodimeric D7 recombinase (Figure 7; 28% recombination efficiency vs. 46% recombination efficiency).

[0237] To compare the ability of monomeric F8 recombinase (H7) and heterodimeric F8 recombinase (D7) to induce F8 inversion in human cells, the recombinase expression plasmid was transfected into HEK293 cells. Genomic DNA was isolated 48 hours after transfection. A PCR reaction was performed to detect genomic inversion. Again, the monomeric H7 recombinase did not function as well as the heterodimeric D7 recombinase (Figure 8).

[0238] To compare the specificities of the H7 monomer and the D7 heterodimer, the recombinases were tested on nine human sequences (SEQ ID NOs: 21-23 and 87-92) that showed the highest similarity to the loxF8 sequence (SEQ ID NO: 17) (Table 2). These nine off-target sites and the loxF8 on-target site were cloned into pEVO vectors carrying the respective recombinases, and the plasmids were grown in E. coli for 24 hours in the presence of 10 μg / ml L-arabinose. DNA extracted from these cultures revealed that the H7 monomer recombined the loxF8 site in addition to the off-target sites HS1, HS2, HS4, and HS8 having the sequences of SEQ ID NOs: 21, 22, 87, and 91, respectively. In contrast, the D7 heterodimer was far more specific, only slightly recombining HS2 and showing robust activity at the loxF8 on-target site (Figure 9). Using the same experimental setup, additional highly active heterodimers (A4) were identified. Based on these cumulative results, the heterodimer D7 recombinase and the heterodimer A4 recombinase were selected for further experiments.

[0239] Example 3: Recombinase activity of fusion heterodimers on symmetric target sites Different linkers for fusing the F8 recombinase heterodimers were designed and tested. Flexible linkers having the sequence (Gly-Gly-Ser)n and iteration numbers (n) of 2, 4, 6, and 8 were designed. Subsequent 5' and 3' oligonucleotides were synthesized with the sequences of "sticky ends" generated by XhoI and BsrGI-HF restriction enzymes, and double-stranded DNA fragments directly used for cloning were obtained by annealing of the sequences.

[0240] First, the non-fused heterodimer was tested on the symmetric site. As expected, the test digest revealed strong recombination activity on the symmetric site (Figures 10 and 12B), which can be explained by the activity of the recombinase monomers on each symmetric site. However, importantly, a single recombinase alone shows no activity on the final loxF8 site (Figure 12A). On the other hand, the fused S2-(G2S)8-S1 heterodimer did not show activity on the symmetric site even at a high induction level (200 μg / ml of L-arabinose) (Figure 12C).

[0241] To compare the recombination activity of the fused heterodimer with that of the non-fused heterodimer, both of them were tested at different induction levels. Subsequently, the calculation of the recombination efficacy based on the band intensity was performed. The test digest revealed that the fused heterodimer has an activity equivalent to that of the non-fused heterodimer.

[0242] (G2S) 2-8 As shown from the results of the fusion using the linker, the linker length affects the recombination activity. To study whether a longer linker can enhance the recombination efficiency without showing activity on the symmetric site, the heterodimer was fused with linkers of 10, 12, and 14 (Gly-Gly-Ser) repeats. The results showed that increasing the linker length did not significantly affect the recombination efficiency, but the heterodimer fused with a longer linker showed activity on the symmetric site. Therefore, the recombinase fused with a linker of more than 30 amino acids loses its specificity.

[0243] Example 4: Design of Linker and Preparation of Linker Library Three linker libraries were designed using degenerate codons RVM encoding Ala, Arg, Asn, Asp, Glu, Gly, Lys, Ser, Thr amino acids. A linker length of 24 amino acids was used. The entire sequence of the (G2S)8 linker (24 amino acids), the middle part (12 amino acids) or the edges (6 amino acids from each side) were changed to RVM codons. Different options were considered when designing the library since it was not known that changes to the linker flexibility could affect the activity of the fusion protein. The (G2S) sequence with known flexibility characteristics was retained either far from the recombinase (in the center of the linker) or close to the recombinase (on the edge of the linker) to examine the effect on activity. The first library was designed to keep the central part of the linker sequence (12 amino acids) the same as the first flexible linker used containing four (G2S) repeats, and change the left and right parts of the linker sequence (6 amino acids from both sides) to (RVM)6. The second library was designed conversely: the central part of 12 amino acids was changed to (RVM) 12 and the left and right parts of 6 amino acids were retained as two (G2S) repeats. In addition, since it was not known whether (G2S) repeats were necessary for the linker, in the third library, the entire 24 amino acid sequence of the linker was changed to (RVM) 24 (Figure 13). Three oligonucleotides were designed and used as templates for three PCR reactions. Two adapters for primer annealing and XhoI and BsrGI-HF restriction sites that allowed the direct use of the linker for cloning were adjacent to the linker sequence. The linker library was cloned into the fusion protein instead of (G2S)8, and three libraries, linker library 1, linker library 2, and linker library 3 were obtained (Figure 13). The fusion protein contained the linker library, but the sequence of the recombinase was not changed.

[0244] Oligonucleotides with degenerate base modifications containing XhoI and BsrGI-HF restriction sites, as well as adapters from each end, were synthesized by Biomers.net GmbH. The oligonucleotide sequences are shown in Table 3. The oligonucleotides were used as templates for high-fidelity PCR, followed by digestion of the PCR products with XhoI and BsrGI-HF restriction enzymes and cloning into pEVO_2rec_link between the S2 and S1 specific recombinases. Freshly prepared competent Escherichia coli XL-1 Blue was transformed with plasmids containing heterodimers fused with a library of linkers. After incubation in SOC medium for 1 hour, 2 μl (1:500) of the electroporation mixture was plated on chloramphenicol-containing agar plates, and the remainder was added to 100 ml of LB medium with 25 μg / ml chloramphenicol and the desired concentration of L-arabinose. Both solid and liquid cultures were incubated at 37 °C for 12 - 16 hours, with the latter incubated with constant shaking at 200 rpm. The library size, evaluated by the number of colonies on the plates, was at least 25,000 linker variants.

[0245] Table 3 - Sequences of oligonucleotides used for linker preparation

Table 4

[0246] Example 5: Selection of Linkers Fusion proteins with the linker library were subcloned into the pEVO_2rec_loxF8 plasmid to test their activities on the loxF8 target site. The digest revealed that heterodimers fused with all three linker libraries were active on the loxF8 site at a high induction level of 200 μg / ml L-arabinose (Figure 14A).

[0247] To select a heterodimer fused with a linker having recombination activity at the loxF8 site, first, digestion was performed with the NdeI and AvrII restriction enzymes. The restriction sites of these enzymes are located between the two loxF8 sites on the pEVO_2rec_LoxF8 plasmid, and thus, only the unrecombined plasmid became linear. Subsequently, high-fidelity PCR was performed using a forward primer (primer 4) that anneals upstream of the fusion protein and a reverse primer (primer 19) that anneals downstream of the second loxF8 site. As a result, the unrecombined plasmid was digested and amplification was inhibited, so only the plasmid-recombined fusion protein was amplified by PCR (Figure 14B).

[0248] Subsequently, to test the activity at the Sym1 site, the obtained PCR product was digested with the SacI-HF and SbfI-HF restriction enzymes and subcloned into the pEVO_2rec_Sym1 plasmid. The test digest revealed that only one of the libraries, S2-link lib3-S1, had activity at the Sym1 site at a high induction level of 200 μg / ml of L-arabinose (Figure 15A). The activities of the heterodimers fused with linker libraries 1 and 2 were not detected by the test digest. To select a fusion heterodimer that does not recombine the Sym1 site, high-fidelity PCR was performed using a forward primer (primer 20) that anneals upstream of the fusion protein and a reverse primer (primer 21) that anneals between the two Sym1 sites (Figure 15B). Thus, only the fusion protein that did not recombine the plasmid was amplified by PCR. To test the activity at the Sym2 site, the obtained PCR product was digested with the SacI-HF and SbfI-HF restriction enzymes and subcloned into the pEVO_2rec_Sym2 plasmid. At a high induction level of 200 μg / ml of L-arabinose, strong activity was not detected at the Sym2 site by the test digest, and selection was performed in the same manner as for the Sym1 site (Figure 15).

[0249] The same procedure was repeated to select linkers with the best properties, and the following selection rounds were performed on loxF8 at 10 μg / ml L-arabinose, Sym1 at 200 μg / ml L-arabinose, Sym2 at 200 μg / ml L-arabinose, loxF8 at 1 μg / ml L-arabinose, Sym1 at 200 μg / ml L-arabinose, Sym2 at 200 μg / ml L-arabinose, loxF8 at 1 μg / ml L-arabinose. During selection at the loxF8 site, low induction levels were used to select fusion proteins with high activity on the final target site. However, during selection at the symmetric sites, the L-arabinose concentration was kept high at 200 μg / ml to select fusion proteins that were inactive at the symmetric sites. After 10 selection rounds, all three fusion proteins with the linker library showed activity on the loxF8 site at the low induction level of 1 μg / ml L-arabinose. However, the activity of the library on the Sym1 and Sym2 sites at the high induction level of 200 μg / ml L-arabinose also increased (Figure 16). The heterodimers fused with all three linker libraries had activity on the Sym2 site, and the heterodimers fused with linker libraries 1 and 3 had activity on the Sym1 site. The increase in activity on the symmetric sites is likely due to the accumulation of active heterodimers in the library.

[0250] Since the library showed activity at a very low induction level on the loxF8 site, single clones were analyzed at this point. The library had some activity on the symmetric sites, but single binding recombinases from those libraries were expected to recombine only loxF8, not Sym1 or Sym2.

[0251] After testing the library on the loxF8 site with 1 μg / ml of L-arabinose, digestion was performed with NdeI and AvrII restriction enzymes, and the undigested plasmids were transformed again. This enabled enrichment of the library for fusion proteins that are active at low induction levels. Next, the fusion proteins with the linker library were subcloned again into the pEVO_2rec_LoxF8 plasmid and grown on plates. Thirty-two colonies from each library were picked, induced with 10 μg / ml of L-arabinose, and tested for recombination using a 3-primer PCR assay. Ninety clones showed recombination activity, and among them, four colonies from each library were further analyzed for activity on the loxF8 on-target (10 μg / ml of L-arabinose) as well as the Sym1 and Sym2 off-targets (200 μg / ml of L-arabinose). All of the clones (L1 - L12) showed recombination activity on the loxF8 site, but the efficiency varied among the clones (Figure 17A). Two clones from library 1 and three clones from library 3 had activity at high induction levels on the Sym1 site (Figure 17B). Two clones from library 2 and all four clones from library 3 had measurable activity at high induction levels on the Sym2 site (Figure 17C). It is notable that none of the clones from library 2 showed activity on the symmetric sites. Thus, it can be inferred that the design of library 2 was more beneficial in preventing activity on the symmetric target sites. The selected clones were sequenced, and the sequencing results demonstrated that the library strategy had functioned (Table 4).

[0252] Table 4: Sequences of single clones from the linker library

Table 5

[0253] Based on the recombination efficiency at the loxF8 site and the absence of activity on the symmetric sites, clone L8 (S2-linker L8-S1) was selected for further analysis.

[0254] Example 6: Recombinase Activity of Recombinase Fusion with L8 Linker Recombinase fusions using the L8 linker were tested at different induction levels on the loxF8 site in parallel with the linker and non-fused heterodimer of the (G2S)8 linker. The test digest revealed that the L8 linker fusion recombined the plasmid more efficiently (Figure 18). Similarly, the L8 linker fusion heterodimer also demonstrated higher recombination efficiency compared to the non-fused heterodimer (Figure 19). Thus, the fusion of the D7 monomer with the L8 linker sequence increases recombination activity.

[0255] Example 7: Off-Target Activity of Fusion Heterodimer A further goal was to examine whether the fusion of recombinase heterodimers increases the specificity of recombination and as a result minimizes potential off-target effects. Based on the remote similarity to the loxF8 sequence, human sequences showing high similarity to the loxF8 sequence, similar to potential asymmetric and symmetric off-target sites for the F8 recombinase heterodimer, were tested in an E. coli excision assay (Table 2, SEQ ID NOs: 19-29).

[0256] Activity on potential off-target sites was tested at high induction levels to identify even weaker recombination activity. To account for the different efficiencies of potential on-target recombination on loxF8, the concentration of L-arabinose was chosen to be 5-fold higher for the fusion L8 heterodimer. Thus, L-arabinose concentrations of 50 μg / ml (for non-fused) and 250 μg / ml (for L8) were used to test heterodimer activity on five asymmetric and four symmetric off-targets.

[0257] For the D7 heterodimer, the test digest revealed that the non-fused heterodimer showed significant recombination activity at one of the five asymmetric off-targets (2LR site, SEQ ID NO: 22 site) tested at high induction levels. In contrast, the fused L8 heterodimer of the D7 recombinase did not show any obvious recombination activity on any of the asymmetric off-target sites (Figure 20A).

[0258] On the symmetric off-target site, the non-fused D7 heterodimer showed significant activity at one of the test sites (2L site, SEQ ID NO: 27 site). Conversely, the fused D7 heterodimer did not show activity against any of these or other test sites (Figure 20A). This result indicates that the fusion of the heterodimer significantly increases the recombination specificity of the designer SSR.

[0259] For the A4 heterodimer, the test digest using the non-fused heterodimer did not show any significant recombination activity against the five asymmetric off-targets tested. The same was true for the fused L8 heterodimer of the A4 recombinase (Figure 20B).

[0260] On the symmetric off-target site, the non-fused A4 heterodimer did not show activity at any of the test sites. The fused A4 heterodimer also did not show activity at any of these or other test sites (Figure 20B).

[0261] Example 8: Recombination Activity in Human Cells The fused L8-recombinase heterodimer showed good activity in bacteria. Therefore, the next step was to examine the recombination efficiency in human cells and compare the activity with that of the non-fused heterodimer. HEK293T cells were transfected with an expression plasmid containing a fusion recombinase or a single recombinase transcriptionally fused to EGFP (Figure 21). Cells transfected with an empty plasmid expressing only EGFP and non-transfected cells were used as positive and negative controls for transfection, respectively.

[0262] Cells were also transfected with a reporter plasmid encoding mCherry with an upstream stop codon located between two loxF8 target sites together with the recombinase expression plasmid. Upon recombination, the reporter plasmid excises the stop cassette, thereby enabling the expression of mCherry, which can be measured by flow cytometry analysis (Figure 22).

[0263] Flow cytometry analysis performed 48 hours after transfection with the D7 recombinase heterodimer (containing the recombinase monomers of SEQ ID NOs: 32 and 33; Figure 23) revealed that the transfection efficiency of cells transfected with both the expression plasmid and the reporter plasmid was 28% for the non-fused heterodimer and 21% for the fused heterodimer, respectively. GFP and mCherry double-positive cells reflect the recombination of the reporter in human cells. To calculate the recombination efficiency of the reporter plasmid in human cells, double-positive cells were normalized to the transfection efficiency (Figure 23).

[0264] Thus, it was shown that the reporter recombined in 76.4% (non-fused) or 92.8% (fused) of the cells, and that the fused L8 recombinase had a higher recombination efficiency in this assay.

[0265] Example 9: Inversion Efficiency at the Genomic Level in Human Cells The recombinases described in this specification were developed to correct a large gene inversion in exon 1 of the F8 gene that causes hemophilia A. To study the inversion efficiency of heterodimers at the genomic level, the described expression construct (Figure 21) was used to express the recombinase in HEK293T cells for 48 hours, after which gDNA was extracted. HEK293T cells do not carry the inverted exon 1 of the F8 gene. However, the recombinase can perform the inversion reaction regardless of the orientation of the genomic DNA fragment between the loxF8 sites. Therefore, the ratio of the inverted exon 1 to the non-inverted exon 1 can be used as a substitute for the inversion efficiency. To calculate the inversion frequency on genomic DNA, a qPCR-based assay using a TaqMan probe specific for exon 1 inversion was developed. A standard curve was calculated using gDNA samples with defined amounts of exon 1 inversion. The inversion efficiency of recombinase-transfected cells can be estimated later using this standard curve. gDNA isolated from samples transfected with non-fused heterodimers showed that inversion was induced in approximately 6.0% of the cells. In samples transfected with fused heterodimers, the inversion frequency was approximately 9.0% (Figure 24). However, since the transfection efficiency is not 100% and GFP expression is transcriptionally linked to recombinase expression, it is necessary to correct the inversion rate of GFP+ cells. Flow cytometry data revealed transfection efficiencies of 29.5% (non-fused) and 21.3% (fused), respectively. Therefore, the genomic-level inversion efficiency in recombinase-expressing cells was 20.3% for non-fused heterodimers and 42.7% for fused L8 heterodimers. Thus, the fusion of the L8 recombinase heterodimer resulted in a two-fold higher inversion rate, demonstrating that the fusion of the recombinase heterodimer increases the recombination activity.

[0266] Example 10: Recombinase Activity in an Integrated Genomic loxF8 Reporter Construct Since recombinases were developed to recombine sequences found in the human genome, a genomic loxF8 reporter cell line was developed (Figure 25A). Upon recombination of the loxF8 target site of the integrated reporter, the cells begin to express the red fluorescent protein mCherry. Recombinases were transfected into the reporter cell line as synthetic mRNA (Figure 25B). The blue fluorescent protein (tagBFP) encoding the synthetic mRNA was co-transfected to estimate the transfection efficiency. Flow cytometry analysis performed 48 hours after transfection with D7 recombinase (non-fused and fused - containing the recombinase monomers of SEQ ID NOs: 32 and 33) and A4 recombinase (non-fused and fused - containing the recombinase monomers of SEQ ID NOs: 93 and 94) along with tagBFP mRNA revealed that the transfection efficiency was 44.5% for D7 non-fused, 48.9% for D7 fused, 55.5% for A4 non-fused, and 55.7% for A4 fused. BFP and mCherry double-positive cells reflect the recombination of the genomic reporter in the human reporter cell line. To calculate the recombination efficiency of the genomic reporter, double-positive (BFP+ and mCherry+) cells were normalized to the transfection efficiency (Figure 25C). Thus, 71.5% of the genomic reporter recombined for D7 non-fused, 50.7% for D7 fused, 24.3% for A4 non-fused, and 21.4% for A4 fused.

[0267] Example 11: Detection of loxF8 Genomic Inversion at the Native loxF8 Locus by PCR HEK293T cells were transfected with synthetic mRNAs encoding recombinase heterodimers of D7 non-fusion, D7 fusion, A4 non-fusion, A4 fusion, and D11 non-fusion (including recombinase monomers of SEQ ID NOs: 32, 33, 93, 94, 99, and 100). Genomic DNA was extracted 48 hours after transfection and analyzed for inversion of the 140 kb native loxF8 locus (Figure 26A). The orientation of the loxF8 locus can be detected by using different combinations of primers for PCR (Figure 26A). Primers P2 and P3 were used to detect the normal orientation of the loxF8 locus, and primers P1 and P3 were used to detect the reverse orientation of the loxF8 locus (Figure 26B). Only cells transfected with synthetic mRNAs encoding recombinase heterodimers showed an inversion-specific band (PCR P1+P3). No bands were visible in the WT control or water control samples. To confirm the integrity of the genomic DNA used in this experiment, normal-orientation PCR (P2+P3) was performed in the same manner. All samples showed PCR bands of the expected size, indicating that the genomic DNA used as a template for all PCR reactions was intact. As expected, the inversion control sample did not show a band in this reaction. These results indicate that recombinase heterodimers of SEQ ID NOs: 32 and 33 (non-fusion D7 recombinase), SEQ ID NOs: 93 and 94 (non-fusion A4 recombinase), and SEQ ID NOs: 99 and 100 (non-fusion D11 recombinase) and fusion proteins of SEQ ID NOs: 72 (fusion D7 recombinase) and SEQ ID NOs: 97 (fusion A4 recombinase) can invert the 140 kb native loxF8 locus.

[0268] References

Table 6

Claims

**Claim 1**: A fusion protein for genome editing, wherein the fusion protein comprises a complex of heterospecific recombinases, the complex comprises at least a first recombinase enzyme, a second recombinase enzyme and at least one linker, the first recombinase enzyme and the second recombinase enzyme specifically recognize the first half-site and the second half-site of the upstream target site and / or the downstream target site of the recombinase; the first recombinase enzyme and the second recombinase enzyme are bound to each other via a linker, the linker has the formula 2: (G 2 S) 2 -X 1 -X 2 -X 3 -X 4 -X 5 -X 6 -X 7 -X 8 -X 9 -X 10 -X 11 -X 12 -(G 2 S) 2 (Formula 2) (wherein G is glycine; S is serine; and X 1 ~X 12 Each of them is independently selected from the group consisting of alanine, arginine, asparagine, aspartic acid, glutamine, glycine, lysine, serine and threonine; the oligopeptide of formula 2 does not consist of only glycine and serine residues) and consists of an oligopeptide having an amino acid sequence of the C-terminus of the first recombinase enzyme is bound to the N-terminus of the second recombinase enzyme via the linker, the fusion protein. **Claim 2** The fusion protein is a heterodimer, and each recombinase enzyme of the heterodimer is a tyrosine site-specific recombinase that independently recognizes the first half-site and the second half-site of the recombinase target site of the tyrosine site-specific recombinase. The fusion protein according to claim 1. **Claim 3** The specific recombinase is selected from the group consisting of Cre-, Dre-, VCre-, SCre-, Vika-, lambda-Int-, Flp-, R-, Kw-, Kd-, B2-, B3-, Nigri- and Panto-recombinases. The fusion protein according to claim 2. **Claim 4** The linker is 【Chemical 1】 selected from the group consisting of. The fusion protein according to any one of claims 1 to 3. **Claim 5** The first recombinase enzyme is a protein having an amino acid sequence having at least 90% sequence identity with the sequence set forth in SEQ ID NO: 30, SEQ ID NO: 32, SEQ ID NO: 93 or SEQ ID NO: 99, and / or the second recombinase enzyme is a protein having an amino acid sequence having at least 90% sequence identity with the sequence set forth in SEQ ID NO: 31, SEQ ID NO: 33, SEQ ID NO: 94 or SEQ ID NO:

100. The fusion protein according to any one of claims 1 to 4. **Claim 6** In the fusion protein, the C-terminus of the second recombinase enzyme that specifically recognizes the second half-site of the recombinase target site is fused to the N-terminus of the first recombinase enzyme that specifically recognizes the first half-site of the recombinase target site using the linker, whether the first recombinase enzyme, the second recombinase enzyme, and the linker are bound to each other; or The fusion protein according to any one of claims 1 to 5, wherein in the fusion protein, the C-terminus of the first recombinase enzyme that specifically recognizes the first half-site of the recombinase target site is fused to the N-terminus of the second recombinase enzyme that specifically recognizes the second half-site of the recombinase target site using the linker, and the first recombinase enzyme, the second recombinase enzyme, and the linker are bound to each other.

7. The fusion protein has an amino acid sequence having at least 90% sequence identity with the sequence set forth in SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 97, or SEQ ID NO: 103, The fusion protein according to any one of claims 1 to 6, wherein each of SEQ ID NO: 71, SEQ ID NO: 72, SEQ ID NO: 97, and SEQ ID NO: 103 contains a linker having the amino acid sequence of SEQ ID NO:

14.

8. A nucleic acid molecule encoding the fusion protein according to any one of claims 1 to 7.

9. i. A nucleic acid encoding a recombinase by the nucleic acid sequence of SEQ ID NO: 34 and / or 35 and a linker by the nucleic acid sequence of SEQ ID NO: 61; or ii. A nucleic acid encoding a recombinase by the nucleic acid sequence of SEQ ID NO: 36 and / or 37 and a linker by the nucleic acid sequence of SEQ ID NO: 61; or iii. A nucleic acid encoding a recombinase by the nucleic acid sequence of SEQ ID NO: 95 and / or 96 and a linker by the nucleic acid sequence of SEQ ID NO: 61; or iv. A nucleic acid encoding a recombinase by the nucleic acid sequence of SEQ ID NO: 101 and / or 102 and a linker by the nucleic acid sequence of SEQ ID NO: 61; or v. A nucleic acid having the sequences of SEQ ID NO: 95 and 96, each encoding a recombinase; or vi. A nucleic acid having the sequences of SEQ ID NO: 101 and 102, each encoding a recombinase; or vii. A nucleic acid having the sequences of SEQ ID NO: 34 and 35, each encoding a recombinase; or viii. A nucleic acid having the sequences of SEQ ID NO: 36 and 37 encoding recombinases respectively The nucleic acid molecule according to claim 8, comprising the same.

10. A polynucleotide molecule comprising the nucleic acid molecule according to claim 8 or 9, and an expression control element operably linked to the nucleic acid to promote its expression.

11. A mammalian cell, insect cell, plant cell or bacterial host cell comprising the nucleic acid molecule according to claim 8 or 9, or the recombinant polynucleotide molecule according to claim 10, or an expression vector comprising the nucleic acid molecule according to claim 8 or 9.

12. A pharmaceutical composition comprising the fusion protein according to any one of claims 1 to 7, or the nucleic acid molecule or recombinant polynucleotide according to any one of claims 8 to 10.

13. The pharmaceutical composition according to claim 12, for use in the treatment of hemophilia A, particularly severe hemophilia A.

14. A method for in vitro inversion of a DNA sequence at the genomic level in a cell, comprising: a) a.i) providing a nucleic acid molecule according to claim 8 or 9 encoding a fusion protein; a.ii) preparing an expression vector by cloning the nucleic acid molecule into an expression vector; a.iii) delivering the expression vector of step a.ii) to a cell containing the DNA sequence to be inverted; a.iv) inverting the DNA sequence on the human chromosome in the cell or b) b.i) providing an RNA molecule encoding a fusion protein according to any one of claims 1 to 7 evolved by directed evolution or rational design; b.ii) delivering the RNA molecule of step b.i) to a cell containing the DNA sequence to be inverted; b.iii) inverting the DNA sequence on the human chromosome in the cell or c) c.i) providing a fusion protein according to any one of claims 1 to 7 evolved by directed evolution or rational design; c.ii) delivering the fusion protein of step c.i) to a cell containing the DNA sequence to be inverted; c.iii) inverting the DNA sequence on the human chromosome in the cell The method as described above.

15. (i) The method further comprises, after step a.iii), expressing the fusion protein comprising a first recombinase enzyme, a second recombinase enzyme and a linker peptide, or (ii) The method according to claim 14, wherein in step a.iii), the expression vector or in step b.ii), the RNA molecule encoding the fusion protein is delivered to the cell.

Citation Information

Patent Citations

  • Substrate-linked directed evolution (slide)

    JP2004518419A

  • A novel method for the integration of foreign DNA into eukaryoticgenomes

    WO1999025840A1

  • Methods and means for genetic alteration of genomes utilizing designer DNA recombining enzymes

    WO2018229226A1