Novel recombinase enzymes for site-specific dna-recombination
Patent Information
- Application Number
- EP2024720781
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-17
- Filing Date
- 2024-04-17
- Publication Date
- 2026-02-25
AI Technical Summary
Current site-specific recombinase systems are limited in their applicability across various cell types due to off-target activity and genomic specificity, necessitating the development of novel recombinase systems for precise genetic manipulation with high activity and low toxicity.
The method involves using proteins with recombinase activity that have amino acid sequences with at least 80% identity to specific sequences (SEQ ID NOs) to catalyze site-specific DNA recombination at recognition sites that are essentially identical or reverse complementary, allowing for precise genetic manipulation in a variety of cell types.
This approach enables efficient and specific site-specific DNA recombination with minimal cell proliferation inhibition, expanding the utility of recombinase systems in genetic manipulation across different cell types and organisms.
Smart Images

Figure IMGF000013_0001 
Figure 00000050_0000 
Figure 00000050_0001
Abstract
Description
[0001] Novel recombinase enzymes for site-specific DNA-recombination
[0002] The present invention relates to the use of a protein with recombinase activity to catalyze a site-specific DNA recombination, as well as to a method for producing a site-specific DNA recombination. The present invention is applicable alone or in combination with other recombinase systems for genetic manipulation, for example in medicine or medical research.
[0003] Background of the invention
[0004] Site-specific recombinases (SSRs) are reliable tools for targeted modification of genomes, with a variety of applications in research, medicine, and biotechnology. SSRs are divided into two evolutionarily and mechanistically distinct families of enzymes: the tyrosine and the serine recombinases (Meinke et al., 2016). Nevertheless, SSRs can catalyze both cleavage and immediate resealing of DNA strands without the help of additional proteins, unlike DNA-altering systems such as CRISPR-Cas and other nuclease-based technologies (Meinke et al., 2016). Although nuclease-based approaches are being intensively developed to expand their utility (z.e., base editors and prime editors (Komor et al., 2016; Gaudelli et al., 2017; Anzalone et al., 2022; Anzalone et al., 2019), these systems still typically introduce DNA-nicks and rely on cell intrinsic repair mechanisms (Meinke et al., 2016). In contrast, SSRs can operate autonomously and the outcome of recombination is very predictable. Furthermore, most SSRs are relatively small in size, rendering them more convenient for different delivery vectors (Meinke et al., 2016).
[0005] Tyrosine recombinases (Y-SSRs) such as Cre and Flp are used extensively for genome engineering due to their simplicity and their ability to conduct efficient genome modifications in heterologous hosts (Sauer et al., 1988). Hence, these SSRs were extensively used to model and understand how these types of enzymes work. The Cre / loxP system is derived from bacteriophage Pl and consists of the recombinase Cre and the 34-bp target sites loxP (Sternberg et al., 1981). loxP sites are palindromic sequences with two 13-bp inverted repeats separated by an 8-bp spacer. Each half-site is bound by one recombinase protomer forming a tetrameric complex on two loxP sites, which together catalyze strand exchange within the spacer region (Duyne and Hamilton, 1981). Depending on the relative orientation of the spacers, Cre is able to perform a variety of reactions such as excision, integration, inversion and translocations (Meinke et al., 2016). Furthermore, when presented with two heterospecific target sites in the genome, Cre alone, or in combination with other SSRs, can perform recombinase-mediated cassette exchange (RMCE) for precise replacement of a DNA fragment (Meinke et al., 2016; Anderson et al., 2012; Minorikawa and Nakayama, 2011).
[0006] In addition to the Cre / loxP and the Flp / FRT system, that are both the most widely used sitespecific recombinase systems of tyrosine class, other recombinase systems are known in the art. US 7,422,889 and US 7,915,037 disclose the so-called Dre / rox system that comprises a Dre recombinase isolated from Enterobacteria phage D6, the recognition site of which is called rox-site. Further known recombinase systems are the VCre / VloxP system isolated from Vibrio plasmid p0908, and the sCre / SloxP system (WO 2010 / 143606 Al; Suzuki and Nakayama, 2011).
[0007] Further site-specific DNA recombinase systems are the Nigri / nox system disclosed in EP 2 877 585 Bl, the Vika / vox system disclosed in EP 2 690 177 Bl, and the Panto / pox system disclosed in EP 3 263 708 Bl.
[0008] The recombinase systems that are known in the art show different activities in cells of different origin. For many applications, such as the production of transgenic animals with conditional gene knockouts, two or more recombinase systems are used in combination with each other. However, emerging complex genetics studies and applications require simultaneous use of multiple recombinases. At the same time, not all well-described site-specific recombinases are equally applicable in all model organisms due to, e.g., genomic specificity (off-target activity on cryptic recognition target sites). For that reason, it is important that an optimal recombinase can be chosen depending on the target organism or experimental setup. Therefore, there is a need in the art for the provision of additional specific recombinase systems that can be used to catalyze a site-specific DNA recombination on short targets in a variety of cell types and with high activity and with low toxicity.
[0009] It is therefore an objective of the present invention to provide novel recombinase systems for site-specific genetic recombination, which can be used in a variety of cell types. Another object of the present invention is the provision of a novel, highly specific recombinase system for site-specific genetic recombination with minor or essentially no inhibition of the cell’s proliferation.
[0010] Summary of the invention
[0011] The objective underlying the present invention is solved by the provision of a method for producing a site-specific DNA-recombination. According to the present invention, the method for producing a site-specific DNA-recombination comprises the steps of a) contacting a nucleic acid comprising at least a first and a second recognition site which are essentially identical or essentially reverse complementary to each other, with a protein having recombinase activity, and b) allowing the protein having recombinase activity to produce the site-specific DNA-recombination. In accordance with said method, a recognition site comprises a first half-site, a spacer and a second half-site, and essentially identical or essentially reverse complementary to each other means that the nucleotide sequence of the first and the second half-site in the first recognition site may deviate in up to two nucleotides from the nucleotide sequence of the first and the second half-site in the second recognition site.
[0012] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 7, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 16 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 16.
[0013] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 3, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 12 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 12.
[0014] According to one aspect of the method of the present invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 1, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 10 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 10.
[0015] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 2, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 11 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 11.
[0016] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 4, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 13 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 13.
[0017] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 5, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 14 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 14.
[0018] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 6, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15.
[0019] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 8, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 17 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 17. According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 9, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15.
[0020] According to one embodiment, the protein having recombinase activity comprises at least two protein monomers.
[0021] According to one embodiment, the nucleic acid sequence that is recombined is present in a cell.
[0022] According to a preferred embodiment, the method further comprises the step of introducing into the cell a nucleic acid encoding the protein having recombinase activity.
[0023] According to another embodiment, the cell comprises a nucleic acid encoding the protein having recombinase activity.
[0024] According to a preferred embodiment, the nucleic acid encoding the protein having recombinase activity comprises a regulatory nucleic acid sequence, and wherein the expression of the nucleic acid encoding the protein having recombinase activity is regulated by the regulatory nucleic acid sequence.
[0025] According to a further embodiment, the cell is a eukaryotic or a bacterial cell.
[0026] According to a further aspect, the present invention provides the use of a protein having at least 80% identity to SEQ ID NO: 7, SEQ ID NO: 3, SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 8, or SEQ ID NO: 9 for producing a site-specific DNA- recombination.
[0027] According to a yet a further aspect, the present invention provides the use of a protein having recombinase activity for catalyzing a site-specific DNA-recombination at recognition sites that essentially are identical or essentially reverse complementary to each other, wherein a recognition site comprises a first half-site, a spacer and a second half-site, and wherein essentially identical or essentially reverse complementary to each other means that the nucleotide sequence of the first and the second half-site in a first recognition site may deviate in up to two nucleotides from the nucleotide sequence of the first and the second half-site in a second recognition site.
[0028] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 7, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 16 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 16.
[0029] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 3, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 12 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 12.
[0030] According to one aspect of the use of the present invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 1, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 10 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 10.
[0031] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 2, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 11 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 11.
[0032] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 4, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 13 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 13.
[0033] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 5, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 14 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 14.
[0034] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 6, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15.
[0035] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 8, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 17 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 17.
[0036] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 9, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15. According to a further aspect, the present invention provides a nucleic acid having a length of not more than 40 base pairs and comprising:
[0037] (i) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 16 or a nucleic acid sequence reverse complementary thereto; or
[0038] (ii) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 12 or a nucleic acid sequence reverse complementary thereto; or
[0039] (iii) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 10 or a nucleic acid sequence reverse complementary thereto; or
[0040] (iv) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 11 or a nucleic acid sequence reverse complementary thereto; or
[0041] (v) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 13 or a nucleic acid sequence reverse complementary thereto; or
[0042] (vi) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 14 or a nucleic acid sequence reverse complementary thereto; or
[0043] (vii) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 15 or a nucleic acid sequence reverse complementary thereto; or
[0044] (viii) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 17 or a nucleic acid sequence reverse complementary thereto.
[0045] According to a further aspect, the present invention provides a vector comprising at least one and preferably at least two essentially identical or essentially reverse complementary nucleic acids of the present invention, wherein a DNA segment is preferably flanked by the two essentially identical or essentially reverse complementary nucleic acids.
[0046] According to one embodiment, the vector further comprises a nucleic acid encoding a protein having recombinase activity, wherein the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to any one of SEQ ID NOs: 1 to 9.
[0047] According to a further aspect, the present invention provides a vector comprising a nucleic acid encoding a protein having recombinase activity, wherein the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to any one of SEQ ID NOs: 1 to 9.
[0048] According to a preferred aspect of the present invention, there is provided the use of the vector of the present invention in a method of the present invention.
[0049] According to a yet another aspect, the present invention provides a protein having at least 80% identity to SEQ ID NO: 7, SEQ ID NO: 3, SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 8 or SEQ ID NO: 9, or a vector according to the invention, for use in medicine. Preferably, the protein or the vector is for use in treating a genetic disease or disorder in a subject. More preferably, the genetic disease or disorder is characterized by modification of the subject’s genome. According to a further aspect, the present invention provides an isolated host cell, comprising the following recombinant DNA fragments: at least one and preferably at least two nucleic acids according to the invention; and / or a vector according to the invention.
[0050] According to one embodiment, the isolated host cell further comprises (i) a nucleic acid encoding for a protein having recombinase activity, wherein the protein comprises an amino acid sequence having at least 80% identity to any one of SEQ ID NOs: 1 to 9; or (ii) a vector according to the invention.
[0051] According to a further aspect, the present invention provides a non-human host organism, comprising: (i) at least one and preferably at least two nucleic acids according to the invention; or (ii) a vector according to the invention.
[0052] According to one aspect, the present invention provides a pharmaceutical composition comprising a protein having at least 80% identity to SEQ ID NO: 7, SEQ ID NO: 3, SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 8, or SEQ ID NO: 9, the nucleic acid according to the invention, a vector according to the invention, the isolated host cell according to the invention, or the non-human host organism according to the invention, and optionally a pharmaceutically acceptable excipient
[0053] Further aspects and embodiments of the invention will become apparent from the appending claims and the following detailed description.
[0054] Description of the drawings
[0055] The invention is further illustrated by the following figures and examples without being limited thereto.
[0056] Fig. 1A shows an overview of the plasmid recombination assay. Important features, such as restriction sites, recombinase coding sequence and target sites (triangles) are shown. A schematic representation of expected recombination products on an agarose gel is also shown. Marker and expected sizes of recombined and unrecombined plasmids are indicated separately or together as usually seen on the gels. Fig. IB shows the plasmid map of pEVO recombination reporter. Protein coding genes, the origin of replication (oriP15A) and the pBAD promoter are depicted as arrows. The protein coding genes include the chloramphenicol resistance gene (cmR), the arabinose regulatory protein (araC) and the genes encoding for the recombinases of interest. Expression of the recombinase is driven by pBAD promoter upon addition of arabinose. Recombination between two lox sites leads to excision of about 700 bp stuffer sequence. BsrGI and Sbfl restriction enzymes are used for cloning of the recombinases as well as for linearization of the plasmids for test digest
[0057] Fig. 2 shows the recombination activity of seventeen tested putative Y-SSR / target site pairs. Each sample was tested with or without L-arabinose (100 pg / ml) added to the growth medium for recombinase expression, indicated with or "+". Recombination is indicated by the band aligned with the single triangle and non-recombined plasmids are indicated by two triangles. Active recombinases are highlighted with a box. M = GeneRuler™ DNA Ladder Mix (Thermo Fisher).
[0058] Fig. 3 shows quantification and reproducibility of recombination activity of new SSRs in E. coli. Recombinase expression was induced with rising concentrations (pg / ml) of L-Arabinose indicated along the x-axis. Vika and Cre were included as positive controls. Recombination was calculated from measuring the ratio of band intensities of unrecombined and recombined bands detected on agarose gels. Bacterial assays were done in triplicates (n = 3).
[0059] Fig. 4A shows the analysis of deep sequencing results. Cross recombination events are displayed by a heatmap of recombination percentage for each possible combination. Target sites are displayed horizontally and ordered based on their similarity. Recombinases tested are aligned based on their homology on the vertical axis. On-target recombination events are boxed with red squares. Fig. 4B shows the validation of cross recombination events by plasmid-based recombination assay. Recombination activity of respective recombinases is shown on their on-target sites, and on off- targets. Recombination was assessed by agarose gel electrophoresis. Each sample was tested with or without L-arabinose to induce recombinase expression, indicated with or "+" (100 pg / ml of L- arabinose). Recombination is indicated by the band aligned with the single triangle and nonrecombined band is indicated by two triangles. M = GeneRulerTM DNA Ladder Mix (Thermo Fisher). Parts of the figure were created with BioRender.com.
[0060] Fig. 5 shows the activity of new recombinases in mammalian cells. Fig. 5A is a graphical representation of the mammalian recombination reporter and expression constructs. Important features are marked in the reporter and expression vectors. Upon recombination of the reporter vector the mCherry cassette will be excised allowing for the expression of GFP (green) from the pCAG promotor (arrow). Black triangles represent different lox sites for each corresponding recombinase. NLS, nuclear localization signal. Fig. 5B shows fluorescence microscopy analysis in Hek293T cells. Transfected cells with the empty expression plasmids and non-recombined reporter plasmids harboring eight new target sites (top panel), or co-transfection of the reporter with the expression plasmids carrying respective recombinases (lower panel) are shown. Vika / vox and Cre / loxP were included as positive controls. Ctrl, negative controls; Rec, recombinase. Fig. 5C shows FACS analysis of the samples shown in Fig. 5B. Darker grey histograms depict control samples where non- recombined reporter plasmids were co-transfected with ‘empty’ expression plasmids, while the lighter gray histograms show samples transfected with corresponding recombinases. Fig. 5D shows the plasmid maps of expression lentiviral vector (left) and reporter vector (right). CMV, PGK and CAG promoters for mammalian expression of viral RNA, Recombinase-P2A-BFP cassette and mCherry cassette, respectively, are shown as white arrows. Features for viral production on pLentiX vector are also depicted. The lentiviral vector was either transfected into HEK293T cells for recombination assay, or used for virus production and infection as to test the effect of continuous expression of the recombinases. BsrGI and Xbal restriction sites used for cloning of the recombinases are marked on the expression vector while Nhel and Hindlll sites labeled on pCAG-lox-mCherry-lox-GFP reporter plasmid were used for cloning of all the target sites. Ori - ColEl origin of replication in bacteria; AmpR - ampicillin resistance gene; bGH poly (A) signal - bovine growth hormone polyadenylation signal; rbGlob-polyA - rabbit P-globin polyadenylation signal.
[0061] Fig. 6 shows the effect on cell growth upon overexpression of the recombinases of the present invention in mammalian cells. Fig. 6A is an overview of the experimental setup. Important steps are indicated by arrows. Cells were transduced at a rate of ca. 50% with a bicistronic lentivirus expression construct, where the expression of respective recombinases was linked via P2A to BFP expression. Cells were analyzed every 72 h by flow cytometry and the percentage of BFP-positive cells was recorded over the course of 2 weeks. Declining percentages of BFP-positive cells are indicative of a proliferation disadvantage of infected cells. Fig. 6B shows the analysis of growth rates. Difference in the percentage of BFP-positive cells between day 3 and day 15 are plotted (biological replicates are shown as dots, n = 3). Error bars represent standard deviation of the mean (SD). Statistical significance relative to BFP control was calculated by a 1-way ANOVA test. (****): P < 0.0001.
[0062] Fig. 7 shows the results of directed evolution of YR9 recombinase Fig. 7A. Plasmid based activity assay of wt YR9 recombinase and three randomly picked clones from the evolved library. Agarose gel pictures to the right show the best clones recombining lox9 in comparison to wt YR9 at four different expression levels (0 pg / ml, 1 pg / ml, 10 pg / ml and 100 pg / ml). To the right, quantification and reproducibility of recombination is shown. Recombination efficiency was calculated by comparing the ratio of recombined and non-recombined band intensities from the gels to the left. Experiments were done in triplicates (n =3). Comparison to YR9 wt was done with a t-test and the p-values were adjusted for multiple comparisons using the Bonferroni method. Significance: (ns) p > 0.05, (*) p < 0.05, (**) p < 0.01, (***) p < 0.001, (****) p < 0.0001 Fig. 7B. Mutation analysis of the YR9 improved clones (one-letter code). The amino acid sequence of wt YR9 recombinase is shown as a reference. Letters represent changes found in the clones YR9.10, YR9.2 and YR9.1 from top to bottom. Predicted C-terminal catalytic domain of DNA breaking-rejoining super family is labeled. Fig. 7C. Activity of the most active clone YR9.10 in mammalian cells. FACS analysis of wt YR9 and YR9.10 recombinase are shown. Darker grey histograms depict control samples where nonrecombined reporter plasmids were co-transfected with ‘empty’ expression plasmids, while the lighter grey histograms show samples transfected with corresponding recombinases.
[0063] List of sequences
[0064] SEP ID NO: 1 (YR6):
[0065] MIENQLSLLGDFSGVRPDDVKAAVQAAQKKGINVAENEQFKAVFDHLLGEFKKREERYSPNT
[0066] LRRLESAWTCFVDWCLAHHRHSLPATPDTVEAFFIERSETLHRNTLSVYRWAISRVHRVAGC PDPCLDIYVEDRLKAISRKKVREGETVKQASPFNEQHLLKLTSLWYLSDKLLLRRNLALLAVA YESMLRAAELANIRVSDLELSGDGTAVLTIPITKTNHSGEPDTCILSQDVVSLLMDYTEAGRLD MRADGYLFVGISKHNTCINPKRDADTGECLHKPITTKTVEGVFYSAWQALELERQGVKPFTA HSARVGAAQDLLKKGYNTLQIQQSGRWSSGTMVARYGRAILARDGAMAHSRVKTRNVSID
[0067] WGSGGSKNTI
[0068] SEP ID NO: 2 (YR1):
[0069] MIENQLSLLGDFTDVRPSDVKTAIEKAQKKGVVVAEDHVFQAAINHLLNEFKKREDRYSPNT LRRLESAWGCFVEWCLDNKRHSLPASPDTTEKFLIYKAESVHRNTLSIYKWAISRVHRVAGCP
[0070] NPCNDVFVEDRYKALVRVKVQSGEAIKQASPFNELHLNALVEKWKQHERVLERRNLALLGV AYESMLRAAELANIKLSDIELAGDGTAILTIPITKTNHSGDPDTCILSHDVVGLIMDYIEAGELH LKQDGYLFTGVSKHNKCTKPKVDKETGEVTYKPITTKTVEGIFKAAWSELELGRQGVKPFTG HSARVGATQDLLRKGYNTLQIQQSGRWSSEVMVARYGRAILARESAMAQSRVKTKNIDLSW
[0071] GSKR
[0072] SEP ID NO: 3 (YR2):
[0073] MNNEIIHSTSNTSLSQYPAEHIQKALANGDIPTDSHLFQSAADHLINEYRSREGLAENTFLALD TGWSLFVDWCVEHNRVSLPASSKTVEDYVKSISKVLRRNTIRVRKWAITKIHKICGLPNPFDS
[0074] EFVTQTISGIYKKKLHEDEITEQASPFNETHLEALELLYADSTLKKRRDLLMMTIAYESLLRSS ELCNIKLKHLRLIGKEIHITIPVTKTNHSGNPDVVALSEHATNQVLEYLNDHSMKLSGDGYLFR
[0075] RLRRNGLAYPSTKQAMSNQSVIDVFNSVHNDLGGSDVLHCEPFTSHSCRVGGAQDLLAAGYS ILQVQQAGRWSDPSMVYRYGRGIFAAKSAMAHFRRNRQKPRN
[0076] SEP ID NO: 4 (YR4):
[0077] MSELLPLTPLTVDRNSDITERLRQFVQDKEAFSPNTWRQLLSVMRICNRWSEDNQRSFLPMSA DDLRDYLSFLAESGRASSTVTSHAALISMLHRNAGLPVPNVSPLVFRTMKKINRVAVINGERA GQAVPFRLSDLLALDEEWSGSDNLQALRDLAFLHVAYATLLRISELSRLRVRDVMRAGDGRII LDVAWTKTIVQTGGLIKALSARSTQRLEEWIEASGLSSQPDAWLFTAVHRSGRPLIAEKPMST
[0078] RALEQIFSRAWRTAGKEGAVKANKNRYTGWSGHSARVGAAQDMADKGYPIARIMQEGTWK KPETLMRYIRHVDAHKGAMVEFMEQYGDPDYPG
[0079] SEP ID NO: 5 (YR8):
[0080] MGKLSPTNQTLPAIQAEEDVLARLKEFVQDKEAFSPNTWRQLMSVMRICHRWSIENSRSFLP
[0081] MLPADLRDYLNWLQESGRASSTIATHGSLISMLHRNAGLIPPNTSPLVFRAVKKINRVAVVTG ERTGQAVPFRLEDLLELDALWSDSISLRHKRDLAFLHVAYSTLLRISELARLRVRDISRATDGR IILNVSYTKTIVQTGGLIKSLSSQSSRRLTEWMSVSGINAEPDAFLFCPVHRSGSATLSVTRPLST PAIESIFAQAWLTIGAGEPIIPNKGRYTAWTGHSARVGAAQDMAGRGYAVAQIMQEGTWKKP ETLMRYIRNLQAHEGAMTDIMEKSTLDHNNTK
[0082] SEP ID NO: 6 (YR9):
[0083] MLAVLHEDLERAAAYKKAARAAATHRAYNSDWIIYTDWCRTRGLEAMPAHPEQIAAFVAN QAASGLKPSTIERRVAAIGHHHRTSNYPAPAAHPEAGGLREALAGIRNEKRAKKTRKEPADA TALRDMLAQIKGDGLRARRDRAALAIGMAAALRRSELVALTLENVGILEHGIELYLGATKTD QAGEGTTIAIPEGTRLRPKALLLDWISAVRVLEAGVVRTPAQEAAVPLFRRLTRSDQLTGEPM SDKAVARLVKRYAGAAGYDAAKFSGHSLRAGFLTEAANQGATIFKMQEVSRHKTVQVLSD YVRSADRFRDHAGERFL
[0084] SEP ID NO: 7 (YR11):
[0085] MVGGMSFVRRDVVVIPDNPDLNDEVIRNLNAFMKDREAFAENTWKQLMMAVRLWCHWCI AKGRPYLPVDADYLRDYLLELHDNGLAPATISNYAAMLNLLHRQAGLIPAGESQKVKRVLK KISRTSIIKGETVGQAIPFRIADLNQVDEAWEASDRLKTIRNLAFLFVAYNTLLRISNIAHLKVK DLAFDHDGSVMLNIGYTKTLVDGKGITKALSPRASARVLKWLHVSGLLDHPDAYLFCKVYR TNKASVTTDKPLTLHPLESIFSEAWAVIHGEKVGIKNKGRYATWTGHSARVGAAQDMTESGY SLAQIMHEGTWKAPKTVLGYTRNLEAKKSVMIDLVG
[0086] SEP ID NO: 8 (YR12):
[0087] MTEMIVANPLLAQFSASDDISAKLASFVRDREAFSSNTWRQLLSVMRICWRWSEENHRSFLP MAPEDLRDYLLHLQCIGRASSTISTHAALISMLHRNAGLVPPNVSPDVFRVVKKINRAAVIAG ERTGQAVPFCRQDLKKLDTAWQGSPRLQQLRDLAFMHVAYSTLLRLSELSRLRVRDISRAAD GRMILDVAWTKTIVQSGGIVKALSTQSSQRLTDWIVAAGLTGEPDAMIFCPVHRSNRMTKKIF SPMSTPCLEDIFLRAREAAGVAALSRTNKGRYAGWSGHSARVGAAQDMARKGFSVAQIMQE GTWTRTETVMRYIRMVEAHKGAMIGLMEEDE
[0088] SEP ID NO: 9 (YR9, 2):
[0089] MLAILHEDLERAAAYKKAARAAATHRAYNSDWIIYTNWCRTRGLEAMPAHPEQIAAFVANQ AASGLKPSTIERRVAAIGHYHRTSNYPAPAAHPEAGGLREVLAGIRNEKRAKKTRKEPADAT ALRDMLAQIKGDGLRARRDRAVLAIGMAAALRRSELVALTLENVGILEHGIELYLGATKTDQ AGEGTTIAIPKGTRLRPKALLLDWISAVRVLEAGVVRTPAQEAAVPLFRRLTRSDQLTGEPMS DKAVARLVKRYAGAAGYDAAKFSGHSLRAGFLTEAANQGATIFKMQEVSRHKTVQVLSDY VRSADRFRDHAGERFL
[0090] Detailed description of the invention
[0091] Before the present invention is described in detail below, it is to be understood that this invention is not limited to the particular methodology, protocols and reagents described herein as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and it is not intended to limit the scope of the present invention which will be limited only by the appended claims. Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.
[0092] Preferably, the terms used herein are defined as described in "A multilingual glossary of biotechnological terms: (IUPAC Recommendations)", Leuenberger, H.G.W, Nagel, B. and Klbl, H. eds. (1995), Helvetica Chimica Acta, CH-4010 Basel, Switzerland).
[0093] Throughout this specification and the claims which follow, unless the context requires otherwise, the word "comprise", and variations such as "comprises" and "comprising", will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. In the following passages, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects unless clearly indicated to the contrary. Any feature indicated as being optional, preferred or advantageous may be combined with any other feature or features indicated as being optional, preferred or advantageous.
[0094] Several documents are cited throughout the text of this specification. Each of the documents cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions etc.), whether supra or infra, is hereby incorporated by reference in its entirety. Nothing herein is to be construed as an admission that the invention is not entitled to antedate such disclosure by virtue of prior invention. Some of the documents cited herein are characterized as being "incorporated by reference". In the event of a conflict between the definitions or teachings of such incorporated references and definitions or teachings recited in the present specification, the text of the present specification takes precedence.
[0095] In the following, the elements of the present invention will be described. These elements are listed with specific embodiments; however, it should be understood that they may be combined in any manner and in any number to create additional embodiments. The variously described examples and preferred embodiments should not be construed to limit the present invention to only the explicitly described embodiments. This description should be understood to support and encompass embodiments which combine the explicitly described embodiments with any number of the disclosed and / or preferred elements. Furthermore, any permutations and combinations of all described elements in this application should be considered disclosed by the description of the present application unless the context indicates otherwise.
[0096] Definitions
[0097] In the following, some definitions of terms frequently used in this specification are provided. These terms will, in each instance of its use, in the remainder of the specification have the respectively defined meaning and preferred meanings.
[0098] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents, unless the content clearly dictates otherwise.
[0099] The "percentage of sequence identity" is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the sequence in the comparison window can comprise additions or deletions (i.e. gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity.
[0100] The term "identical" is used herein in the context of two or more nucleic acids or polypeptide sequences, to refer to two or more sequences or subsequences that are the same, i.e. that comprise the same sequence of nucleotides or amino acids. Sequences are "identical" to each other if they have a specified percentage of nucleotides or amino acid residues that are the same. According to the present invention, at least 60% identical includes at least at least 61%, at least at least 62%, at least at least 63%, at least at least 64%, at least at least 65%, at least at least 66%, at least at least 67%, at least at least 68%, at least at least 69%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity over the specified sequence, when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection. These definitions also refer to the complement of a test sequence. Accordingly, the term "at least XY% sequence identity" is used throughout the specification with regard to polypeptide and polynucleotide sequence comparisons. This expression preferably refers to a sequence identity of at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% to the respective reference polypeptide or to the respective reference polynucleotide.
[0101] In the context of the present invention, a protein having recombinase activity and comprising an amino acid sequence having at least 80% identity to a given SEQ ID NO preferably means that said protein has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 94%, at least 96%, at least 98% or at least 99% sequence identity to the given SEQ ID NO.
[0102] Likewise, in the context of the present invention, a nucleic acid sequence having at least 60% sequence identity to a given SEQ ID NO or a nucleic acid sequence reverse complementary thereto preferably means that said nucleic acid has a sequence having at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 96%, at least 98% or at least 99% sequence identity to the given SEQ ID NO or a nucleic acid sequence reverse complementary to said SEQ ID NO.
[0103] The term "sequence comparison" is used herein to refer to the process wherein one sequence acts as a reference sequence, to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, if necessary, subsequence coordinates are designated, and sequence algorithm program parameters are designated. Default program parameters are commonly used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters. In case where two sequences are compared and the reference sequence is not specified in comparison to which the sequence identity percentage is to be calculated, the sequence identity is to be calculated with reference to the longer of the two sequences to be compared, if not specifically indicated otherwise. If the reference sequence is indicated, the sequence identity is determined on the basis of the full length of the reference sequence indicated by one of the SEQ ID NOs of the present invention, if not specifically indicated otherwise.
[0104] Methods of alignment of sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be conducted, for example, by the local homology algorithm of Smith and Waterman (Adv. Appl. Math. 2:482, 1970), by the homology alignment algorithm of Needleman and Wunsch (J. Mol. Biol. 48:443, 1970), by the search for similarity method of Pearson and Lipman (Proc. Natl. Acad. Sci. USA 85:2444, 1988), by computerized implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by manual alignment and visual inspection (see, e.g., Ausubel et al., Current Protocols in Molecular Biology (1995 supplement)). Algorithms suitable for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al. (Nuc. Acids Res. 25:3389-402, 1977), and Altschul et al. (J. Mol. Biol. 215:403-10, 1990), respectively. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ). This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a wordlength (W) of 11, an expectation (E) of 10, M=5, N=-4 and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a wordlength of 3, and expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89: 10915, 1989) alignments (B) of 50, expectation (E) of 10, M=5, N=-4, and a comparison of both strands. The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90:5873-87, 1993). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the smallest sum probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.2, typically less than about 0.01, and more typically less than about 0.001.
[0105] The term "nucleic acid" and "nucleic acid molecule" are used synonymously herein and are understood as well -accepted in the art, i.e. as single or double-stranded oligo- or polymers of deoxyribonucleotide or ribonucleotide bases or both. The term "nucleic acids" as used herein includes not only deoxyribonucleic acids (DNA) and ribonucleic acids (RNA), but also all other linear polymers in which the bases adenine (A), cytosine (C), guanine (G) and thymine (T) or uracil (U) are arranged in a corresponding sequence (nucleic acid sequence). The invention also comprises the corresponding RNA sequences (in which thymine is replaced by uracil), complementary sequences and sequences with modified nucleic acid backbone or 3 'or 5 '-terminus. Nucleic acids in the form of DNA are however preferred.
[0106] The term "protein having recombinase activity" comprises any enzyme that is capable of manipulating the structure of a genome. More specifically, the term refers to respective enzymes that can catalyse a recombination selected among an excision, integration, inversion and translocation reaction. Proteins having recombinase activity are well known in the art (cf. discussion of the background above) and specifically include DNA recombinases such as site-specific recombinases and more specifically tyrosine recombinases. A protein having recombinase activity can be present as a monomer, a dimer or a tetramer. Dimers and tetramers can comprise two or four monomers of the same protein having recombinase activity (homodimer or homotetramer), respectively, or alternatively two or more different monomers having recombinase activity (heterodimer or heterotetramer).
[0107] The term "recognition site" (sometimes also referred to as "target site") as used herein refers to a specific nucleotide sequence which a recombinase enzyme recognizes, and at which DNA breakage and strand exchange occur. These sequences typically range between 30 and 200 base pairs in length and are comprised of two inversely repeated recombinase binding regions flanking a central spacer sequence (Meinke et al., 2016). An example of such a recognition site can be seen in the SSR Cre / loxP binding complex, where the Cre recombinase is bound to the 34 base pair loxP target sequence. The loxP recognition site comprises two 13 base pair inverted repeat Cre binding elements flanking an 8 base pair spacer region. The left half-site is the 13 base pair binding element to the left of the spacer and the right half-site is the 13 base pair binding element to the right of the spacer. Depending on the number and relative orientation of the recognition sites and their spacers, the DNA recombining enzyme either performs an excision, an integration, an inversion or a replacement of genetic content (reviewed in Meinke et al., 2016). Therefore, a "recognition site" according to the invention is a nucleotide sequence comprising a first half-site, a second half-site, and a spacer separating the first and the second half-site.
[0108] For a recombination event to occur, the recombinase complex recognizes a first recognition site and a second recognition site on a DNA double strand. The recognition sites are also referred to as upstream and downstream recognition sites, depending on their location on the DNA double strand.
[0109] In symmetric recognition sites, the first half-site (e.g. the left half-site) and the second half-site (e.g. the right half-site) are identical and palindromic (reverse complement). In asymmetric recognition sites, the first half-site (e.g. the left half-site) and the second half-site (e.g. the right halfsite) are not identical and not palindromic, i.e. they differ from each other in at least one nucleotide.
[0110] The term "functional mutant" as used herein means that one or more nucleic acids can be added to, inserted, deleted or substituted from the nucleic acid sequence to which the term refers. In cases of polypeptides and proteins, the term "functional mutant" means that one or more amino acids can be added to, inserted, deleted or substituted from the amino acid sequence to which the term refers. Preferred mutations are point mutations or an exchange of a particular nucleic acid or amino acid. Specifically in cases of functional mutants of a recognition site, it is known that the exchange of the spacer region of a recognition site does not influence its activity as a target for the specific recombinase. Therefore, functional mutants of recognition sites explicitly include mutations in the spacer region up to a full substitution of the spacer region. Functional mutants of recognition sites also include mutations in the half sites of a recognition site. For example, a functional mutant of a recognition site as disclosed herein may include one, two, three, four, five, six, seven or eight mutations in its nucleotide sequence compared to its reference nucleotide sequence, from which the functional mutant is derived. According to a preferred embodiment of the present invention, a functional mutant of a recognition site includes one or two mutations in the nucleotide sequence of the first half site, in the nucleotide sequence of the second half site, or in the nucleotide sequences of both half sites, compared to its reference nucleotide sequence, from which the functional mutant is derived.
[0111] The term "reporter vector" as used herein includes a plasmid, virus or other nucleic acid carriers, that comprise a nucleic acid sequence according to the invention by genetic recombination (recombinantly), e.g. by insertion or incorporation of said nucleic acid sequence. Prokaryotic vectors as well as eukaryotic vectors, for example artificial chromosomes, such as YAC (yeast artificial chromosomes), are applicable for the invention.
[0112] The term "therapeutically effective amount" as used herein, means that amount of active compound or pharmaceutical agent that elicits the biological or medicinal response in a tissue system, animal or human being sought by a researcher, veterinarian, medical doctor or other clinician, which includes alleviation of the symptoms of the disease or disorder being treated.
[0113] The term "pharmaceutical composition" as used herein refers to a substance and / or a combination of substances being used for the identification, prevention or treatment of a disease or tissue status. The pharmaceutical composition is formulated to be suitable for administration to a patient in order to prevent and / or treat a disease. Further a pharmaceutical composition refers to the combination of an active agent with a carrier, inert or active, making the composition suitable for therapeutic use. Such a carrier is also referred to as being pharmaceutically acceptable. Pharmaceutical compositions can be formulated for oral, parenteral, topical, inhalative, rectal, sublingual, transdermal, subcutaneous or vaginal application routes according to their chemical and physical properties. Pharmaceutical compositions comprise solid, semisolid, liquid, transdermal therapeutic systems (TTS). Solid compositions are selected from the group consisting of tablets, coated tablets, powder, granulate, pellets, capsules, effervescent tablets or transdermal therapeutic systems. Also comprised are liquid compositions, selected from the group consisting of solutions, syrups, infusions, extracts, solutions for intravenous application, solutions for infusion or solutions of the carrier systems of the present invention. Semisolid compositions that can be used in the context of the invention comprise emulsion, suspension, creams, lotions, gels, globules, buccal tablets and suppositories. As used herein, the term "pharmaceutically acceptable" embraces both human and veterinary use: For example, the term "pharmaceutically acceptable" embraces a veterinarily acceptable compound or a compound acceptable in human medicine and health care.
[0114] The term "subject" as used herein, refers to an animal, preferably a mammal, most preferably a human.
[0115] Description of embodiments
[0116] According to the present invention, the method for producing a site-specific DNA- recombination comprises the steps of a) contacting a nucleic acid comprising at least a first and a second recognition site which are essentially identical or essentially reverse complementary to each other, with a protein having recombinase activity, and b) allowing the protein having recombinase activity to produce the site-specific DNA-recombination.
[0117] In accordance with said method and as defined above, a recognition site is a nucleotide sequence comprising a first half-site, a second half-site, and a spacer separating the first and the second half-sites.
[0118] As used in the context of the present invention, the term "essentially identical" or "essentially reverse complementary to each other" means that the nucleotide sequence of the first- and the second half-site in the first recognition site may deviate in up to two nucleotides from the nucleotide sequence of the first and the second half-site in the second recognition site. As an example, it is referred to the recognition sites as identified herein and as shown in Fig. 4B. Taking the recombinase YR6 (SEQ ID NO: 1) as an example, its recognition site has the sequence of TCAATTTCCGAGA ATGACAGT TCTCAGAAATTAA (SEQ ID NO: 10), with the spacer sequence denoted in bold and underlined. In cases where the first and the second recognition sites are identical, both recognition sites have the sequence of SEQ ID NO: 10. In cases where the recognition sites are reverse complementary to each other, one of the two recognition sites has the sequence of SEQ ID NO: 10, and the other has the sequence of TTAATTTCTGAGA ACTGTCAT TCTCGGAAATTGA (SEQ ID NO: 19). In cases of essentially identical sequences, one of the recognition sites has SEQ ID NO: 10, and the other has SEQ ID NO: 10 but with up to two diverging nucleotides in the sequences of the first and the second half-site. This means that either the first half-site of a recognition site comprises a nucleotide sequence that differs from the nucleotide sequence of the first half-site of the other recognition site by two nucleotides, and the second half site and the spacer are identical in both recognition sites, the first halfsite of a recognition site comprises a nucleotide sequence that differs from the nucleotide sequence of the first half-site of the other recognition site by one nucleotide, the second half site of the same recognition site comprises a nucleotide sequence that differs from the nucleotide sequence of the second half-site of the other recognition site by one nucleotide, and the spacer sequences identical in both recognition sites, or the second half-site of a recognition site comprises a nucleotide sequence that differs from the nucleotide sequence of the second half-site of the other recognition site by two nucleotides, and the first half site and the spacer are identical in both recognition sites. The same principle applies to cases where the two recognition sites comprise sequences that are essentially reverse complementary to each other. In such cases, the reverse complement sequence must be compared to the non-reverse complement sequence. Using the recognition sites for YR6 as an example again, a first recognition site has the sequence of SEQ ID NO: 10, and the other recognition site is the reverse complement thereto, i.e. SEQ ID NO: 19. In cases where the two recognition sites are essentially reverse complementary to each other, the first recognition site may have the sequence of SEQ ID NO: 10, and the second recognition site may have the sequence of SEQ ID NO: 19, with the proviso that SEQ ID NO: 19 comprises up to two diverging nucleotides in the sequences for the first and the second half-sites, but not in the sequence of the spacer. In analogy to the example above for essentially identical recognition sites, the two diverging nucleotides may either be in the first half-site, the second half-site, or there is one different nucleotide in each the first and the second half-site.
[0119] According to one embodiment of the present invention, the first and second recognition sites are identical or reverse complementary to each other in their half site sequences, meaning that the first and second recognition sites do not have any sequence deviations between their first half sites and between their second half sites. In such an embodiment, the spacer sequences of the first and the second recognition sites do not have to be identical or reverse complementary to each other. According to a further embodiment of the present invention, the first and second recognition sites are identical or reverse complementary to each other over their entire sequence including the first and second half sites and the spacer region, meaning that the first and second recognition sites do not have any sequence deviations between each other.
[0120] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 7, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 16 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 16.
[0121] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 3, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 12 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 12.
[0122] According to one aspect of the method of the present invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 1, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 10 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 10. According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 2, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 11 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 11.
[0123] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 4, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 13 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 13.
[0124] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 5, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 14 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 14.
[0125] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 6, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15.
[0126] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 8, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 17 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 17.
[0127] According to another aspect of the method of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 9, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15.
[0128] A nucleic acid encoding the recognition sites according to the invention comprises a maximum of 40, preferably 34, base pairs. This nucleic acid according to the invention includes the recognition sites of the recombinase proteins of the present invention. Accordingly, the present invention provides a nucleic acid having a length of not more than 40 base pairs and comprising:
[0129] (i) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 16 or a nucleic acid sequence reverse complementary thereto; or (ii) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 12 or a nucleic acid sequence reverse complementary thereto; or
[0130] (iii) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 10 or a nucleic acid sequence reverse complementary thereto; or
[0131] (iv) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 11 or a nucleic acid sequence reverse complementary thereto; or
[0132] (v) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 13 or a nucleic acid sequence reverse complementary thereto; or
[0133] (vi) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 14 or a nucleic acid sequence reverse complementary thereto; or
[0134] (vii) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 15 or a nucleic acid sequence reverse complementary thereto; or
[0135] (viii) a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 17 or a nucleic acid sequence reverse complementary thereto.
[0136] In the methods of the present invention, a protein having recombinase activity and comprising an amino acid sequence having at least 80% identity to a given SEQ ID NO preferably means that said protein has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 94%, at least 96%, at least 98% or at least 99% sequence identity to the given SEQ ID NO.
[0137] Likewise, in the methods of the present invention, a nucleic acid sequence having at least 60% sequence identity to a given SEQ ID NO or a nucleic acid sequence reverse complementary thereto preferably means that said nucleic acid has a sequence having at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 96%, at least 98% or at least 99% sequence identity to the given SEQ ID NO or a nucleic acid sequence reverse complementary to said SEQ ID NO.
[0138] In the method according to the invention, the recombinase protein is contacted with at least two of its recognition sites preferably inside a cell. Upon binding of the recombinase protein to the recognition sites, site-specific DNA-recombination occurs. For example, the recombinase YR6 having SEQ ID NO: 1 is contacted with at least two lox6 sites having SEQ ID NO: 10 inside a cell. Upon binding of the recombinase protein to said lox6 sites, site-specific DNA-recombination occurs.
[0139] The method according to the invention can be carried out in vitro or in vivo. In case the invention is carried out in an animal (including humans), it can be carried out for a non-therapeutic use. According to one embodiment, the method is not for the therapeutic treatment of a human being or an animal.
[0140] The method is applicable in all areas where site specific recombinases are conventionally used (including inducible knock-out or knock-in mice and other transgenic animal models). In preferred methods according to the invention, the site-specific recombination results in integration, deletion, inversion, translocation or exchange of DNA. According to one embodiment of the invention, the method is used for creating animal models, which are useful for biomedical research, e.g. as models for human diseases.
[0141] According to one embodiment of the present invention, the protein having recombinase activity is present as a monomer. According to a preferred embodiment of the present invention, the protein having recombinase activity comprises at least two protein monomers, i.e. is in a dimeric form (dimer). Such a dimer can comprise two monomers of the same type (i.e. two identical protein monomers, homodimer) or two monomers of a different type (i.e. two different protein monomers, heterodimer). According to a further preferred embodiment of the present invention, the protein having recombinase activity comprises at least four protein monomers, i.e. is in a tetrameric form (tetramer). Such a tetramer can comprise four monomers of the same type (i.e. four identical protein monomers, homotetramer), or monomers of a different type such as two, three or four different monomers (heterotetramer). According to a specifically preferred embodiment, the protein having recombinase activity comprises a heterodimer or heterotetramer comprising different protein monomers, i.e. is a heterodimer or a heterotetramer.
[0142] According to a specifically preferred embodiment of the present invention, the protein having recombinase activity is a recombinase, more preferably a DNA recombinase.
[0143] According to one preferred embodiment, the nucleic acid sequence that is to be recombined is already present in a cell. Alternatively, the nucleic acid sequence to be recombined can be introduced into the cell by conventional means known to the skilled person, such as by recombinant techniques.
[0144] According to one embodiment, the nucleic acid sequence encoding the recombinase protein of the present invention is either already present in the cell or is introduced into the cell by conventional means known to the skilled person, such as by recombinant techniques. Such a method thus further includes the step of introducing into the cell a nucleic acid encoding for the recombinase protein of the invention.
[0145] Alternatively, according to one embodiment of the present invention, the cell already comprises a nucleic acid encoding the protein of the invention having recombinase activity.
[0146] For activation of the expression of the nucleic acid encoding for the protein having recombinase activity, the nucleic acid encoding for the recombinase protein further comprises a regulatory nucleic acid sequence, preferably a promoter region. Hence, expression of the nucleic acid encoding for the protein with recombinase activity is initiated or regulated by activating the regulatory nucleic acid sequence. Accordingly, to induce a DNA recombination, the regulatory nucleic acid sequence (preferably the promoter region) is activated to express the gene encoding for the recombinase protein. Preferably, the regulatory nucleic acid sequence (preferably the promoter region) is either introduced into the cell, preferably together with the sequence encoding for the recombinase protein, or the regulatory nucleic acid sequence is already present in the cell in the beginning of the method according to the invention. In the second case, merely the nucleic acid encoding for the recombinase protein is introduced into the cell (and placed under the control of the regulatory nucleic acid sequence).
[0147] The term "regulatory nucleic acid sequence" as used herein refers to gene regulatory regions of DNA. In addition to promoter regions, this term encompasses operator regions more distant from the gene as well as nucleic acid sequences that influence the expression of a gene, such as ciselements, enhancers or silencers. The term "promoter region" as used herein refers to a nucleotide sequence on the DNA allowing a regulated expression of a gene. The promoter region allows regulated expression of the nucleic acid encoding for the respective protein. The promoter region is located at the 5'-end of the gene and thus before the RNA coding region. Both, bacterial and eukaryotic promoters are applicable for the invention.
[0148] According to one embodiment of the invention, the recognitions sites are either included in the cell or introduced into the cell, preferably by recombinant techniques. Accordingly, the methods of the present invention may include the steps of introducing into a cell the following nucleic acids: a) a first nucleic acid (first recognition site) comprising a nucleic acid sequence according to or reverse complementary to one of SEQ ID NOs: 10 to 17; or a nucleic acid sequence that is a functional mutant thereof; b) a second nucleic acid (second recognition site) comprising a nucleic acid sequence essentially identical or essentially reverse complementary to the nucleic acid sequence of the first nucleic acid (first recognition site).
[0149] According to one preferred embodiment of the present invention, a nucleic acid encoding for the recombinase protein of the invention and at least two recognition sites are introduced into the cell. This method includes the following steps: a) introducing into a cell the following nucleic acids:
[0150] (i) a first nucleic acid encoding for the recombinase protein, wherein the nucleic acid is introduced into the DNA such that a regulatory nucleic acid sequence (preferably a promoter region) controls the expression of the nucleic acid encoding for the recombinase protein,
[0151] (ii) a second nucleic acid comprising a nucleic acid sequence according to or reverse complementary to one of SEQ ID NOs: 10 to 17; or a nucleic acid sequence that is a functional mutant thereof;
[0152] (iii) a third nucleic acid comprising a nucleic acid sequence essentially identical or essentially reverse complementary to the nucleic acid sequence defined in (ii), and b) activating the regulatory nucleic acid sequence (preferably the promoter region) to induce expression of the first nucleic acid for the synthesis of the protein with recombinase activity.
[0153] According to this method, the nucleic acid sequence encoding for the recombinase protein is introduced into a cell and at least two recognition sites are introduced into the genomic or episomal DNA of the cell. The steps (i) to (iii) can be performed in any order. The introduction of the nucleic acids into the cells is performed using techniques of genetic manipulation known by a person skilled in the art. Among suitable methods are cell transformation, transfection or viral infection, whereby a nucleic acid sequence encoding the protein is introduced into the cell as a component of a vector or part of virus-encoding DNA or RNA. The cell culturing is carried out by methods known to a person skilled in the art for the culture of the respective cells. Therefore, cells are preferably transferred into a conventional culture medium, and cultured at temperatures and in a gas atmosphere that is conducive to the survival of the cells.
[0154] The present invention also includes nucleic acid sequences or polynucleotides in which the coding sequence for the recombinase protein is fused in the same reading frame to a polynucleotide sequence which aids in expression and secretion of a protein from a host cell. For example, a leader sequence which functions as a secretory sequence for controlling transport of a polypeptide from the cell may be fused to the sequence encoding the recombinase protein. The polypeptide or protein having such a leader sequence is termed a pre-protein or a pre-proprotein and may have the leader sequence cleaved by the host cell to form the mature form of the protein. These polynucleotides may have a 5' extended region so that it encodes a proprotein, which is the mature protein plus additional amino acid residues at the N-terminus. The expression product having such a pro-sequence is termed a pro-protein, which is an inactive form of the mature protein; however, once the pro-sequence is cleaved, an active mature protein remains. The additional sequence may also be attached to the protein and be part of the mature protein. Thus, for example, the polynucleotides of the present invention may encode polypeptides, or proteins having a pro-sequence, or proteins having both, a pro-sequence and a pre-sequence (such as a leader sequence).
[0155] The nucleic acids of the present invention may also have the coding sequence fused in frame to a marker sequence, which allows for purification of the proteins of the present invention. The marker sequence may be an affinity tag or an epitope tag such as a polyhistidine tag, a streptavidin tag, a Xpress tag, a FLAG tag, a cellulose or chitin binding tag, a glutathione-S transferase tag (GST), a hemagglutinin (HA) tag, a c-myc tag or a V5 tag.
[0156] The HA tag would correspond to an epitope obtained from the influenza hemagglutinin protein (Wilson et al., 1984), and the c-myc tag may be an epitope from human Myc protein (Evans et al., 1985).
[0157] If the nucleic acid of the invention is a mRNA, in particular for use as a medicament, the delivery of mRNA therapeutics can be facilitated by the significant progress that has been achieved in maximizing the translation and stability of mRNA, preventing its immune-stimulatory activity and the development of in vivo delivery technologies. The 5' cap and 3' poly(A) tail are the main contributors to efficient translation and prolonged half-life of mature eukaryotic mRNAs. Incorporation of cap analogs such as ARCA (anti-reverse cap analogs) and poly(A) tail of 120-150 bp into in vitro transcribed (IVT) mRNAs has markedly improved expression of the encoded proteins and mRNA stability. New types of cap analogs, such as 1,2-dithiodiphosphate-modified caps, with resistance against RNA decapping complex, can further improve the efficiency of RNA translation. Replacing rare codons within mRNA protein-coding sequences with synonymous frequently occurring codons, so-called codon optimization, also facilitates better efficacy of protein synthesis and limits mRNA destabilization by rare codons, thus preventing accelerated degradation of the transcript. Similarly, engineering 3' and 5' untranslated regions (UTRs), which contain sequences responsible for recruiting RNA-binding proteins (RBPs) and miRNAs, can enhance the level of protein product. Interestingly, UTRs can be deliberately modified to encode regulatory elements (e.g., K-tum motifs and miRNA binding sites), providing a means to control RNA expression in a cell-specific manner. Some RNA base modifications such as Nl-methyl-pseudouridine have not only been instrumental in masking mRNA immune-stimulatory activity but have also been shown to increase mRNA translation by enhancing translation initiation. In addition to their observed effects on protein translation, base modifications and codon optimization affect the secondary structure of mRNA, which in turn influences its translation. Respective modifications of the nucleic acid molecules of the invention are also contemplated by the invention.
[0158] The RNA or plurality of RNAs preferably encode the recombinase enzyme of the present invention or any of its subunits. Specific methods for delivering and expressing nucleic acids and specifically RNAs are disclosed e.g. in EP2590676 and EP3115064, which are herein incorporated by reference. The RNA may be present in a particle and is preferably self-replicating. After in vivo administration of the particles, RNA is released from the particles and is translated inside a cell to provide the DNA recombining enzyme or any of its monomeric subunits.
[0159] A self-replicating RNA molecule (replicon) can, when delivered to a vertebrate cell even without any proteins, lead to the production of multiple daughter RNAs by transcription from itself (via an antisense copy which it generates from itself). These daughter RNAs, as well as collinear sub- genomic transcripts, may be translated by themselves to provide in situ expression of an encoded polypeptide, or may be transcribed to provide further transcripts with the same sense as the delivered RNA which are translated to provide in situ expression of the polypeptide. The overall results of this sequence of transcriptions is a huge amplification in the number of the introduced replicon RNAs and so the encoded polypeptide becomes a major polypeptide product of the cells.
[0160] A preferred self-replicating RNA molecule encodes (i) an RNA-dependent RNA polymerase which can transcribe RNA from the self-replicating RNA molecule, and (ii) a recombinase protein of the present invention. The polymerase can be an alphavirus replicase e.g. comprising one or more of alphavirus proteins nsPl, nsP2, nsP3 and nsP4. It is preferred that the self-replicating RNA molecules of the invention do not encode alphavirus structural proteins. Thus, a preferred self-replicating RNA can lead to the production of genomic RNA copies of itself in a cell, but not to the production of RNA-containing virions. A self-replicating RNA molecule useful in the context of the present invention may have two open reading frames. The first (51) open reading frame encodes a replicase, and the second (31) open reading frame encodes a polypeptide of the present invention. In some embodiments, the RNA may have additional (e.g. downstream) open reading frames e.g. for further encoding accessory polypeptides.
[0161] Such RNA is particularly suitable for the general use in gene therapy, and specifically for use in the treatment of genetic disorder or disease.
[0162] The method according to the invention can be performed using eukaryotic and prokaryotic cells. Preferred prokaryotic cells are bacterial cells. Particularly preferred prokaryotic cells are cells of Escherichia coli. Preferred eukaryotic cells are yeast cells (preferably Saccharomyces cerevisiae). insect cells, non-insect invertebrate cells, amphibian cells, or mammalian cells (preferably somatic or pluripotent stem cells, including embryonic stem cells and other pluripotent stem cells, like induced pluripotent stem cells, and other native cells or established cell lines, including NIH3T3, CHO, HeLa, HEK293, hiPS). In case of human embryonic stem cells, cells are preferably obtained without destroying human embryos, e.g. by outgrowth of single blastomeres derived from blastocysts as described by Chung et al., 2008, by parthenogenesis, e.g. from a one-pronuclear oocyte as described by Lin et al. 2007, or by parthenogenetic activation of human oocytes as described by Mai et al. 2007. Also preferred are cells of a non-human host organism, preferably non-human germ cells, somatic or pluripotent stem cells, including embryonic stem cells, or blastocytes.
[0163] According to one embodiment, the cell is a eukaryotic or a bacterial cell.
[0164] According to another embodiment, the cell is not a human germ cell.
[0165] According to a further aspect, the present invention provides the use of a protein having at least 80% identity to SEQ ID NO: 7, SEQ ID NO: 3, SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 8, or SEQ ID NO: 9 for producing a site-specific DNA- recombination.
[0166] In the context of the present invention, a protein having at least 80% identity to a given SEQ ID NO preferably means that said protein has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 94%, at least 96%, at least 98% or at least 99% sequence identity to the given SEQ ID NO.
[0167] According to a yet a further aspect, the present invention provides the use of a protein having recombinase activity for catalyzing a site-specific DNA-recombination at recognition sites that essentially are identical or essentially reverse complementary to each other, wherein a recognition site comprises a first half-site, a spacer and a second half-site, and wherein essentially identical or essentially reverse complementary to each other means that the nucleotide sequence of the first and the second half-site in a first recognition site may deviate in up to two nucleotides from the nucleotide sequence of the first and the second half-site in a second recognition site.
[0168] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 7, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 16 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 16.
[0169] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 3, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 12 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 12.
[0170] According to one aspect of the use of the present invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 1, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 10 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 10.
[0171] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 2, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 11 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 11.
[0172] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 4, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 13 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 13.
[0173] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 5, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 14 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 14.
[0174] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 6, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15.
[0175] According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 8, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 17 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 17. According to another aspect of the use of the invention, the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 9, and the at least one two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15.
[0176] In the uses of the present invention, a protein having recombinase activity and comprising an amino acid sequence having at least 80% identity to a given SEQ ID NO preferably means that said protein has an amino acid sequence having at least 85%, at least 90%, at least 92%, at least 94%, at least 96%, at least 98% or at least 99% sequence identity to the given SEQ ID NO.
[0177] Likewise, in the uses of the present invention, a nucleic acid sequence having at least 60% sequence identity to a given SEQ ID NO or a nucleic acid sequence reverse complementary thereto means that said nucleic acid has a sequence having at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 96%, at least 98% or at least 99% sequence identity to the given SEQ ID NO or a nucleic acid sequence reverse complementary to said SEQ ID NO.
[0178] According to a further aspect, the present invention provides a vector comprising at least one and preferably at least two essentially identical or essentially reverse complementary nucleic acids of the present invention, wherein a DNA segment is preferably flanked by the two essentially identical or essentially reverse complementary nucleic acids.
[0179] According to one embodiment, the vector further comprises a nucleic acid encoding a protein having recombinase activity, wherein the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to any one of SEQ ID NOs: 1 to 9.
[0180] According to a further aspect, the present invention provides a vector comprising a nucleic acid encoding a protein having recombinase activity, wherein the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to any one of SEQ ID NOs: 1 to 9.
[0181] According to a preferred aspect of the present invention, there is provided the use of the vector of the present invention in a method of the present invention.
[0182] Additionally, the invention provides vectors comprising the nucleic acid sequences of the present invention. According to one embodiment, such vector comprises at least one, preferably at least two essentially identical or essentially reverse complementary nucleic acid sequences as described herein for the recognition sites. According to a further embodiment, a vector according to the present invention comprises a nucleic acid encoding for a protein having recombinase activity, preferably wherein the protein comprises an amino acid sequence exhibiting at least 80%, preferably at least 85%, at least 90%, at least 95%, and most preferably at least 99% sequence identity to one of SEQ ID NOs: 1 to 9.
[0183] The present invention further provides a vector (also referred to herein as "reporter vector") comprising at least one nucleic acid comprising a nucleic acid sequence according to or reverse complementary to one of SEQ ID NOs: 10 to 17, or a nucleic acid sequence that is a functional mutant thereof. Typically, an expression vector comprises an origin of replication, a promoter, as well as specific gene sequences that allow phenotypic selection of host cells comprising the reporter vector. In a preferred embodiment of the invention, the vector comprises at least two recognition-sites, i.e. at least two nucleic acids that independently of each other exhibit a nucleic acid sequence according to or reverse complementary one of SEQ ID NOs: 10 to 17, or a nucleic acid sequence that is a functional mutant thereof. The at least two recognition sites are preferably not located consecutively in the vector. Rather, the at least two recognition sites are positioned such that they are flanking a DNA segment of interest, which upon recognition of the recognition sites by the respective recombinase protein is recombined, preferably excised or inverted. The DNA segment of interest can preferably contain a gene or a promoter region. The DNA segment of interest is excised when it is flanked by two recognition sites having the same orientation (same nucleic acid sequence). An inversion of the DNA segment of interest is catalyzed by a recombinase protein when the DNA segment is flanked by two recognition sites arranged in opposite orientations, i.e. the recognition sites comprise nucleic acid sequences that are reverse complementary to one another.
[0184] The use of any of the vectors according to the invention in a method for producing a sitespecific DNA-recombination is also included in the invention.
[0185] According to a further aspect, the present invention provides an isolated host cell, comprising the following recombinant DNA fragments: at least one and preferably at least two nucleic acids according to the invention; and / or a vector according to the invention.
[0186] According to one embodiment, the isolated host cell further comprises a nucleic acid encoding for a protein having recombinase activity, wherein the protein comprises an amino acid sequence having at least 80% identity to any one of SEQ ID NOs: 1 to 9.
[0187] According to a further embodiment, the isolated host cell further comprises a vector according to the invention.
[0188] Preferably excluded are isolated host cells that comprise the nucleic acids or the vector as mentioned herein naturally. According to one embodiment, the invention concerns only those isolated host cells that comprise the above mentioned nucleic acids or vectors recombinantly and not naturally, i.e. by genetic modification of the host cell.
[0189] According to a further aspect, the present invention provides a non-human host organism, comprising at least one and preferably at least two nucleic acids according to the invention. According to a further aspect, the present invention provides a non-human host organism, comprising the vector according to the invention.
[0190] Further, the invention includes an isolated host cell or an isolated host organism comprising (i) at least one, preferably at least two, nucleic acids according to the invention comprising a recognition site as defined above (preferably two nucleic acids according to the invention that include a recognition site, respectively, and which flank a further DNA segment) and / or a nucleic acid according to the invention encoding for a protein with recombinase activity, as defined above, or (ii) a vector according to the invention comprising at least two nucleic acids comprising a recognition site as defined above (preferably two nucleic acids according to the invention that include a recognition site, respectively, and which flank a further DNA segment) and / or a vector according to the invention comprising a nucleic acid encoding for a protein with recombinase activity as defined above.
[0191] The present invention also includes an isolated host cell comprising the following recombinant DNA fragments: (i) at least one, preferably at least two, nucleic acids according to the invention comprising a recognition site and / or a nucleic acid according to the invention encoding for a recombinase protein or (ii) a vector according to the invention comprising at least two nucleic acids comprising a recognition site (preferably two nucleic acids according to the invention that include a recognition site and flank a further DNA segment of interest) and / or a vector according to the invention comprising a nucleic acid encoding for a recombinase protein.
[0192] The invention concerns only those isolated host cells that comprise the above mentioned nucleic acids or vectors recombinantly and not naturally, i.e. by genetic modification of the host cell.
[0193] Particularly preferred are isolated host cells that contain both, a nucleic acid encoding for the recombinase protein of the present invention and at least two of its recognition sites (which are either oriented in the same or in opposite direction).
[0194] A host cell within the meaning of the invention is a naturally occurring cell or a cell line (optionally transformed or genetically modified) that comprises at least one vector according to the invention or a nucleic acid according to the invention recombinantly, as described above. Thereby, the invention includes transient transfectants (e.g. by mRNA injection) or host cells that include at least one expression vector according to the invention as a plasmid or artificial chromosome, as well as host cells in which an expression vector according to the invention is stably integrated into the genome of said host cell.
[0195] Suitable host cells in the context of the present invention are in particular eukaryotic cells, including stem cells like hematopoietic stem cells, neuronal stem cells, adipose tissue derived stem cells, fetal stem cells, umbilical cord stem cells, induced pluripotent stem cells and embryonic stem cells. In case of human embryonic stem cells, there are preferably not derived from the destruction of embryos. Further, the modification of the human germline and of human gametes as host cells is preferably excluded.
[0196] Using the present invention, it is also possible to induce tissue-specific or site-specific recombination in host organisms, such as mammals. Therefore, the present invention also includes a non-human host organism comprising the following recombinant DNA fragments: at least one, preferably at least two, nucleic acids according to the invention comprising a recognition site (preferably two nucleic acids according to the invention that include a recognition site, respectively, which flank a further DNA segment of interest) and / or a nucleic acid according to the invention encoding for a recombinase protein. Explicitly included are non-human host organisms that only comprise a recombinant nucleic acid encoding for a recombinase protein of the present invention (and which do not comprise a nucleic acid including a recognition site respectively).
[0197] Furthermore, the invention includes non-human host organisms that only comprise at least one, preferably at least two recognition sites (and which do not comprise a nucleic acid encoding for a recombinase protein of the invention). Upon cross-breeding of two non-human host organisms, wherein a first host organism comprises a recombinant nucleic acid encoding for a recombinase protein of the invention, and a second host organism comprises at least two recombinant recognition sites preferably flanking a further DNA segment of interest, the offspring includes host organisms expressing the recombinase protein and further including the recognition sites, so that a site-specific DNA-recombination, like a tissue-specific conditional knock-out, is possible.
[0198] Also provided are non-human host organisms may comprise a vector according to the invention or a nucleic acid according to the invention as described above that is, respectively, stably integrated into the genome of the host organism or individual cells of the host organism.
[0199] Preferred host organisms according to the present invention are plants, invertebrates and vertebrates, particularly Bovidae. Drosophila melanogaster, Caenorhahditis elegans. Xenopus laevis. medaka, zebrafish, or Mus musculus, or embryos of these organisms.
[0200] The present invention also provides a novel recombinase system suitable for producing a sitespecific recombination in cells of various cell types. Such a system includes the respective recombinase protein (one of SEQ ID NOs: 1 to 9) and at least two respective recognition sites as identified herein. With such a system, a diverse range of genetic manipulations can be realized, particularly rearrangements of the DNA fragments flanked by the recognition sites of the present invention in same orientation (excision), opposite orientation (inversion), or when one specific recognition site is present on each of two DNA molecules with one - if in circular form - in any orientation (integration). Exemplary manipulations are the excision of a DNA segment that is flanked by two recognition sites oriented in the same direction mediated by the respective recombinase protein as identified herein. Amongst others, the recombinase systems according to the invention provide the possibility to excise a target DNA, such as a stopper DNA fragment, flanked by two recognition sites, which target DNA is located 5’ of the gene and 3’ of the corresponding to the gene promoter. Without recombination, the stopper sequence prevents gene expression, whereas upon recombinase-mediated excision of the stopper via the two flanking recognition sites, the gene is located to the proximity of the promoter and will therefore be expressed. Different types of promoter regions that regulate the expression of the recombinase allow, inter aha, conditional DNA recombination, when for example a tissue or organism-specific or inducible promoter region is used to express the recombinase protein.
[0201] The recombinase systems according to the invention are applicable for use in combination with other recombinase systems and become a particular valuable tool for genetic experiments where multiple recombinases are required simultaneously or sequentially. The present invention further provides the use of a protein with recombinase activity, wherein the protein comprises an amino acid sequence exhibiting at least 85%, at least 90%, at least 95%, preferably at least 95%, even more preferably at least 99% amino acid sequence identity to one of the amino acid sequences according to SEQ ID NOs: 1 to 9 to catalyze a site-specific DNA recombination. The aforementioned site-specific recombinase is thus used for site-specific DNA recombination at least one and preferably at least two recognition sites of the present invention that are essentially identical or essentially reverse complementary to each other. The at least one recognition site comprises a nucleic acid sequence according to or reverse complementary to the nucleic acid sequence according to SEQ ID NOs: 10 to 17, respectively, or a nucleic acid sequence that is a functional mutant thereof.
[0202] According to a further aspect, the present invention provides a protein having at least 80% identity to SEQ ID NO: 7, SEQ ID NO: 3, SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 8 or SEQ ID NO: 9, or a vector according to the invention, for use in medicine. According to a preferred embodiment, the protein or the vector is for use in treating a genetic disease or disorder in a subject. Preferably, the genetic disease or disorder is characterized by modification of the subject’s genome.
[0203] The present invention also provides a pharmaceutical composition comprising the recombinase protein of the present invention, the vector or the host cell of the present invention, or one or more nucleic acids according to the present invention, and optionally a pharmaceutically acceptable carrier.
[0204] The pharmaceutical compositions that contain a therapeutically active agent according to the invention may be in any form that is suitable for the selected mode of administration.
[0205] In one embodiment, a pharmaceutical composition of the present invention is administered parenterally.
[0206] The phrases "parenteral administration" and "administered parenterally" as used herein means modes of administration other than enteral and topical administration, usually by injection, and include epidermal, intravenous, intramuscular, intraarterial, intrathecal, intracapsular, intraorbital, intracardiac, intradermal, intraperitoneal, intratendinous, transtracheal, subcutaneous, subcuticular, intraarticular, subcapsular, subarachnoid, intraspinal, intracranial, intrathoracic, epidural and intrastemal injection and infusion.
[0207] The therapeutically active agents as referred to herein include but are not limited to the recombinase proteins of the present invention and the recognition sites of the present invention. The therapeutically active agents of the invention can be administered, as sole active agent, or in combination with other active agents, in a unit administration form, as a mixture with conventional pharmaceutical supports, to animals and human beings.
[0208] In further embodiments, the pharmaceutical compositions contain carriers (also termed vehicles) which are pharmaceutically acceptable for a formulation capable of being injected. These may be in particular isotonic, sterile, saline solutions (monosodium or disodium phosphate, sodium, potassium, calcium or magnesium chloride and the like or mixtures of such salts), or dry, especially freeze-dried compositions which upon addition, depending on the case, of sterilized water or physiological saline, permit the constitution of injectable solutions.
[0209] The pharmaceutical forms suitable for injectable use include sterile aqueous solutions or dispersions; formulations including sesame oil, peanut oil or aqueous propylene glycol; and sterile powders for the extemporaneous preparation of sterile injectable solutions or dispersions. In all cases, the form must be sterile and must be fluid. It must be stable under the conditions of manufacture and storage and must be preserved against the contaminating action of microorganisms, such as bacteria and fungi.
[0210] Solutions comprising the therapeutically active agents as free base or pharmacologically acceptable salts can be prepared in water suitably mixed with a surfactant, such as hydroxypropylcellulose. Dispersions can also be prepared in glycerol, liquid polyethylene glycols, and mixtures thereof and in oils. Under ordinary conditions of storage and use, these preparations contain a preservative to prevent the growth of microorganisms.
[0211] The therapeutically active agents can be formulated into a composition in a neutral or salt form. Pharmaceutically acceptable salts include the acid addition salts (formed with the free amino groups of the protein) and which are formed with inorganic acids such as, for example, hydrochloric or phosphoric acids, or such organic acids as acetic, oxalic, tartaric, mandelic, and the like. Salts formed with the free carboxyl groups can also be derived from inorganic bases such as, for example, sodium, potassium, ammonium, calcium, or ferric hydroxides, and such organic bases as isopropylamine, trimethylamine, histidine, procaine and the like.
[0212] The carrier can also be as solvent or dispersion medium containing, for example, water, ethanol, polyol (for example, glycerol, propylene glycol, and liquid polyethylene glycol, and the like), suitable mixtures thereof, and vegetables oils. The proper fluidity can be maintained, for example, by the use of a coating, such as lecithin, by the maintenance of the required particle size in the case of dispersion and by the use of surfactants. The prevention of the action of microorganisms can be brought about by various antibacterial and antifungal agents, for example, parabens, chlorobutanol, phenol, sorbic acid, thimerosal, and the like. In many cases, it will be preferable to include isotonic agents, for example, sugars or sodium chloride. Prolonged absorption of the injectable compositions can be brought about by the use in the compositions of agents delaying absorption, for example, aluminum monostearate and gelatin.
[0213] Sterile injectable solutions are prepared by incorporating the active polypeptides in the required amount in the appropriate solvent with several of the other ingredients enumerated above, as required, followed by fdtered sterilization. Generally, dispersions are prepared by incorporating the various sterilized active ingredients into a sterile vehicle which contains the basic dispersion medium and the required other ingredients from those enumerated above. In the case of sterile powders for the preparation of sterile injectable solutions, the preferred methods of preparation are vacuum-drying and freeze-drying techniques which yield a powder of the active ingredient plus any additional desired ingredient from a previously sterile-fdtered solution thereof.
[0214] Upon formulation, solutions can be administered in a manner compatible with the dosage formulation and in such amount as is therapeutically effective. The formulations are easily administered in a variety of dosage forms, such as the type of injectable solutions described above, but drug release capsules and the like can also be employed. Multiple doses can also be administered. As appropriate, the therapeutically active agents described herein may be formulated in any suitable vehicle for delivery. For instance, they may be placed into a pharmaceutically acceptable suspension, solution or emulsion. Suitable mediums include saline and liposomal preparations. More specifically, pharmaceutically acceptable carriers may include sterile aqueous of non-aqueous solutions, suspensions, and emulsions. Examples of non-aqueous solvents are propylene glycol, polyethylene glycol, vegetable oils such as olive oil, and injectable organic esters such as ethyl oleate. Aqueous carriers include but are not limited to water, alcoholic / aqueous solutions, emulsions or suspensions, including saline and buffered media. Intravenous vehicles include fluid and nutrient replenishers, electrolyte replenishers (such as those based on Ringer's dextrose), and the like.
[0215] Preservatives and other additives may also be present such as, for example, antimicrobials, antioxidants, chelating agents, and inert gases and the like.
[0216] A colloidal dispersion system may also be used for targeted gene delivery. Colloidal dispersion systems include macromolecule complexes, nanocapsules, microspheres, beads, and lipid- based systems including oil-in-water emulsions, micelles, mixed micelles, and liposomes.
[0217] An appropriate therapeutic regimen can be determined by a physician, and will depend on the age, sex, weight, of the subject, and the stage of the disease. As an example, for delivery of a nucleic acid sequence encoding a genetically engineered DNA recombining enzyme of the invention using a viral expression vector, each unit dosage of the genetically engineered DNA recombining enzyme expressing vector may comprise a composition including a viral expression vector in a pharmaceutically acceptable fluid at a concentration ranging from 1011to 1016viral genomes per ml, for example.
[0218] The effective dosages and the dosage regimens for administering a genetically engineered DNA recombining enzyme of the invention or of its subunits in the form of a recombinant polypeptide depend on the disease or condition to be treated and may be determined by the persons skilled in the art.
[0219] A physician or veterinarian having ordinary skill in the art may readily determine and prescribe the effective amount of the pharmaceutical composition required. For example, the physician or veterinarian could start doses of the therapeutically active agents of the invention employed in the pharmaceutical composition at levels lower than that required in order to achieve the desired therapeutic effect and gradually increase the dosage until the desired effect is achieved. In general, a suitable daily dose of a composition of the present invention will be that amount of the delivery system which is the lowest dose effective to produce a therapeutic effect. Such an effective dose will generally depend upon the factors described above. Administration may e.g. be intravenous, intramuscular, intraperitoneal, or subcutaneous, and for instance administered proximal to the site of the target. If desired, the effective daily dose of a pharmaceutical composition may be administered as two, three, four, five, six or more sub-doses administered separately at appropriate intervals throughout the day, optionally, in unit dosage forms. While it is possible for a delivery system of the present invention to be administered alone, it is preferable to administer the delivery system as a pharmaceutical composition as described above.
[0220] Further provided are kits comprising a therapeutically active agent as described herein. In one embodiment, the kit provides the therapeutically active agents prepared in one or more unitary dosage forms ready for administration to a subject, for example in a preloaded syringe or in an ampoule. In another embodiment, the therapeutically active agents are provided in a lyophilized form.
[0221] According to a further aspect, the present invention provides a method for treating or preventing a disease, the method comprising administering to a subject in need thereof a therapeutically effective amount of a protein having at least 80% identity to SEQ ID NO: 7, SEQ ID NO: 3, SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 8, or SEQ ID NO: 9, of the nucleic acid according to the invention, of the vector according to the invention, of the isolated host cell according to the invention, of the non-human host organism according to the invention, or of the pharmaceutical composition according to the invention.
[0222] The present inventors screened over 500 putative recombinase candidates of which a selection of 17 candidates were experimentally tested and eight novel recombinase systems were molecularly characterized in detail. To optimally use different recombinase systems, it is essential to characterize their applied properties. The determination of the in vivo recombinase activity may be crucial for planning an experiment. The eight newly characterized recombinases showed varying activity on their predicted target sites in bacteria, from over 90% in the case of YR1, YR4, YR6 to as low as around 15% in the case of YR8 and YR9. Possible explanations for this low activity may include the requirement for additional cofactors, or optimal temperature for more efficient recombination of the target site. The specificity of the recombinases was profiled, providing a valuable overview of possible cross recombination events on all target sites. For the applied use in higher organisms, the new recombinases were tested for their activity and compatibility in mammalian cells. All the recombinases showed high activity in a plasmid-based assay in HEK293T cells. Interestingly, YR8 showed high recombination rates in mammalian cells, whereas this recombinase had only weak activity in bacteria, suggesting that activity profiles of recombinases can vary in heterologous hosts.
[0223] For the applied use in heterologous cells, it is also important to consider pseudo-sites that may exist in the host genome. The human and mouse genomes were screened for lox-like sites for all known Cre-type recombinases. This information is useful in two ways: i) it provides an estimation on potential off-target sites that could potentially compromise an experiment and ii) it delivers potential endogenous target sites that might be usefully employed for genome engineering exercises (for instance for targeted delivery of DNA cargo into a safe harbor locus).
[0224] The disclosure and characterization of different naturally occurring recombinases is particularly useful to accelerate the development of novel enzymes via directed evolution. It has been demonstrated that shuffling of related genes can speed up the evolution process (Crameri et al., 1998). Hence, new recombinases with desired properties might become available in shorter time through family shuffling.
[0225] Examples
[0226] Example 1: Identification of putative recombinases and their target sites
[0227] Potential new Y-SSRs were identified by using tblastn from BLAST+ 2.10.1 (https: / / www.ncbi.nlm.nih.gov / books / NBK131777 / ) with the protein sequences of Cre, Vika (Karimova et al., 2012), Nigri, and Panto (Karimova et al. 2016) as references to search the NCBI nucleotide collection database (v5). The results were filtered for below 90% identity and a sequence length of 300 to 400 amino acids with GNU awk. Protein sequences were acquired with efetch (https: / / dataguide.nlm.nih.gov / edirect / efetch.html). Full genome sequences of the potential SSRs were gathered using bastdbcmd (part of BLAST+). Potential target sites were identified by searching the genome sequences 1000 bp upstream and downstream of the potential Y-SSRs for palindromes. Palindrome search was performed with EMBOSS palindrome (v6.6.0.0) (Rice et al., 2000) with a minimum palindromic length of 13 to 15 base pairs and a gap limit of 8 with one mismatch allowed. The program output was then converted to a tabular format with GNU awk and combined with the protein data in R with the dplyr package. Potential SSRs with a Eevenshtein distance (stringdist R package, https: / / joumal.r-project.org / archive / 2014 / RJ-2014-011 / index.html) below 10 to the references were removed. Potential SSRs were clustered with complete hierarchical clustering (base R) based on Eevenshtein distances and cluster groups were formed with a cut-off distance of 11. The same clustering method was also used on the potential half-sites of the palindromic sequences, here the cluster cut-off was a distance of 2. The top candidates for testing were chosen considering the clustering, their organism they were found in and their distance to the reference recombinases.
[0228] The phylogeny tree of known and putative recombinases was generated by performing an all- against-all pairwise sequence alignment of the protein sequences using EMBOSS needle (Needleman and Wunsch, 1970), followed by complete hierarchical clustering of the sequence dissimilarities. Visualization of the tree was done with R packages tidygraph and ggraph. For brevity, the tree was cut at the 98% sequence similarity, to represent almost identical proteins as a single node. Human and mouse genomic sequences with high similarity to potential target sites were identified using PatMaN (Priifer et al., 2008). The search was performed on half-sites only, allowing for up to 2 mismatches. If two genomic sequences matching the same half-site were found to be located on opposite strands, with a distance of 8 bp between them, they were called a potential target site of the respective recombinase. Genomic coordinates manipulation and sequence extraction steps were performed with the BEDTools suite (Quinlan and Hall, 2010).
[0229] Example 2: Plasmid construction
[0230] For expression in E. coli. codon-optimized DNA sequences of 17 predicted candidates were synthetized by Twist Biosciences and cloned into the pEVO vector (Fig. IB) via BsrGI and Xbal restriction sites (Buchholz and Stewart, 2001). Target sites were introduced via primers that were designed to carry the desired target sites and an overlap with the pEVO vector. The PCR fragment, that was generated when using the pEVO vector as template, was then cloned via Cold Fusion into a Bglll digested pEVO backbone (System Biosciences).
[0231] For expression of recombinases in mammalian cells, a lentiviral PGK-NLS-BFP plasmid (Fig. 5D) was used. Recombinase sequences were amplified from pEVO vectors and cloned via BsrGI and Xbal restriction sites. The before mentioned lentiviral vector, harboring recombinases was either transfected into HEK293T cells for transient expression, or used for virus production and infection for continuous expression.
[0232] For the construction of the recombination reporters, pCAG-loxP-mCherry-loxP-GFP ‘traffic light’ vector was used (Karpinski et al. 2016). Oligonucleotides containing respective target sites were used to amplify mCherry cassette and the fragment was later ligated to the pCAG plasmid via Nhel and Hindlll restriction sites.
[0233] Example 3: Recombination reporter assays
[0234] To visualize the recombination activity of recombinases on their predicted target sites, a plasmid-based assay was used as previously described (Karimova et al., 2012 and 2016) (Fig. 1A). In short, expression of the recombinases from the pBAD promoter was induced with L-arabinose (Sigma- Aldrich Chemie GmbH). Single clones containing the pEVO plasmid with the recombinase and recombination target sites were cultured overnight in 6 ml LB medium with 25 pg / ml Cm and either 0 or 100 pg / ml L-arabinose at 37°C and 200 rpm. The recombinase mediated excision event was detected by agarose gel electrophoresis after digestion with BsrGI and Sbfl restriction enzymes. The recombined plasmid is smaller in size compared to the non-recombined plasmid. Therefore, after gel electrophoresis a slower migrating non-recombined band (~5.0 kb), and a faster migrating (~4.3 kb) band for recombined plasmids can be seen (Fig. 1A).
[0235] To compare the recombination efficiency of the active recombinases on their native target sites, recombinase expression was induced with increasing concentrations of L-arabinose (0, 1, 10 or 100 pg / ml medium) overnight in 6 ml culture volume. The test digest was prepared for each induction level and recombination efficiency was estimated by agarose gel electrophoresis. To quantify the recombinase activity, the ratio of band intensities was determined using Fiji-ImageJ for image processing. The quantified recombination was plotted in R 4.0.3 with dplyr vl.0.7 and visualized with ggplot2 v3.3.5. All test digests were done in triplicates (n = 3).
[0236] For the mammalian recombination reporter assay, HEK293T cells were plated at a density of 2 x 105cells per well in 24-well dishes and cultured in glucose Dulbecco’s Modified Eagle’s Medium (DMEM, Gibco®), supplemented with 10% fetal bovine serum (Invitrogen), 1% Penicillin- Streptomycin (10,000 U / ml, Thermo Fisher). At a confluency of 70-80%, cells were co-transfected with pPGK-NLS-Recombinase-P2A-BFP plasmids expressing a recombinase and pCAG-lox- mCherry-lox-GFP traffic light reporters using Lipofectamine® 2000 Transfection Reagent (Invitrogen) according to manufacturer's instructions. Per well 0.5 pg of DNA (0.25 pg of each plasmid) and 2.5 pl of Lipofectamine® 2000 reagent diluted in 100 pl Opti-MEM® Reduced Serum Media each were used. On the next day, the media was changed and the cells were further cultured at 37 °C and 5% CO2. Upon recombination between the target sites, the mCherry cassette is excised and CAG promotor starts driving the expression of downstream green fluorescent protein (GFP). The cells were analyzed two days after transfection with fluorescent activated cell analysis and were then imaged with a fluorescent microscope (EVOS FL imaging system; Thermo Fisher Scientific)
[0237] Example 4: Fluorescent activated cell analysis
[0238] HEK293T were washed once with PBS and then detached using Trypsin (Gibco). The cells were then resuspended in Dulbecco’s Modified Eagle’s Medium (DMEM, Gibco®) and analyzed with the MACSQuant® VYB Flow Cytometer (Miltenyi). Analysis of the data was performed using Flow Jo™ 10 (BD).
[0239] Example 5: Overexpression studies in mammalian cells
[0240] For viral delivery, pPGK-Recombinase-P2A-BFP vectors were used to produce lentiviral particles as described previously (Suriin et al., 2020). NIH / 3T3 mouse fibroblasts were seeded at a density of 4 x 104cells per well in 24-well plates and grown at 37°C and 5% CO2. The next day, fibroblasts were transduced with the different lentiviruses with a MOI of 0,5 in order to achieve about 50% of infection rate. The percentage of BFP expressing cells from at least 2 x 104cells was tracked over the course of 15 days using MACSQuant® VYB Flow Cytometer (Miltenyi). The difference in the percentage of BFP cells at the last time point (day 15) and first time point was calculated and visualized with GraphPad Prism and statistical significance relative to BFP control was calculated by doing 1-way ANOVA test with 95% confidence interval (CI). Example 6: Cross-recombination assay: Nanopore sequencing
[0241] The recombinases and Cre, Vika, Panto, Dre (Anastassiadis et al., 2009), and VCre (Suzuki and Nakayama, 2011) were amplified from pEVO vectors, cleaned using the Isolate II PCR and Gel Cleanup Kit (Bioline) and mixed together in a 1: 1 ratio. All respective target sites were cloned into the pEVO vectors with Cold Fusion Cloning kit as previously described and the resulting vectors were also mixed with equal molar ratio. Both, the mix of recombinases and pEVO backbones were digested with BsrGI and Sbfl and ligated in a single reaction, thus creating a library of different recombinase / target site pairs. Plasmids were transformed in XL 1 -Blue electrocompetent E. coli cells and grown overnight with 100 pg / ml L-arabinose to induce recombinase expression. On the next day, plasmids were linearized with BsrGI and Seal and fragments carrying the recombinase sequence and target sites were isolated by agarose gel excision using the Isolate II PCR and Gel Cleanup Kit (Bioline). These DNA fragments were then prepared for nanopore sequencing with the SQK-LSK110 Kit according to the “Amplicons by Ligation” protocol on a MinlON R9.4.1 Flow Cell (Oxford Nanopore Technologies). Base calling of the sequence data was performed with guppy v5.0.7 on the high accuracy model (Oxford Nanopore Technologies). The sequence reads were then filtered for a read length of at least 1800 bp and a minimum mean phred score of ten with filtlong (https: / / github.com / rrwick / Filtlong). To identify the recombinases, the reads were aligned to the reference recombinase sequences with minimap2 v2.17 (Li, 2018). The target sites were identified using exonerate v2.2.0 using the affmedocal model (https: / / www.ebi.ac.uk / about / vertebrate- genomics / software / exonerate). The read ID and the matching references were then extracted from both alignments and combined in R with the dplyr package. Visualization of the data was performed with the R package ggplot2.
[0242] Example 7: New recombinases recombine their predicted target sites in bacteria
[0243] The activity of candidate recombinase was tested on the predicted target sequence. In brief, candidates were selected and their coding sequence individually cloned into the L-arabinose inducible pEVO recombination reporter vector (Buchholz and Stewart, 2001) harboring two copies of the respective predicted target sites (Fig. IB). The plasmids were then transformed into E. coli and cultured overnight in medium containing L-arabinose to induce recombinase expression. Upon expression, successful recombination leads to excision of a -700 bp DNA fragment from the plasmid. This size difference was visualized by agarose electrophoresis of linearized plasmids (Fig. 1A). Eight out of seventeen candidates showed activity on their predicted target site, evident by the appearance of -4 kb recombination bands (Fig. 2). Five of the candidates already showed efficient recombination in the samples without addition of L-arabinose to the medium (YR1, YR2, YR4, YR6 and YR12), indicating that these enzymes are active even when expressed at very low levels. Other recombinases (YR8, YR9, and YR11) recombined the plasmid only when L-arabinose was present in the growth medium, suggesting that they require a higher induction to become active in this assay. These results demonstrate successful identification of novel Y-SSRs and their respective target sites.
[0244] To characterize the active Y-SSRs in more detail, their recombination efficiencies was quantified at different expression levels in order to investigate dose response and to obtain a better side-by-side comparison of their efficiencies compared to the well-established recombinases Cre and Vika (Karimova et al., 2012). pEVO plasmids with desired recombinase / target site pairs were transformed into E. coli and were grown over night at different concentrations of L-arabinose to induce expression of the recombinase (Guzman et al., 1995). Extracted plasmid DNA was then assayed for recombination on agarose gels. Quantification of band intensities revealed that the new recombinases have different activity profiles on their respective target sites (Fig. 3). Although, most of the recombinases were highly active when grown at high L-arabinose concentrations, showing recombination rate between 87 and 100% (YRl, YR2, YR4, YR6 and YR11), they behaved quite differently when expressed with low L-arabinose concentrations. While YRl, YR2, YR4 and YR11 showed low recombination rates when induced at 1 or 10 pg / ml of L-arabinose, YR6 was highly active even at these low induction levels, with its activity profile resembling Cre and Vika (Fig, 3). Interestingly, the YRl 2 recombinase showed a mostly constant recombination rate raging from -50% at 0 pg / ml L-arabinose and peaking at 70% when induced with 100 pg / ml of L-arabinose. YR8 and YR9, on the other hand, showed the weakest activity and only recombined their target sites to 25% and 17%, respectively, at the highest L-arabinose concentration.
[0245] Recombinase YR9 was further subjected to substrate-linked directed evolution (Buchholz and Stewart, 2001), thereby managing to increase the activity of YR9 recombinase significantly. The best performing clone named YR9.2 (SEQ ID NO: 9) showed about seven-fold improvement in activity in bacteria (Fig. 7A), and eight-fold improvement in mammalian cells (Fig. 7C).
[0246] Example 8: Profiling target-site selectivity of Cre-like recombinases
[0247] SSRs with different sequence specificity are frequently used in combination to allow sophisticated genomic or synthetic biology experiments (Fed, 2007; Sheets et al., 2020; Merrick et al., 2018; Livet et al., 2007; Snippert et al., 2010). For such applications, it is important to know the specificity of the enzymes and to consider possible cross reactivity (Fenno et al., 2014; Weinberg et al., 2017). In order to test all possible combinations, high-throughput sequencing approach was developed in which the activity of known (Panto, Dre, Cre, Vika and VCre) and new Y-SSRs (YRl, YR2, YR4, etc.) can be quantified on all target sites in a single experiment. A two-step cloning scheme for producing all combinations of 13 recombinases and their respective 13 target sites on 169 (13x13) individual vectors was established. The 13 target sequences were cloned individually into the pEVO vector. The resulting constructs were then pooled and linearized to clone in a pool of the 13 recombinase coding sequences in one ligation reaction (Fig. 4A). After an overnight culture and induction of recombinase expression, plasmid DNA was retrieved and fragments carrying the recombinase sequence on the 3’-end and target site(s) on the 5’-end were cut out. Using the Oxford Nanopore Technologies’ long read sequencing platform, a total of 417,769 reads was obtained containing both, the specified recombinases and target sites. All possible 169 combinations of Y-SSRs and target sites were identified with a minimum coverage of 224 reads. Using this data, the recombination rates for the individual recombinases on all the target sites was calculated, providing a specificity profile for each recombinase (Fig. 4B).
[0248] Example 9: Activity of recombinases in human cells
[0249] The activity of the new Y-SSRs was tested in a human cell line. In brief, HEK293T cells were co-transfected with recombinase expression plasmids (as described in Example 3 for the mammalian recombination assay) alongside recombination reporter plasmids harboring the corresponding target sites (Fig. 5A). In the reporter plasmids, the mCherry cassette, driven by a CAG promoter, is flanked by the target (lox) sites. Upon recombination, the mCherry cassette is deleted from the plasmid and the CAG promoter then drives the expression of a GFP cassette (Fig. 5A). Hence, the activity of the recombinases on their predicted target sites can be visualized by fluorescent microscopy and quantified by flow cytometry. When co-transfection experiments were analyzed, all of the recombinases tested displayed GFP positive cells (Fig. 5B), demonstrating that these recombinases are active in HEK293T cells, whereas no GFP-positive cells were observed when co-transfections were done with an “empty” expression vector, lacking recombinase coding sequences (Fig. 5B). Flow cytometry analyses revealed that in almost all samples more than 90% of the cells that were cotransfected with the recombinase expression plasmid and the reporter (cells that were double, BFP and mCherry positive) were also GFP positive, indicating that these recombinases are highly active in this setting (Fig. 5C).
[0250] Example 10: Influence of recombinase expression on cell proliferation
[0251] Investigations of SSRs in heterologous hosts are important to define their applied properties. Recombinases could recognize cryptic (pseudo) recombination sites in a genome that might be recombined and lead to off-target effects, potentially resulting in growth arrest or apoptosis. Indeed, active pseudo-loxP sites have been described in the human and mouse genome (Thyagarajan et al., 2000). Consequently, impairment of cell proliferation may occur when Cre is overexpressed in human or mouse cells (Loonstra et al., 2001; Schmidt et al., 2000; Pugach et al., 2015).
[0252] To test for potential effects on cell proliferation of the newly identified Y-SSRs when they are overexpressed, lentiviral vectors were constructed, which allow co-expression of the recombinases and tagBFP (Fig. 6A). As controls, viral particles for overexpression of either Cre, an inactive Cre variant (CreY324F), Vika and tagBFP alone were used. NIH3T3 cells were infected and the change in percentage of BFP -positive cells was monitored for 15 days. A reduction in BFP-positive cells over time indicates a negative effect on cell proliferation due to recombinase overexpression (Schmidt et al., 2000; Pugach et al., 2015). The number of BFP-positive cells progressively dropped when Cre recombinase was tested in this assay, while the catalytically inactive version of Cre had no effect (Fig. 6B). In comparison to Cre, YR4 and YR8 showed a less pronounced decrease in the percentage of BFP positive cells (-15%; p<0.0001, and -17%; p<0.0001, respectively), suggesting that overexpression of these recombinases slightly inhibits cell proliferation (Fig. 6B). In contrast, the percentage of BFP-positive cells did not significantly change in cells expressing the other recombinases (Fig. 6B), indicating that overexpression of these recombinases is well tolerated in the cells.
[0253] Cited non-patent literature
[0254] Anastassiadis,K., Fu,J., Patsch,C., Hu,S., Weidlich,S., Duerschke,K., Buchholz, F., Edenhofer,F. and Stewart, A. F. (2009) Dre recombinase, like Cre, is a highly efficient site-specific recombinase in E. coli, mammalian cells and mice. Dis Model Meeh, 2, 508-515.
[0255] Anderson, R.P., Voziyanova,E. and Voziyanov,Y. (2012) Flp and Cre expressed from Flp-2A-Cre and Flp-IRES-Cre transcription units mediate the highest level of dual recombinase-mediated cassette exchange. Nucleic Acids Res, 40, e62-e62.
[0256] Anzalone, A.V., Gao,X.D., Podracky,C.J., Nelson, A.T., Koblan,L.W., Raguram,A., LevyJ.M., Mercer, J.A.M. and Liu,D.R. (2022) Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing. Nat Biotechnol, 40, 731-740.
[0257] Anzalone, A.V., Randolph, P.B., Davis, J. R., Sousa, A.A., Koblan,L.W., LevyJ.M., Chen,P.J., Wilson, C., Newby, G.A., Raguram,A., et al. (2019) Search-and-replace genome editing without double-strand breaks or donor DNA. Nature, 576, 149-157.
[0258] Buchholz, F., and Stewart, A.F. (2001). Alteration of Cre recombinase site specificity by substrate- linked protein evolution. Nat Biotechnol 19, 1047-1052.
[0259] Crameri,A., Raillard, S. -A., Bermudez, E. and Stemmer,W.P.C. (1998) DNA shuffling of a family of genes from diverse species accelerates directed evolution. Nature, 391, 288-291.
[0260] Duyne,G.D.V. (2001) A Structural View of Cre- loxP Site-Specific Recombination. Annu Rev Bioph Biom, 30, 87-104.
[0261] Feil,R. (2007) Conditional Mutagenesis: An Approach to Disease Models. Handb Exp Pharmacol, 10.1007 / 978-3-540-35109-2_l.
[0262] Fenno,L.E., Mattis, J., Ramakrishnan,C., Hyun,M., Lee,S.Y., He,M., TucciaroneJ., Selimbeyoglu,A., Berndt, A., Grosenick,L., et al. (2014) Targeting cells with single vectors using multiple-feature Boolean logic. Nat Methods, 11, 763-772.
[0263] Gaudelli,N.M., Komor,A.C., Rees,H.A., Packer, M.S., Badran, A.H., Bryson, D. I. and Liu,D.R. (2017) Programmable base editing of A«T to G*C in genomic DNA without DNA cleavage. Nature, 551, 464-471.
[0264] Guzman, L.M., Belin, D., Carson, M. J. and Beckwith, J. (1995) Tight regulation, modulation, and high- level expression by vectors containing the arabinose PBAD promoter. J Bacteriol, 177, 4121— 4130.
[0265] Karimova, M., Abi-Ghanem,J., Berger, N., Surendranath,V., Pisabarro,M.T. and Buchholz, F. (2012) Vika / vox, a novel efficient and specific Cre / loxP-like site-specific recombination system. Nucleic Acids Res, 41, e37-e37.
[0266] Karimova, M., Splith,V., Karpinski, J., Pisabarro,M.T. and Buchholz, F. (2016) Discovery of Nigri / nox and Panto / pox site-specific recombinase systems facilitates advanced genome engineering. Sci Rep-uk, 6, 30130.
[0267] Karpinski, J., Hauber,!., Chemnitz,!., Schafer, C., Paszkowski-Rogacz,M., Chakraborty,D.,
[0268] Beschomer,N., Hofmann-Sieber, H., Lange,U.C., Grundhoff,A., et al. (2016) Directed evolution of a recombinase that excises the provirus of most HIV-1 primary isolates with high specificity. Nat Biotechnol, 34, 401-409.
[0269] Komor,A.C., Kim,Y.B., Packer, M.S., ZurisJ.A. and Liu,D.R. (2016) Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature, 533, 420-424.
[0270] Li,H. (2018) Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics, 34, 3094-3100.
[0271] Livet ., Weissman, T.A., Kang,H., Draft, R.W., Lu, J., Bennis, R.A., SanesJ.R. and LichtmanJ.W. (2007) Transgenic strategies for combinatorial expression of fluorescent proteins in the nervous system. Nature, 450, 56-62.
[0272] Loonstra,A., Vooijs,M., Beverloo,H.B., Allak,B.A., Drunen,E. van, Kanaar,R., Berns, A. and Jonkers,!. (2001) Growth inhibition and DNA damage induced by Cre recombinase in mammalian cells. Proc National Acad Sci, 98, 9209-9214.
[0273] Meinke,G., Bohm, A., Hauber, J., Pisabarro,M.T. and Buchholz, F. (2016) Cre Recombinase and Other Tyrosine Recombinases. Chem Rev, 116, 12785-12820.
[0274] Merrick, C. A., Zhao, J. and Rosser, S. J. (2018) Serine Integrases: Advancing Synthetic Biology. Acs Synth Biol, 7, 299-310.
[0275] Minorikawa,S. and Nakayama, M. (2011) Recombinase-mediated cassette exchange (RMCE) and BAC engineering via VCre / VloxP and SCre / SloxP systems. Biotechniques, 50, 235-246.
[0276] Needleman, S.B. and Wunsch,C.D. (1970) A general method applicable to the search for similarities in the amino acid sequence of two proteins. J Mol Biol, 48, 443-453.
[0277] Priifer,K., Stenzel, U., Dannemann,M., Green, R.E., Lachmann,M. and Kelso, J. (2008) PatMaN: rapid alignment of short sequences to large databases. Bioinformatics, 24, 1530-1531.
[0278] Pugach,E.K., Richmond,? .A., AzofeifaJ.G., Dowell, R.D. and Leinwand,L.A. (2015) Prolonged Cre expression driven by the a-myosin heavy chain promoter can be cardiotoxic. J Mol Cell Cardiol, 86, 54-61.
[0279] Quinlan, A. R. and Hall, I. M. (2010) BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics, 26, 841-842.
[0280] Sauer, B. and Henderson, N. (1988) Site-specific DNA recombination in mammalian cells by the Cre recombinase of bacteriophage Pl. Proc National Acad Sci, 85, 5166-5170.
[0281] Schmidt, E.E., Taylor, D.S., PriggeJ.R., Barnett, S. and Capecchi,M.R. (2000) Illegitimate Cre- dependent chromosome rearrangements in transgenic mouse spermatids. Proc National Acad Sci, 97, 13702-13707.
[0282] Sternberg, N. and Hamilton, D. (1981) Bacteriophage Pl site-specific recombination I. Recombination between loxP sites. J Mol Biol, 150, 467-486.
[0283] Suriin,D., Schneider, A., Mircetic ., Neumann, K., Lansing, F., Paszkowski-Rogacz,M., Hanchen,V., Lee-Kirsch, M.A. and Buchholz, F. (2020) Efficient Generation and Correction of Mutations in Human iPS Cells Utilizing mRNAs of CRISPR Base Editors and Prime Editors. Genes-basel, 11, 511.
[0284] Rice,P., Longden,!. and Bleasby,A. (2000) EMBOSS: The European Molecular Biology Open Software Suite. Trends Genet, 16, 276-277.
[0285] Sheets, M.B., Wong,W.W. and Dunlop, M. J. (2020) Light-Inducible Recombinases for Bacterial Optogenetics. Acs Synth Biol, 9, 227-235.
[0286] Snippert,H.J., Flier, L.G. van der, Sato,T., Es,J.H. van, Bom,M. van den, Kroon-Veenboer,C., Barker, N., Klein, A.M., Rheenen . van, Simons, B.D., et al. (2010) Intestinal Crypt Homeostasis Results from Neutral Competition between Symmetrically Dividing Lgr5 Stem Cells. Cell, 143, 134-144.
[0287] Suzuki, E. and Nakayama, M. (2011) VCre / VloxP and SCre / SloxP: new site-specific recombination systems for genome engineering. Nucleic Acids Res, 39, e49-e49.
[0288] Thyagarajan,B., Guimaraes, M.J., Groth, A. C. and Calos, M.P. (2000) Mammalian genomes contain active recombinase recognition sites. Gene, 244, 47-54.
[0289] Weinberg, B.H., Pham,N.T.H., Caraballo, L.D., Lozanoski,T., Engel, A., Bhatia, S. and Wong,W.W. (2017) Large-scale design of robust genetic circuits with multiple inputs and outputs for mammalian cells. Nat Biotechnol, 35, 453-462.
Claims
Claims1. A method for producing a site-specific DNA-recombination, the method comprising the steps of: a) contacting a nucleic acid comprising at least a first and a second recognition site which are essentially identical or essentially reverse complementary to each other with a protein having recombinase activity, and b) allowing the protein having recombinase activity to produce the site-specific DNA- recombination, wherein a recognition site comprises a first half-site, a spacer and a second half-site, and wherein essentially identical or essentially reverse complementary to each other means that the nucleotide sequence of the first and the second half-site in the first recognition site may deviate in up to two nucleotides from the nucleotide sequence of the first and the second half-site in the second recognition site, wherein(i) the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 7, and wherein the at least two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 16 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 16; or(ii) the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 3, and wherein the at least two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 12 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 12; or(iii) the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 1, and wherein the at least two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 10 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 10; or(iv) the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 2, and wherein the at least two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 11 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 11; or(v) the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 4, and wherein the at least two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 13 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 13; or(vi) the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 5, and wherein the at least two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 14 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 14; or(vii) the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 6, and wherein the at least two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15; or(viii) the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 8, and wherein the at least two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 17 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 17; or(ix) the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 9, and wherein the at least two recognition sites comprise a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or to a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15.
2. The method according to claim 1, wherein the protein having recombinase activity comprises at least two protein monomers.
3. The method according to claim 1 or 2, wherein the nucleic acid sequence that is recombined is present in a cell, preferably further comprising the step of introducing into the cell a nucleic acid encoding the protein having recombinase activity, or wherein the cell comprises a nucleic acid encoding the protein having recombinase activity.
4. The method according to claim 3, wherein the nucleic acid encoding the protein having recombinase activity comprises a regulatory nucleic acid sequence, and wherein the expression of thenucleic acid encoding the protein having recombinase activity is regulated by the regulatory nucleic acid sequence, and / or wherein the cell is a eukaryotic or a bacterial cell.
5. Use of a protein having at least 80% identity to SEQ ID NO: 7, SEQ ID NO: 3, SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 8, or SEQ ID NO: 9 for producing a site-specific DNA-recombination.
6. Use of a protein having recombinase activity for catalyzing a site-specific DNA- recombination at recognition sites that are essentially identical or essentially reverse complementary to each other, wherein a recognition site comprises a first half-site, a spacer and a second half-site, and wherein essentially identical or essentially reverse complementary to each other means that the nucleotide sequence of the first and the second half-site in the first recognition site may deviate in up to two nucleotides from the nucleotide sequence of the first and the second half-site in the second recognition site, wherein:(i) the protein comprises an amino acid sequence having at least 80% identity to SEQ ID NO:
7. and wherein the at least one recognition site comprises a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 16 or a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 16; or(ii) the protein comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 3, and wherein the at least one recognition site comprises a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 12 or a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 12; or(iii) the protein comprises an amino acid sequence having at least 80% identity to SEQ ID NO:1, and wherein the at least one recognition site comprises a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 10 or a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 10; or(iv) the protein comprises an amino acid sequence having at least 80% identity to SEQ ID NO:2, and wherein the at least one recognition site comprises a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 11 or a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 11; or(v) the protein comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 4, and wherein the at least one recognition site comprises a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 13 or a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 13; or(vi) the protein comprises an amino acid sequence having at least 80% identity to SEQ ID NO:5, and wherein the at least one recognition site comprises a nucleic acid sequence according to orreverse complementary to SEQ ID NO: 14 or a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 14; or(vii) the protein comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 6, and wherein the at least one recognition site comprises a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15; or(viii) the protein comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 8, and wherein the at least one recognition site comprises a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 17 or a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 17; or(ix) the protein comprises an amino acid sequence having at least 80% identity to SEQ ID NO: 9, and wherein the at least one recognition site comprises a nucleic acid sequence according to or reverse complementary to SEQ ID NO: 15 or a functional mutant thereof, wherein the functional mutant comprises a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15.
7. A nucleic acid having a length of not more than 40 base pairs and comprising:(i) a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 16 or a nucleic acid sequence reverse complementary thereto; or(ii) a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 12 or a nucleic acid sequence reverse complementary thereto; or(iii) a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 10 or a nucleic acid sequence reverse complementary thereto; or(iv) a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 11 or a nucleic acid sequence reverse complementary thereto; or(v) a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 13 or a nucleic acid sequence reverse complementary thereto; or(vi) a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 14 or a nucleic acid sequence reverse complementary thereto; or(vii) a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 15 or a nucleic acid sequence reverse complementary thereto; or(viii) a nucleic acid sequence having at least 60% sequence identity to SEQ ID NO: 17 or a nucleic acid sequence reverse complementary thereto.
8. A vector comprising at least one and preferably at least two essentially identical or essentially reverse complementary nucleic acids according to claim 7, wherein a DNA segment is preferably flanked by the two identical or reverse complementary nucleic acids, the vector preferably further comprises a nucleic acid encoding a protein having recombinase activity, wherein the protein havingrecombinase activity comprises an amino acid sequence having at least 80% identity to any one of SEQ ID NOs: 1 to 9.
9. A vector comprising a nucleic acid encoding a protein having recombinase activity, wherein the protein having recombinase activity comprises an amino acid sequence having at least 80% identity to any one of SEQ ID NOs: 1 to 9.
10. Use of the vector according to any one of claims 8 or 9 in a method according to any one of claims 1 to 4.
11. A protein having at least 80% identity to SEQ ID NO: 7, SEQ ID NO: 3, SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: SEQ ID NO: 8, or SEQ ID NO: 9, or a vector according to claim 8 or 9 for use in medicine, preferably for use in treating a genetic disease or disorder in a subject, more preferably wherein the genetic disease or disorder is characterized by modification of the subject’s genome.
12. An isolated host cell, comprising the following recombinant DNA fragments:(i) at least one and preferably at least two nucleic acids according to claim 7; and / or(ii) a vector according to any one of claims 8 or 9.
13. The isolated host cell according to claim 12, further comprising:(i) a nucleic acid encoding for a protein having recombinase activity, wherein the protein comprises an amino acid sequence having at least 80% identity to any one of SEQ ID NOs: 1 to 9; or(ii) a vector according to any one of claims 8 or 9.
14. A non-human host organism, comprising:(i) at least one and preferably at least two nucleic acids according to claim 7; or(ii) a vector according to any one of claims 8 or 9.
15. A pharmaceutical composition comprising a protein having at least 80% identity to SEQ ID NO: 7, SEQ ID NO: 3, SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 8, or SEQ ID NO: 9, the nucleic acid according to claim 7, a vector according to any one of claims 8 or 9, the isolated host cell according to claim 12 or 13, or the non-human host organism according to claim 14, and optionally a pharmaceutically acceptable excipient.