Novel recombinant tet3 enzymes

ZA202608519APending Publication Date: 2026-09-30QUGEN GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
ZA202608519
Authority / Receiving Office
ZA · ZA
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-22
Filing Date
2026-08-26
Publication Date
2026-09-30

AI Technical Summary

Technical Problem

Current methods for sequencing 5-methylcytidine (mdC) in nucleic acids, such as bisulfite sequencing and 3rd-generation sequencing, face challenges like DNA fragmentation, cumbersome protocols, and limited oxidation capabilities of existing TET enzymes, making accurate and efficient mdC sequencing difficult, especially for early cancer diagnostics.

Method used

A recombinant Ten Eleven Translocation 3 (TET3) enzyme with specific amino acid sequences and a heterologous connecting domain, optimized for efficient oxidation of mdC to 5-carboxycytidine (cadC), enabling accurate sequencing without bisulfite addition and providing high selectivity and quantitative oxidation.

Benefits of technology

The recombinant TET3 enzyme achieves selective and quantitative oxidation of mdC, facilitating accurate sequencing with reduced DNA fragmentation and streamlined protocols, suitable for early cancer diagnostics and other applications.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

NOT VISIBLE DUE TO STATUS OF PATENT
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Novel recombinant TET3 enzymes

[0002] Field of the invention

[0003] The present invention refers to a recombinant Ten Eleven Translocation 3 (TET3) enzyme, a nucleic acid molecule encoding the TET3 enzyme, a vector comprising the nucleic acid molecule and a host cell transformed or transfected with the vector, the use of the recombinant TET3 enzyme for sequencing of a nucleic acid substrate, a method for sequencing a nucleic acid substrate and a reagent for sequencing a nucleic acid substrate.

[0004] Background of the invention

[0005] DNA methylation is one of the important modifications of genomes in eukaryotic cells and plays a key role in mammalian cell development, differentiation, genomic imprinting and regulation of gene expression. It has important biological significance and is associated with human diseases (including tumor diseases) that are closely related to the occurrence and development processes. Therefore, it is of great value and significance to discover and establish genomic DNA modification profiles related to disease processes.

[0006] The presence of the 5thnucleoside 5-methyldeoxycytidine (mdC) either in promoter regions or in the gene body influences the transcriptional state of the corresponding gene

[0001] , The presence of the methylated nucleobase mdC in promoter regions typically silences the gene, while unmethylated promoters stand for more active transcription. The identification of mdC within genes allows consequently to characterize the transcriptional state of the gene of interest, which is for example important in order to characterize and identify tumor cells [2-3], in which oncogenes are typically wrongly switched on and tumor suppressor genes are aberrantly switched off. Sequencing of mdC with as little input material as possible is consequently highly desired in order to establish a new area of early tumor diagnostic, called liquid biopsy [4]- i Since sequencing of mdC positions in the genome is of paramount importance, especially for early cancer diagnostics in order to determine incorrect expression of genes, many attempts have been made to provide an accurate sequencing of mdC.

[0007] Until today, mdC sequencing is predominantly performed with bisulfite. Treatment of genomic DNA with bisulfite at 64°C converts all non-methylated cytidines into uracil, while mdC remains intact. Following PCR and sequencing and by comparing the obtained reads with a reference genome therefore allows to determine the position of the mdCs in the genome. The main problem associated with this method, however, is that a large amount of the genomic input DNA does not survive the harsh bisulfite treatment conditions due to extensive DNA fragmentation. Thus, this method cannot be applied for limited input samples.

[0008] This limitation may be overcome by extensive PCR-based amplification of the nondegraded DNA. For example, C. Liu et al. describe a mdC-specific whole-genome amplification system for simultaneous DNA amplification and methylation in a one-pot, primer-free reaction [5], This method is specifically directed to samples having limited input materials in order to increase DNA amount prior to sequencing. After amplification, the amplified products are subjected to standard bisulfite sequencing.

[0009] Another caveat is that the handling of the bisulfite sequencing protocol is cumbersome and error prone. Milder methods that are currently being developed like, EM-seq, make use of the deaminating enzyme AP0BEC3A (A3A), which deaminates dC to dU [6], Deamination of all dC bases to dU, however, creates the problem that it decreases the genome complexity from a four letter code to a now only three nucleobase code (dA, dG and dU), which, plus the base mdC, makes the mathematical sequence assembly challenging.

[0010] An alternative approach for sequencing mdC is based on 3rdgeneration sequencing where one directly reads out sequences without the need for a PCR step. Today, all 3rdgeneration single molecule sequencing tools such as Nanopore or SMRT sequencing indeed allow a direct read-out of mdC [7], These methods, however, are still at an early stage and the obtained signal differences between the sequencing signals obtained for dC and mdC are often very small. This further requires cumbersome deconvolution of the data with substantial bioinformatics [8-9],

[0011] CN 111220760 A describes a method for determining the concentration of mdC in plasma DNA. The method adopts a liquid chromatography-mass spectrometry system for determination, and comprises the following steps: taking a DNA sample extracted from plasma; adding a quantity of formic acid and carrying out reacting in a closed reaction bottle; carrying out cooling to room temperature, carrying out volatilizing in vacuum, adding methanol for dissolving, and carrying out centrifuging at a high speed; and taking supernatant for chromatographic column separation, and carrying out detecting by using a mass spectrum detector. However, the method needs heating the DNA samples to 70°C. Further, a strong acidic environment is needed, which finally hydrolyses mdC groups.

[0012] Mild epigenetic mdC sequencing methods, which would circumvent cytidine deamination are consequently highly desirable in order to provide a non-invasive, accurate, time, cost and resource saving early tumor diagnosis.

[0013] Limitations for mdC sequencing could be overcome by 5-carboxycytidine (cadC) sequencing. cadC exhibits in contrast to dC an additional carboxy group, which is negatively charged under neutral pH-conditions. This may provide a large signal difference between the neutral mdC and the negatively charged cadC. T. A. Clark et al. describe a method of detecting cadC instead of mdC by SMRT sequencing

[0010] ,

[0014] This oxidation possibility has also been employed to develop new sequencing methods for mdC and 5-hydroxymethylcytidine (hmdC) [11 -13] providing milder reaction alternatives to the harsh conditions in bisulfite sequencing

[0014] , Mapping of mdC and hmdC is required in order to characterize the epigenetic state of genes, which is for example not only the basis for determining the age and differentiation of tissues, but it is also considered to enable the detection of tumor DNA in blood samples in a process called liquid biopsy in the future

[0015] , For sequencing of mdC and hmdC a method called TET-assisted pyridine borane sequencing (TAPS) was developed. Here mdC, hmdC and also 5-formylcytidine (fdC) bases in the genome are oxidized with a TET enzyme to cadC, which is followed by conversion of cadC with pyridine borane to give dihydrouridine (DHU)

[0016] , For identification, hmdC can be protected from the TET- mediated-oxidation through glycosylation (=5 gmdC) by p-glycosyltransferases (0-GT). After the enzymatic oxidation of mdC to cadC by TET, cadC is subsequently either converted with bisulfite to deoxyuracil (dll) (TET-assisted bisulfite sequencing (TAB- seq))

[0017] or it is again reduced with pyridine borane to DHU (TAPSp)

[0018] , The position of dU / DHU are then decoded as “T” in the subsequent sequencing step, while 5gmdC is read as “C”. This all together provides positional information of mdC of hmdC in the genome. In an alternative approach called enzymatic methyl sequencing (EM-Seq), mdC is again first oxidized by TET enzymes to a give mixture of hmdC, fdC and cadC. Finally, the unconverted dC is deaminated to dU using Apolipoprotein B mRNA editing enzymes (APOBEC, method 3). The position of dU is next identified by sequencing and comparison to a reference genomic sequence, without the need to use bisulfite or pyridine borane [19, 20], For identification, hmC can be again protected by glycosylation, prior to TET-oxidation and APOBEC-mediated deamination (E5hmC- seq) [17, 20],

[0015] For all these methods, modified TET enzymes are required for the most efficient oxidation of mdC. This is a challenge because the available and described TET enzymes, which are mostly derived from TET2, have limited oxidation capabilities, making the quantitative oxidation of all mdC in a genome to cadCs one of the most difficult steps. For example, in the EM-Seq

[0019] method it would be desirable to have a TET protein that converts all the mdCs present in the genome to be sequenced with high efficiency into cadC. This is difficult to achieve with the currently available TET enzymes.

[0016] Summary of the Invention

[0017] The present disclosure provides a truncated but robust, recombinant Ten Eleven Translocation 3 (TET3) enzyme comprising: a) a first TET3 domain comprising an amino acid sequence having amino acids 696-1048 or amino acids 697-1048 of UniProt No. A0A5K1VVP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 85% over the whole length thereof; b) a second TET3 domain comprising an amino acid sequence having amino acids 1509-1599 of UniProt No. A0A5K1VVP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 85% over the whole length thereof, and c) a connecting domain located between domain (a) and domain (b) comprising an amino acid sequence of at least 5 amino acids which is heterologous to a TET3 enzyme.

[0018] A further aspect of the present invention relates to a nucleic acid molecule encoding a TET3 enzyme as described above, optionally linked to an expression control sequence.

[0019] Still a further aspect of the invention relates to a vector comprising a nucleic acid molecule as described above.

[0020] Still a further aspect of the invention relates to a host cell transformed or transfected with a vector as described above.

[0021] Still a further aspect of the invention relates to a use of a recombinant TET3 enzyme as described above for the sequencing of a nucleic acid substrate, particularly for the sequencing of a nucleic acid substrate comprising mdC and / or 5-methylcytdine (mC) nucleotides.

[0022] Still a further aspect of the invention relates to a method for sequencing a nucleic acid substrate comprising mdC and / or mC nucleotides, wherein the method comprises the steps:

[0023] (A) contacting the nucleic acid substrate with a recombinant TET3 enzyme as described above under conditions, where mdC and / or mC nucleotides in the nucleic acid substrate are oxidized, e.g., to 5-hydroxymethyl-2'-deoxycytidine (hmdC), 5-formyl-2'-deoxycytidine (fdC), cadC, 5-hydroxymethylcytidine (hmC), 5-formylcytidine (fC), and / or 5-carboxycytidine (caC), (B) optionally isolating the oxidized nucleic acid substrate, and

[0024] (C) determining the amount and / or position of mdC and / or mC nucleotides in the nucleic acid substrate.

[0025] A further aspect of the present invention relates to a reagent for sequencing a nucleic acid substrate comprising mdC and / or mC nucleotides comprising:

[0026] (A) a recombinant TET3 enzyme as described above, and

[0027] (B) an oxidizing agent.

[0028] Embodiments of the invention

[0029] In the following, specific embodiments of the invention are disclosed as follows:

[0030] 1 . A recombinant Ten Eleven Translocation 3 (TET3) enzyme, comprising: a) a first TET3 domain comprising an amino acid sequence having amino acids 696-1048 or amino acids 697-1048 of UniProt No. A0A5K1WP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 85% over the whole length thereof; b) a second TET3 domain comprising an amino acid sequence having amino acids 1509-1599 of UniProt No. A0A5K1WP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 85% over the whole length thereof, and c) a connecting domain located between domain (a) and domain (b) comprising an amino acid sequence of at least 5 amino acids which is heterologous to a TET3 enzyme.

[0031] 2. The recombinant TET3 enzyme of embodiment 1 , wherein domain (a) further comprises C-terminal extension of up to 5 amino acids, particularly with an amino acid sequence selected from IQKEK (SEQ ID NO: 5), LQKEK (SEQ ID NO. 6) or any partial sequence thereof. The recombinant TET3 enzyme of embodiment 1 or 2, wherein domain (b) further comprises C-terminal extension of up to 5 amino acids, particularly with an amino acid sequence selected from AARLG (SEQ ID NO: 7) or any partial sequence thereof. The recombinant TET3 enzyme of any one of embodiments 1 -3, wherein domain

[0032] (a) comprises:

[0033] (i) an amino acid sequence having amino acids 696-1048 or amino acids 697- 1048 of UniProt No. A0A5K1 VVP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 90% over the whole length thereof;

[0034] (ii) an amino acid sequence having amino acids 831-1183 of UniProt No. Q8BG87 (SEQ ID NO: 2) or an amino acid sequence having an identity of at least 90% over the whole length thereof;

[0035] (iii) an amino acid sequence having amino acids 824-1175 of UniProt No. 043151 (SEQ ID NO: 3) or an amino acid sequence having an identity of at least 90% over the whole length thereof; or

[0036] (iv)an amino acid sequence having amino acids 953-1304 of UniProt No. A0JP82 (SEQ ID NO: 4) or an amino acid sequence having an identity of at least 90% over the whole length thereof. The recombinant TET3 enzyme of any one of embodiments 1 -4, wherein domain

[0037] (b) comprises:

[0038] (i) an amino acid sequence having amino acids 1509-1599 of UniProt No. A0A5K1 WP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 90% over the whole length thereof;

[0039] (ii) an amino acid sequence having amino acids 1644-1734 of UniProt No. Q8BG87 (SEQ ID NO: 2) or an amino acid sequence having an identity of at least 90% over the whole length thereof;

[0040] (iii) an amino acid sequence having amino acids 1635-1726 of UniProt No. 043151 (SEQ ID NO: 3) or an amino acid sequence having an identity of at least 90% over the whole length thereof; or

[0041] (iv)an amino acid sequence having amino acids 1742-1833 of UniProt No. A0JP82 (SEQ ID NO: 4) or an amino acid sequence having an identity of at least 90% over the whole length thereof. 6. The recombinant TET3 enzyme of any one of embodiments 1 -5, wherein the connecting domain has a length of 5-100, preferably 5-30, more preferably 10-20, and most preferably about 15 amino acids.

[0042] 7. The recombinant TET3 enzyme of any one of embodiments 1 -6, wherein the connecting domain consists of amino acids selected from G, S, A, T, E, N, D, and / or K.

[0043] 8. The recombinant TET3 enzyme according to any one of embodiments 1 -7, wherein the connecting domain has the amino acid sequence

[0044] [(Gn)Xm]r, wherein X is in each occurrence independently selected from S, E or D, n is 1 -5, particularly 2-4, m is 0-5, particularly 1 -3, r is 1 -5, particularly 2-4.

[0045] 9. The recombinant TET3 enzyme of any one of embodiments 1 -8, wherein the connecting domain has the amino acid sequence of SEQ ID NO: 8: GGGGSGGGGSGGGGS or SEQ ID NO: 9: GGGSGGGGSGGGGE or SEQ ID NO: 10 GGGGSGGGGSGGGGD.

[0046] 10. The recombinant TET3 enzyme of any one of embodiments 1 -9, which does not contain amino acid portions of a TET3 enzyme having a length of 20 amino acids or more, of 10 amino acids or more, and particularly of 6 amino acids or more outside domain (a) and domain (b).

[0047] 11 . The recombinant TET3 enzyme of any one of embodiments 1 -10, which does not contain amino acid portions of SEQ ID NO: 1 having a length of

[0048] 20 amino acids or more, of 10 amino acids or more, and particularly of 6 amino acids or more outside amino acids 696-1048 or amino acids 697-1048 and amino acids 1509-1599 of UniProt No. A0A5K1 WP6; and / or which does not contain amino acid portions of SEQ ID NO: 2 having a length of 20 amino acids or more, of 10 amino acids or more, and particularly of 6 amino acids or more outside amino acids 831 -1183 and amino acids 1644-1734 of UniProt No. Q8BG87; and / or which does not contain amino acid portions of SEQ ID NO: 3 having a length of 20 amino acids or more of 10 amino acids or more, and particularly of 6 amino acids or more outside amino acids 824-1175 and amino acids 1635-1726 of UniProt No. 043151 ; and / or which does not contain amino acid portions of SEQ ID NO: 4 having a length of 20 amino acids or more, of 10 amino acids or more, and particularly of 6 amino acids or more outside amino acids 953-1304 and amino acids 1742-1833 of UniProt No. A0JP82.

[0049] 12. The recombinant TET3 enzyme of any one of embodiments 1 -11 , which does not contain:

[0050] (1 ) the amino acids corresponding to positions 1 -695 or positions 1 -696 of UniProt No. A0A5K1 WP6 or an amino acid sequence having an identity of at least 90% over the whole length thereof;

[0051] (2) the amino acids corresponding to positions 1049-1508 of UniProt No. A0A5K1 WP6 or an amino acid sequence having an identity of at least 90% over the whole length thereof; and / or

[0052] (3) the amino acids corresponding to positions 1600-1668 of UniProt No. A0A5K1 WP6 or an amino acid sequence having an identity of at least 90% over the whole length thereof.

[0053] 13. The recombinant TET3 enzyme of any one of embodiments 1 -12, which does not contain:

[0054] (1 ) the amino acids corresponding to positions 1 -830 of UniProt No. Q8BG87 or an amino acid sequence having an identity of at least 90% over the whole length thereof;

[0055] (2) the amino acids corresponding to positions 1184-1643 of UniProt No. Q8BG87 or an amino acid sequence having an identity of at least 90% over the whole length thereof; and / or

[0056] (3) the amino acids corresponding to positions 1735-1803 of UniProt No. Q8BG87 or an amino acid sequence having an identity of at least 90% over the whole length thereof. The recombinant TET3 enzyme of any one of embodiments 1 -13, which does not contain:

[0057] (1 ) the amino acids corresponding to positions 1 -823 of UniProt No. 043151 or an amino acid sequence having an identity of at least 90% over the whole length thereof;

[0058] (2) the amino acids corresponding to positions 1176-1634 of UniProt No. 043151 or an amino acid sequence having an identity of at least 90% over the whole length thereof; and / or

[0059] (3) the amino acids corresponding to positions 1727-1795 of UniProt No. 043151 or an amino acid sequence having an identity of at least 90% over the whole length thereof. The recombinant TET3 enzyme of any one of embodiments 1 -14, which does not contain:

[0060] (1 ) the amino acids corresponding to positions 1 -952 of UniProt No. A0JP82 or an amino acid sequence having an identity of at least 90% over the whole length thereof;

[0061] (2) the amino acids corresponding to positions 1305-1741 of UniProt No. A0JP82 or an amino acid sequence having an identity of at least 90% over the whole length thereof; and / or

[0062] (3) the amino acids corresponding to positions 1834-1901 of UniProt No. A0JP82 or an amino acid sequence having an identity of at least 90% over the whole length thereof. The recombinant TET3 enzyme of any one of embodiments 1 -15, wherein amino acid T at position 940 of UniProt No. A0A5K1WP6 (SEQ ID NO: 1 ) is substituted with another amino acid, particularly with an amino acid selected from A, G, K and / or V; or amino acid T at position 1075 of UniProt No. Q8BG87 is substituted with another amino acid, particularly with an amino acid selected from A, G, K, and / or V; or amino acid T at position 1067 of UniProt No. 043151 is substituted with another amino acid, particularly with an amino acid selected from A, G, K and / or V; or amino acid T at position 1196 of UniProt No. A0JP82 is substituted with another amino acid, particularly with an amino acid selected from A, G, K and / or V.

[0063] 17. The recombinant TET3 enzyme of any one of embodiments 1 -16, wherein amino acid Y at position 1567 of UniProt No. A0A5K1WP6 (SEQ ID NO: 1 ) is substituted with another amino acid, particularly with an amino acid selected from F, M and / or W; or amino acid Y at position 1702 of UniProt No. Q8BG87 is substituted with another amino acid, particularly with an amino acid selected from F, M and / or W; or amino acid Y at position 1694 of UniProt No. 043151 is substituted with another amino acid, particularly with an amino acid selected from F, M and / or W; or amino acid Y at position 1901 of UniProt No. A0JP82 is substituted with another amino acid, particularly with an amino acid selected from F, M and / or W.

[0064] 18. The recombinant TET3 enzyme of any one of embodiments 1 -17, which catalyzes oxidation of 5-methyl-2'-deoxycytidine (mdC) to 5-hydroxymethyl-2'-deoxycytidine (hmdC), 5-formyl-2'-deoxycytidine (fdC) and / or 5-carboxy-2'-deoxycytidine (cadC), and which optionally catalyzes oxidation of 5-methylcytdine (mC) to 5- hydroxymethylcytidine (hmC), 5-formylcytidine (fC), and / or 5-carboxycytidine (caC).

[0065] 19. The recombinant TET3 enzyme of any one of embodiments 1 -18, which catalyzes oxidation of mdC selectively to cadC.

[0066] 20. A nucleic acid molecule encoding the TET3 enzyme of any one of embodiments 1 - 19, optionally linked to an expression control sequence.

[0067] 21. The nucleic acid molecule of embodiment 20 having a codon-optimized sequence for expression in bacteria.

[0068] 22. A vector comprising a nucleic acid molecule of any one of any one of embodiments 20-21.

[0069] 23. A host cell transformed or transfected with a vector of embodiment 22. 24. The host cell of embodiment 23, which is a prokaryotic or eukaryotic cell, preferably a prokaryotic cell.

[0070] 25. The host cell of embodiment 23 or 24, which is an E. coli cell.

[0071] 26. Use of a recombinant TET3 enzyme of any one of embodiments 1 -19 for the sequencing of a nucleic acid substrate, particularly for the sequencing of a nucleic acid substrate comprising mdC and / or mC nucleotides.

[0072] 27. The use of embodiment 26, wherein mdC nucleotides are oxidized to hmdC, fdC and / or cadC, and / or mC nucleotides are oxidized to hmC, fC, and / or caC.

[0073] 28. The use of any one of embodiments 26-27, wherein the sequencing procedure does not involve bisulfite addition.

[0074] 29. The use of any one of embodiments 26-28, wherein the sequencing procedure is a single molecule-sequencing procedure.

[0075] 30. The use of any one of embodiments 26-29 in diagnostics, preferably (early) cancer diagnostics, such as cancer screening.

[0076] 31 . The use of any one of embodiments 26-30 in biological assays, such as oxygenase assays, enzyme kinetic assays, drug screening assays, and / or drug selectivity profiling.

[0077] 32. A method for sequencing a nucleic acid substrate comprising mdC and / or mC nucleotides, wherein the method comprises the steps:

[0078] (A) contacting the nucleic acid substrate with a recombinant TET3 enzyme of any one of embodiments 1 -19 under conditions, where mdC and / or mC nucleotides in the nucleic acid substrate are oxidized, e.g. to hmdC, fdC, cadC, hmC, fC, and / or caC

[0079] (B) optionally isolating the oxidized nucleic acid substrate, and

[0080] (C) determining the amount and / or position of mdC and / or mC nucleotides in the nucleic acid substrate. The method of embodiment 32, wherein the oxidation in step (A) is performed with an oxidizing agent, particularly a hypervalent iodine reagent such as IBX, nitroxyl radical-generating agent such as TEMPO, Mn02, chromate, or any combination thereof. The method of any one of embodiments 32-33, which is a SMRT-sequencing method. The method of any one of embodiments 32-34, wherein the determining step (C) comprises detecting oxidized nucleotides by chromatography, e.g., a HPLC-based chromatography, optionally in combination with MS. The method of any one of embodiments 32-35, wherein the oxidation in step (A) takes place under conditions, wherein mdC and / or mC nucleotides are selectively oxidized to cadC and / or caC nucleotides, particularly in an amount of at least 90%, preferably at least 95%, more preferably at least 99%, even more preferably at least 99.9% based on the total amount of oxidized mdC and / or mC nucleotides. The method of embodiment 36, wherein the oxidation in step (A) is performed at an ion strength corresponding to 50-120 mM sodium chloride, preferably corresponding to 60-100 mM sodium chloride. The method of any one of embodiments 32-35, wherein the oxidation in step (A) takes place under conditions, wherein mdC and / or mC nucleotides are selectively oxidized to 5-hmdC, 5-hmC, 5-fdC and / or 5-fC nucleotides, particularly in an amount of at least 50%, preferably at least 70%, more preferably at least 90% based on the total amount of oxidized mdC and / or mC nucleotides. The method of embodiment 38, wherein the oxidation in step (A) is performed at an ion strength corresponding to 130-160 mM sodium chloride, preferably corresponding to 140-150 mM sodium chloride. The method of any one of embodiments 32-35, wherein the oxidation in step (A) takes place under conditions, wherein mdC and / or mC are selectively oxidized to 5-fdC and / or 5-fC, particularly to at least 50%, preferably at least 70%, more preferably more than 90% based on the total amount of oxidized mdC and / or mC nucleotides.

[0081] 41. The method of embodiment 40, wherein the oxidation in step (A) is performed at an ion strength corresponding to 180-250 mM sodium chloride, preferably corresponding to 190-220 mM sodium chloride, even more preferably corresponding to 200-210 mM.

[0082] 42. A reagent for sequencing a nucleic acid substrate comprising mdC and / or mC nucleotides comprising:

[0083] (A) a recombinant TET3 enzyme of any one of embodiments 1 -19 and

[0084] (B) an oxidizing agent.

[0085] 43. The reagent of embodiment 42, wherein the oxidizing agent is selected from the group consisting of a hypervalent iodine reagent such as IBX, nitroxyl radicalgenerating agent such as TEMPO, MnO2, chromate, or any combination thereof.

[0086] Detailed description

[0087] TET enzymes are alpha-ketoglutarate dependent Fe2+dioxygenases that catalyze dioxygenase-mediated oxidation of mdC to 5-hydroxymethyl-2'-deoxycytidine (hmdC), 5-formyl-2'-deoxycytidine (fdC) and / or 5-carboxy-2' -deoxycytidine (cadC). So far, those oxidations were performed with difficult to overexpress TET1 - and TET2-derived enzymes [6, 17, 21 -24], Further, the known enzymes were not able to quantitatively oxidize the mdC substrate with high selectivity and turnover rates.

[0088] Surprisingly, a recombinant TET3 enzyme has been found, which can be efficiently produced by a cost, time and energy saving process. Moreover, the recombinant TET3 enzyme of the invention surprisingly provides highly selective and quantitative oxidation capability towards a mdC and / or mC containing nucleic acid substrate.

[0089] “TET3 enzyme” in the sense of the present invention refers to any vertebrate TET 3 enzyme, e.g., to a mammalian, amphibious or reptile TET3 enzyme, in particular human (Homo sapiens), mouse (Mus musculus), or frog (Xenopus tropicalis) TET3 enzyme. TET3 plays a key role in epigenetic chromatin reprogramming during embryonic development and is implicated in demethylation in many biological processes such as zygote formation, embryogenesis, axon regeneration, and synaptic transmission.

[0090] In the human adult brain, the most prevalent TET3 enzyme is TET3, having the amino acid sequence of SEQ ID NO: 3 (UniProt No. 043151 ). In mice, four isoforms exist, the longest isoform, having the amino acid sequence of SEQ ID NO: 2 (UniProt No. Q8BG87) and three shorter isoforms, having the amino acid sequence of SEQ ID NO: 1 (UniProt No. A0A5K1 WP6), UniProt No. Q8BG87-2 and UniProt No. Q8BG87-4. In frogs, in particular in Xenopus tropicalis, the TET3 enzyme of the amino acid sequence SEQ ID NO: 4 (UniProt No. A0JP82) is present.

[0091] Thus, a first aspect of the present invention is directed to a TET3 enzyme, comprising: a) a first TET3 domain comprising an amino acid sequence having amino acids 696-1048 or amino acids 697-1048 of UniProt No. A0A5K1WP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 85%, at least 90%, at least 95% or at least 98% over the whole length thereof; b) a second TET3 domain comprising an amino acid sequence having amino acids 1509-1599 of UniProt No. A0A5K1WP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 85%, at least 90%, at least 95% or at least 98% over the whole length thereof, and c) a connecting domain located between domain (a) and domain (b) comprising an amino acid sequence of at least 5 amino acids which is heterologous to a TET3 enzyme.

[0092] The TET3 enzyme of the invention comprises a first TET3 domain (a) and a second TET3 domain (b) linked by a heterologous connecting domain. The first and second TET3 domain comprise the indicated portions of the amino acid sequence of a TET3 enzyme as described above or an amino acid sequence having an identity of at least 85%, at least 90%, at least 95% or at least 98% to the mouse TET3 amino acid sequence UniProt No. A0A5K1VVP6 (SEQ ID NO: 1 ) over the whole length thereof. "Percent (%) amino acid sequence identity" with respect to a peptide or polypeptide sequence is defined as the percentage of amino acid residues in a candidate sequence that are identical with the amino acid residues in the specific peptide or polypeptide sequence, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity. Alignment for purposes of determining percent amino acid sequence identity can be achieved in various ways that are within the skill in the art, for instance using publicly available computer software such as BLAST.

[0093] All UniProt sequences cited in connection with the present invention refer in each case to the UniProt sequences as they have been listed on the filing day of the priority application EP24172176.0 (April 24, 2024).

[0094] In certain embodiments, domain (a) further comprises a C-terminal extension of up to 5 amino acids, particularly an amino acid sequence selected from IQKEK (SEQ ID NO: 5), LQKEK (SEQ ID NO: 6) or any partial sequence thereof. In certain embodiments, domain (b) further comprises a C-terminal extension of up to 5 amino acids, particularly an amino acid sequence selected from AARLG (SEQ ID NO: 7) or any partial sequence thereof.

[0095] In certain embodiments, domain (a) comprises an amino acid sequence having amino acids 696-1048 or amino acids 697-1048 of UniProt No. A0A5K1 WP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof.

[0096] In certain embodiments, domain (a) comprises an amino acid sequence having amino acids 831 -1183 of UniProt No. Q8BG87 (SEQ ID NO: 2) or an amino acid sequence having an identity of at least at least 90%, 95%, 98% or 99% over the whole length thereof,

[0097] In certain embodiments, domain (a) comprises an amino acid sequence having amino acids 824-1175 of UniProt No. 043151 (SEQ ID NO: 3) or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof. In certain embodiments, domain (a) comprises an amino acid sequence having amino acids 953-1304 of UniProt No. A0JP82 (SEQ ID NO: 4) or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof.

[0098] In certain embodiments, domain (b) comprises an amino acid sequence having amino acids 1509-1599 of UniProt No. A0A5K1VVP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof.

[0099] In certain embodiments, domain (b) comprises an amino acid sequence having amino acids 1644-1734 of UniProt No. Q8BG87 (SEQ ID NO: 2) or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof.

[0100] In certain embodiments, domain (b) comprises an amino acid sequence having amino acids 1635-1726 of UniProt No. 043151 (SEQ ID NO: 3) or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof.

[0101] In certain embodiments, domain (b) comprises an amino acid sequence having amino acids 1742-1833 of UniProt No. A0JP82 (SEQ ID NO: 4) or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof.

[0102] The connecting domain connects domain (a) with domain (b). Preferably, the connecting domain has a length of 5-100, more preferably 5-30, even more preferably 10-20, and most preferably about 15 amino acids. In certain embodiments, the connecting domain consists of amino acids selected from G, S, A, T, E, N, D, and / or K.

[0103] Preferably, the connecting domain has the amino acid sequence

[0104] [(Gn)Xm]r, wherein X is in each occurrence independently selected from S, E or D, n is 1 -5, particularly 2-4, m is 0-5, particularly 1 -3, r is 1 -5, particularly 2-4. In a preferred embodiment, the connecting domain has the amino acid sequence of SEQ ID NO: 8: -GGGGSGGGGSGGGGS-. In a further preferred embodiment, the connecting domain has the amino acid sequence of SEQ ID NO: 9: -GGGGSGGGGSGGGGE- or SEQ ID NO: 10: -GGGGSGGGGSGGGGD-.

[0105] The recombinant TET3 enzyme may further comprise a further N- and / or C-terminal peptide domain heterologous to a TET3 enzyme as herein described above. In certain embodiments, the N- and / or C-terminal peptide domain has a length of 5-100 amino acids, preferably 5-50 amino acids, even more preferably 10-30 amino acids. The peptide domain may be an affinity tag for purification, e.g., a Strep-tag, a poly-H istidine- tag (e.g., a Hise-tag), a chitin binding protein, a maltose binding protein, an albuminbinding protein, a cellulose binding protein, a choline-binding domain, FLAG-tag, or glutathione-S-transferase (GST)-tag.

[0106] In particular embodiments, the N- and / or C-terminal peptide domain comprises a Strep-tag, more preferably a Strep-tag II, e.g., MASWSHPQFEK (SEQ ID NO: 11 ).

[0107] The recombinant TET3 enzyme may further comprise a protease cleavage site for removing the N- and / or C-terminal heterologous peptide domain as described above. In a preferred embodiment the protease cleavage site is a TEV protease cleavage site, e.g., SEQ ID NO: 12, which is susceptible to cleavage by a TEV protease (Tobacco Etch Virus nuclear-inclusion-a endopeptidase).

[0108] In another preferred embodiment, the protease cleavage site is a PreScission protease cleavage site, e.g., SEQ ID NO: 13, which is susceptible to cleavage by a PreScission protease, a genetically engineered fusion protein of human rhinovirus 3C protease and glutathione S transferase (GST) or a Hise-tagged HRV 3C protease cleavage site. Preferably, the protease cleavage site is N-terminally fused to domain (a) of the recombinant TET3 enzyme, e.g., to amino acid 696 or amino acid 697 of UniProt No. A0A5K1WP6, to amino acid 831 of UniProt No. Q8BG87, to amino acid 824 of UniProt No. 043151 or to amino acid 953 of UniProt No. A0JP82.

[0109] In certain embodiments, the recombinant TET3 enzyme does not contain amino acid portions of a TET3 enzyme having a length of 20 amino acids or more, of 10 amino acids or more, or of 6 amino acids or more outside domain (a) and domain (b), preferably wherein the recombinant TET3 enzyme does not contain amino acid portions of any vertebrate TET 3 enzyme, including mammalian, amphibious or reptile TET3 enzymes, more preferably wherein the recombinant TET3 enzyme does not contain amino acid portions of a human (Homo sapiens), mouse (Mus musculus), or frog (Xenopus tropicalis) TET3 enzyme.

[0110] In certain embodiments, the recombinant TET3 enzyme does not contain amino acid portions of SEQ ID NO: 1 having a length of 20 amino acids or more, of 10 amino acids or more, particularly of 6 amino acids or more outside amino acids 696-1048 or amino acids 697-1048 and amino acids 1509-1599.

[0111] In certain embodiments, the recombinant TET3 enzyme does not contain amino acid portions of SEQ ID NO: 2 having a length of 20 amino acids or more, of 10 amino acids or more, particularly of 6 amino acids or more outside amino acids 831 -1183 and amino acids 1644-1734.

[0112] In certain embodiments, the recombinant TET3 enzyme does not contain amino acid portions of SEQ ID NO: 3 having a length of 20 amino acids or more, of 10 amino acids or more, or of 6 amino acids or more outside amino acids 824-1175 and amino acids 1635-1726.

[0113] In certain embodiments, the recombinant TET3 enzyme does not contain amino acid portions of SEQ ID NO: 4 having a length of 10 amino acids or more, particularly of 6 amino acids or more outside amino acids 953-1304 and amino acids 1742-1833.

[0114] In particular embodiments, the recombinant TET3 enzyme comprises at least one amino acid substitution compared to a naturally occurring TET3 enzyme that alters enzyme specificity and / or activity.

[0115] In certain embodiments, the amino acid T corresponding to amino acid T at position 940 of UniProt No. A0A5K1WP6, amino acid T at position 1075 of UniProt No. Q8BG87, amino acid T at position 1067 of UniProt No. 043151 , and amino acid T at position 1196 is substituted with another amino acid, particularly with an amino acid selected from A, G, K and / or V, more particularly with A.

[0116] In a further preferred embodiment, the amino acid Y corresponding to amino acid Y at position 1567 of UniProt No. A0A5K1VVP6, amino acid Y at position 1702 of UniProt No. Q8BG87, amino acid Y at position 1694 of UniProt No. 043151 or amino acid Y at position 1901 of UniProt No. A0JP82 is substituted with another amino acid, particularly with F or W.

[0117] In a specific embodiment, the recombinant TET3 enzyme comprises a) a first TET3 domain comprising an amino acid sequence having amino acids 696-1048 or amino acids 697-1048 of SEQ ID NO: 1 or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof; b) a second TET3 domain comprising an amino acid sequence having amino acids 1509-1599 of SEQ ID NO: 1 or an amino acid sequence having an identity of at least at least 90%, 95%, 98% or 99% over the whole length thereof, and c) a connecting domain located between domain (a) and domain (b) comprising an amino acid sequence of at least 5 amino acids which is heterologous to a TET3 enzyme, wherein amino acid T at position 940 is substituted with another amino acid, preferably with an amino acid selected from A, G, K and / or V, more preferably A; and / or amino acid Y at position 1567 is substituted with another amino acid, preferably with an amino acid selected from F.

[0118] In a further specific embodiment, the recombinant TET3 enzyme comprises a) an amino acid sequence having amino acids 831 -1183 of SEQ ID NO: 2 or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof; b) an amino acid sequence having amino acids 1644-1734 of SEQ ID NO: 2 or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof; and c) a connecting domain located between domain (a) and domain (b) comprising an amino acid sequence of at least 5 amino acids which is heterologous to a TET3 enzyme, wherein amino acid T at position 1075 is substituted with another amino acid, preferably with an amino acid selected from A, G, K and / or V, more preferably A; and / or amino acid Y at position 1702 is substituted with another amino acid, preferably with an amino acid selected from F or W.

[0119] In a further specific embodiment, the recombinant TET3 enzyme comprises a) an amino acid sequence having amino acids 824-1175 of SEQ ID NO: 3 or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof; b) an amino acid sequence having amino acids 1635-1726 SEQ ID NO: 3 or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof; and c) a connecting domain located between domain (a) and domain (b) comprising an amino acid sequence of at least 5 amino acids which is heterologous to a TET3 enzyme, wherein amino acid T at position 1067 is substituted with another amino acid, preferably with an amino acid selected from A, G, K and / or V, more preferably A; and / or amino acid Y at position 1694 is substituted with another amino acid, preferably with F or W.

[0120] In a further specific embodiment, the recombinant TET3 enzyme comprises a) an amino acid sequence having amino acids 953-1304 of SEQ ID NO: 4 or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof; b) an amino acid sequence having amino acids 1743-1833 of SEQ ID NO: 4 or an amino acid sequence having an identity of at least 90%, 95%, 98% or 99% over the whole length thereof; and c) a connecting domain located between domain (a) and domain (b) comprising an amino acid sequence of at least 5 amino acids which is heterologous to a TET3 enzyme, wherein amino acid T at position 1196 is substituted with another amino acid, preferably with an amino acid selected from A, G, K and / or V, more preferably; and / or amino acid Y at position 1901 of UniProt No. A0JP82 is substituted with another amino acid, preferably with F or W.

[0121] In a preferred embodiment, the recombinant TET3 enzyme has a calculated molecular weight of at least 45 kDa, preferably at least 50 kDa, more preferably 50-55 kDa, most preferably about 52 kDa.

[0122] The recombinant TET3 enzyme catalyzes the oxidation of 5-methyl-2'-deoxycytidine (mdC). In certain embodiments, mdC is oxidized to 5-hydroxymethyl-2'-deoxycytidine (hmdC), 5-formyl-2'-deoxycytidine (fdC) and / or 5-carboxy-2'-deoxycytidine (cadC). Optionally, the recombinant TET3 enzyme catalyzes oxidation of 5-methylcytdine (mC). In certain embodiments, mC is oxidized to 5-hydroxymethylcytidine (hmC), 5- formylcytidine (fC), and / or 5-carboxycytidine (caC).

[0123] In certain embodiments, the TET3 enzyme catalyzes preferably oxidation of mdC selectively to cadC.

[0124] In a second aspect, the present invention is directed to a nucleic acid molecule encoding the recombinant TET3 enzyme, optionally in operative linkage to an expression control sequence.

[0125] “Expression control sequence” in the sense of the present invention refers to a regulatory sequence, i.e. a segment of a nucleic acid molecule which is capable of increasing or decreasing the expression of specific genes within a host cell, e.g., a prokaryotic or eukaryotic host cell. Suitable expression control sequences are well known in the art. For expression in a prokaryotic host cell such as E. coli, a T7 polymerase, lactose / IPTG, arabinose (PBAD) and / or anhydrotetracycline responsive expression control sequence may be used. Preferably, the nucleic acid molecule has a codon-optimized sequence for expression in bacteria. Codon optimization improves the translation efficiency of a target gene in order to maximize protein expression from the DNA sequence. Preferably, the following codon optimization strategy was used: Choice of preferentially most prevalent codons in E. coli, while avoiding repetitive sequence elements that might impact on vector stability.

[0126] In a third aspect, the present invention is directed to a vector, comprising the nucleic acid molecule of above.

[0127] “Vector” in the sense of the present invention refers to an extrachromosomal vehicle, e.g. a plasmid or virus designed for introduction of a nucleic acid molecule into a host cell encoding the recombinant TET3 enzyme and optionally allowing expression of the nucleic acid molecule. The vector may contain an expression control sequence as described above in operative linkage to the nucleic acid molecule encoding the recombinant TET3 enzyme.

[0128] In a further aspect, the present invention is directed to a host cell, transformed or transfected with the vector of above. The host cell is preferably a prokaryotic or eukaryotic cell, more preferably a prokaryotic cell. In a preferred embodiment, the host cell is an E. coli cell or Bacillus cell, most preferably an E. coli cell. In a further preferred embodiment, the host cell is a yeast cell, e.g. Saccharomyces cerevisiae, or an animal cell, particularly an insect or a mammalian cell, e.g. a human or a hamster cell.

[0129] The recombinant TET3 enzyme may be recombinantly produced in a suitable host cell, e.g. in a prokaryotic or eukaryotic host cell. For this purpose, a vector comprising a nucleic acid molecule encoding the recombinant TET3 enzyme may be introduced into the host cell and the host cell is cultivated under conditions allowing expression of the recombinant TET3 enzyme.

[0130] In a further aspect, the present invention is directed to a use of the recombinant TET3 enzyme for the sequencing of a nucleic acid substrate, e.g., a DNA substrate or an RNA substrate, comprising methylated nucleotides such as mdC and / or mC nucleotides. The nucleic acid substrate may be a single-or double stranded nucleic acid molecule, e.g., a DNA or an RNA molecule. In certain embodiments, the nucleic acid substrate is genomic DNA which may be obtained from a subject, e.g., a human subject.

[0131] In a preferred embodiment, the recombinant TET3 enzyme is incubated with the nucleic acid substrate under conditions wherein mdC nucleotides in the nucleic acid substrate are oxidized to hmdC, fdC and / or cadC, and / or mC nucleotides are oxidized to hmC, fC, and / or caC. Preferably, the total amount of oxidized nucleotides in the substrate is least about 50%, at least about 70% least about 90% or least about 99%.

[0132] In a preferred embodiment, a sequencing procedure using the recombinant TET3 enzyme does not involve bisulfite addition.

[0133] In a preferred embodiment, the nucleotide sequence of the oxidized substrate is determined by single molecule sequencing, e.g., by single molecule real-time (SMRT) sequencing. SMRT sequencing is a parallelized single molecule DNA sequencing method, which may utilize a zero-mode waveguide (ZMW). A single DNA polymerase enzyme molecule is immobilized on a ZMW support and contacted with a single molecule of DNA as a template. The ZMW is a structure that creates an illuminated observation volume that is small enough to observe only a single nucleotide of DNA being incorporated by DNA polymerase. Each of the four DNA bases is attached to one of four different fluorescent dyes. When a nucleotide is incorporated by the DNA polymerase, the fluorescent tag is cleaved off and diffuses out of the observation area of the ZMW where its fluorescence is no longer observable. A detector detects the fluorescent signal of the nucleotide incorporation, and the base call is made according to the corresponding fluorescence of the dye.

[0134] In a particular aspect, the present invention is directed to a method for sequencing a nucleic acid substrate comprising mdC and / or mC nucleotides, wherein the method comprises the steps:

[0135] (A) contacting the nucleic acid substrate with a recombinant TET3 enzyme as described above under conditions, where mdC and / or mC nucleotides in the nucleic acid substrate are oxidized, e.g., oxidized to hmdC, fdC, cadC, hmC, fC, and / or caC,

[0136] (B) optionally isolating the oxidized nucleic acid substrate, and

[0137] (C) determining the location and / or amount of oxidized mdC and / or mC nucleotides in the oxidized nucleic acid substrate.

[0138] Preferably, the nucleic acid substrate comprises a 5-methyl-2'-deoxycytidine (mdC) and / or a 5-methylcytidine (mC) nucleotide or an analogue thereof according to formula (I) wherein

[0139] Ri , and R2 are linkages to adjacent nucleotides; and

[0140] R3 is H or OH.

[0141] In a preferred embodiment, the recombinant TET3 enzyme oxidizes the nucleotide of formula (I) to hmC, hmdC and / or an analogue thereof according to formula (II) wherein

[0142] R1, and R2 are linkages to adjacent nucleotides; and R3 is H or OH. or to fC, fdC and / or an analogue thereof according to formula (III), wherein

[0143] Ri, and F are linkages to adjacent nucleotides; and

[0144] R3 is H or OH. or to caC, cadC or an analogue thereof according to formula (IV) wherein

[0145] R1, and R2 are linkages to adjacent nucleotides; and

[0146] R3 is H or OH. or any combination thereof.

[0147] The oxidation in step (A) is preferably performed in the presence of an oxidizing agent which is capable of oxidizing a nucleotide of formula (I) without substantially disrupting the nucleic substrate. In certain embodiments, the oxidizing agent is selected from a hypervalent iodine reagent such as IBX, a nitroxyl radical generating compound such as 2,2,6,6-Tetramethyl-1 -piperidinyloxy (TEMPO), MnO2, chromate, or any combination thereof.

[0148] In certain embodiments, the oxidation in step (A) may take place under conditions, wherein mdC and / or mC nucleotides are selectively oxidized to cadC and / or caC nucleotides, particularly in an amount of at least 90%, preferably at least 95%, more preferably at least 99%, even more preferably at least 99.9% based on the total amount of oxidized mdC and / or mC nucleotides. Preferably, under these conditions the reaction takes place at an ion strength, i.e., an amount of monovalent cations and anions corresponding to 50-120 mM sodium chloride, preferably corresponding to 60- 100 mM sodium chloride. The ion strength presumably regulates the on and off rates of the enzymes with the oligonucleotide target and for each oxidation a new a- ketoglutarate needs to be loaded into the enzyme.

[0149] Determining cadC nucleotides is particularly preferred in the context of SMRT sequencing by increasing the kinetic parameters IPD (inter-pulse duration) and PW (pulse width) and improving data evaluation. This may open the door for mild epigenetic cadC sequencing.

[0150] In certain embodiments, the oxidation in step (A) may take place under conditions, wherein mdC and / or mC nucleotides are selectively oxidized to 5-hmdC, 5-hmC, 5-fdC and / or 5-fC nucleotides, particularly in an amount of at least 50%, preferably at least 70%, more preferably at least 90% based on the total amount of oxidized mdC and / or mC nucleotides. Preferably, under these conditions the reaction takes place at an ion strength, i.e. an amount of monovalent cations and anions corresponding to 130-160 mM sodium chloride, preferably corresponding to 140-150 mM sodium chloride.

[0151] In certain embodiments, the oxidation in step (A) may take place under conditions, wherein mdC and / or mC are selectively oxidized to 5-fdC and / or 5-fC, particularly to at least 50%, preferably to at least 70%, more preferably to at least 90% based on the total amount of oxidized mdC and / or mC nucleotides. Preferably, under these conditions the reaction takes place at an ion strength, i.e. an amount of monovalent cations and anions corresponding to 180-250 mM sodium chloride, preferably corresponding to 190-220 mM sodium chloride, even more preferably corresponding to 200-210 mM sodium chloride.

[0152] In certain embodiments, the oxidation in step (A) may take place under conditions, wherein the undesired oxidative side products 8-oxodG and hmdll are generated in minor amounts. Thus, preferably, deoxyguanosine (dG) and / or deoxythymidine (dT) are oxidized to 8-oxodG and / or 5-hydroxymethyl-dU (hmdll), respectively, to at most 5%, more preferably to at most 2%, more preferably to at most 1 %, more preferably to at most 0.05%, even more preferably to at most 0.01 %, based on the total amount of oxidized nucleotides.

[0153] In certain embodiments, step (C) comprises determining the amount of oxidized nucleotides in the nucleic acid substrate without sequence analysis, e.g., by cleaving the nucleic acid substrate into individual nucleotides and detecting oxidized nucleotides. Preferably, the detection is performed quantitatively thereby allowing to determine the quantitative amount of methylated nucleotides in the nucleic acid substrate. The detection may be carried out by chromatography, e.g., HPLC-based chromatography, optionally in combination with MS. Preferably, quantitative triple quadruple mass spectrometry (UHPLC-QQQ-MS / MS) is used.

[0154] In certain embodiments, step (C) comprises determining the positions of oxidized nucleotides in the nucleic acid substrate involving a sequencing procedure, e.g., a SMRT sequencing method as described above.

[0155] Sequencing of mdC positions in the genome is of paramount importance for diagnostic applications such as early cancer diagnostics in order to determine incorrect expression of genes.

[0156] Thus, in a preferred embodiment, the recombinant TET3 enzyme is used in diagnostics, preferably (early) cancer diagnostics, such as cancer screening or the detection of the age status of a certain tissue or an organism. Further, the recombinant TET3 enzyme may be used in biological assays, drug screening assays, and / or drug selectivity profiling and / or assays in which the methylation state of a genetic element needs to be determined.

[0157] A further aspect of the present invention is directed to a reagent for sequencing a nucleic acid substrate comprising mdC and / or mC nucleotides comprising:

[0158] (A) a recombinant TET3 enzyme as described above, and

[0159] (B) an oxidizing agent.

[0160] Further, the invention is described in detail by the following figures and examples. Figure Legends

[0161] Figure 1 : Summary of TET3 catalytic domains and constructs.

[0162] The Figure shows the preferred TET3 enzyme constructs according to the invention, based on UniProt No. A0A5K1WP6 (SEQ ID NO: 1 ) (Mus musculus, hereinafter hpMmTET3 ) UniProt No. 043151 (SEQ ID NO: 3) (Homo sapiens, herein after “hpHsTET3”) or UniProt No. A0JP82 (SEQ ID NO: 4) (Xenopus tropicalis, herein after hpXtTET3 )

[0163] Figure 2: Purification of hpTET3 enzymes.

[0164] Figure 2 a): Depiction of the domain structure of recombinant hpMmTET3 enzyme in comparison to the native isoform TET3 (on basis of SEQ ID NO: 1 (UniProt No. A0A5K1 WP6) or SEQ ID NO: 2 (Q8BG87)).

[0165] Figure 2 b): Schematic outline of recombinant TET3 enzymes. Domain structure of TET3 and shortened TET3 version (“hpTET3”). Numbering corresponds to the full- length human TET3. DSBH: double-stranded 0-helix domain, LCI: low-complexity insert; a-KG: a-ketoglutarate.

[0166] Figure 2 c): Conservation of TET3 homologs mapped onto the surface of the predicted structure of the catalytic domain (residues 821 -1721 , without LCI) of human TET3 (Alphafold: AF-O43151 -F1 -model_v4). The LCI is indicated by the dashed line. For residues 1 -825 and 1721 -1795 as well as the LCI no structure was predicted due to likely flexibility and / or disorder of these regions.

[0167] Figure 2 d): SDS-PAGE analysis of the purified hpTET3 proteins. H. sapiens hpHsTET3, Mus musculus hpMmTET3; Xenopus tropicalis hpXtTET3 after two steps..

[0168] Figure 3: Two-step purification of hpMmTET3 enzyme expressed in E. coli.

[0169] N-terminally Strep(ll)-tagged TET3 enzyme (hpMmTET3) was first purified by affinity chromatography on a StrepTrap XT resin. Contaminating DnaK and DnaJ chaperons were removed from hpMmTET3 by cation exchange chromatography using a HiTrap Heparin HP or HiTrap CaptoS column. Representative chromatograms and Coomassie-stained SDS-polyacrylamide gel electrophoresis analysis of the StrepTrap XT affinity chromatography elution fractions are shown in (A) and after Heparin chromatography in (B).

[0170] Figure 4: Purified recombinant hpMmTET3 enzyme.

[0171] Purified hpMmTET3 protein was loaded on a 4-15% gradient SDS-PAGE gel and stained with Coomassie Blue.

[0172] Figure 5: Recombinant TET3 enzyme protein activity.

[0173] The figure shows the influence of various salt concentrations on the distribution of the oxidation products hmdC, fdCd and cadC. At 150 mM NaCI preferentially fdC is formed. At 200 mM NaCI a mixture of hmdC and fdC is generated.

[0174] Figure 6: Oxidation of genomic DNA.

[0175] Figures 6 a) and b): 8-oxo-dG (a) and hmdll (b) levels per nucleosides (dN) of various genomic DNA before and after treatment with hpTET3 as quantified by UHPLC-QQQ- MS. Depicted are biological mean values with the respective ±SD.

[0176] Figures 6 c) and d): Analysis of the oxidation byproducts 8-oxodG and hmdll in synthetic and genomic DNA using the depicted isotope standards for quantification.

[0177] Amount of 8-oxodG per dT of ODN and gDNA (c), and levels of 5hmdU per dG prior (control and gDNA / ODN) (d) and following treatment with hpTET3 or solely with reaction buffer.

[0178] Figure 7: Modification Frequency of LMD-dC-, LMD-mdC- and LMD-cadC-CpGs.

[0179] The bars (x-axis) display the percentage of CpGs with a modification (methylation) frequency of 0 -10% (top plot, <=10%) and 90 - 100% (bottom plot, >= 90%) for the individual data sets, unmodified (LMD_dC), methylated (LMD-mdC) and carboxylated (LMD_cadC) obtained from applying the three trained deep-learning algorithm (y-axis).

[0180] Figure 8: Comparison of bisulfite sequencing and SMRT sequencing.

[0181] 8 a) Depiction of bisulfite sequencing method and 8 b) of the mdC to cadC oxidation chemistry as the basis for the sequencing method described in this publication. 8 c) Depiction of the SMRT sequencing concept with a circular sequencing DNA template bound to the polymerase and fluorescent labelled triphosphates. 8 d) Fluorescent signal during SMRT sequencing recorded by the Sequel He system. Y-axis shows the fluorescent intensity, x-axis the passing time. The kinetics or the nucleotide incorporation by the polymerase can be described by tow time constants. The inter pulse duration (IPD), the time between two light pulses (nucleotide incorporation) and the pulse width (PW) the time the polymerase needs to form the phosphodiester bond. The presence of non-canonical bases increase IPD and PW.

[0182] Figure 9: Oxidation of genomic DNA with hpTET3.

[0183] 9 a) M.Sssl methylated ADNA(dam-,dcm-) was digested to single nucleosides, isotope standards were added and the mixture analyzed by isotope dilution triple quadrupole mass spectrometry giving quantitative data for all nucleosides (left panel). The quantitative analysis of the nucleoside composition of genomic DNA was repeated after treatment of the genomic DNA with hpTET3 (right panel). 9 b) This analysis was performed on various human and mouse genomes with natural methylation levels.

[0184] Figure 10: Application of hpTET3 for SMRT sequencing.

[0185] 10 a) sequencing and model training workflow; dem & dam negative DNA from lambda phage were sequenced using the Sequel He system, HiFi reads were aligned and IPD as well as PW values (features) extracted using ccsmeth; using ccsmeth we trained a 5mdC detecting model, as well as a cadC detecting model. 10 b) mean IPD & PW values (z-score normalized) extracted by ccsmeth across a 21 kmer centred around a given for unmodified lambda DNA (LMD_dC, grey), the methylated lambda DNA (LMD_mdC, red) and the carboxylated lambda DNA (LMD_cadC, blue). 10 c) CpG model training parameters of the 5mdC and cadC model respectively. 10 d) density plot of the called unmodified cytosines (LMD_dC) and modified CpGs (LMD_5mdC or LMD_cadC).

[0186] Figure 11 : Depiction of the oxidation of mdC to hmdC, fdC and cadC.

[0187] While not-mutated hpTET3 generates exclusively and with high yield cadC (>99.9%), double-mutated hpTET3 (T940A and Y1567F) produces a mixture of hmdC and fdC with little cadC.

[0188] Figure 12: Structural comparison of Homo sapiens TET2 with the AlphaFold prediction of Homo sapiens TET3 (UniProt No. 043151, AFDB: AF-O43151 -F1 - v4). a) Superposition of X-ray crystal structure of the catalytic domain of HsTET2 in complex with dsDNA containing 5hmdC (PDB code 5DELI, grey) and the Alphafold2 model of human TET3 (black). The location of the GS-linker that replaces the low complexity insert (LCI) is indicated as dashed line, b) Zoom in the active site: highlighted as stick representation are the active site residues that are lining the active site and coordinating the iron (shown as sphere), the N-oxalylglycine (dark grey), substituting the 2-oxoglutarate cofactor, as well as the 5hmdC (light grey) flipped-out of the dsDNA duplex and positioned within the catalytic centre. Numbering in grey and black corresponds to HsTET2 and HsTET3, respectively. Figure 13: Sequence alignment of hpTET3 homologs and comparison with HsTET2.

[0189] Secondary structure annotation and numbering corresponds to the crystallized construct of HsTET2 (PDB code 5DELI). Highlighted in blue is the GS-linker sequence that replaces the LCI. The orange and green dots mark the residues coordinating the iron and a-KG, respectively. Sequences were aligned using Clustal5 and the alignment was annotated using ESPript 3.0.6.

[0190] Figure 14: Stability of different TET3cd homologs.

[0191] Temperature dependent unfolding profile curves (fluorescence ratio 350 nm / 330 nm) and inflection temperatures (Ti).

[0192] Examples

[0193] 1. Design of recombinant TET3 enzymes

[0194] Functional recombinant TET3 enzymes were designed which still comprise minimum regions necessary for catalysis (A means that this region of the sequence was deleted):

[0195] • Mus musculus TET3: amino acids 696 - 1604 or amino acids 697-1048 of UniProt No. A0A5K1WP6 (SEQ ID NO: 1 ), A1049 - 1508.

[0196] • Mus musculus TET3: amino acids 831 - 1734 of UniProt No. Q8BG87 (SEQ ID NO: 2), A1189-1641.

[0197] • Homo sapiens TET3: amino acids 824 - 1726 of UniProt No. 043151 (SEQ ID NO: 3), A1181-1634.

[0198] • Xenopus tropicalis TET3: amino acids 953 - 1833 of UniProt No. A0JP82 (SEQ ID NO: 4), A1310-1741.

[0199] 2. Cloning, expression and purification of TET3 enzymes in E. coli

[0200] 2.1 hpMmTET3

[0201] The for E. coli expression codon optimized synthetic sequence encoding a N-terminally Strep(ll)-tagged truncated mouse TET3 protein (amino acids 696 - 1604 or amino acids 697-1048 of UniProt No. A0A5K1VVP6 (SEQ ID NO: 1 ) with residues 1049 - 1507 replaced by a 15-residue GS-linker GGGGSGGGGSGGGGS, SEQ ID NO: 8, named hereinafter “hpMmTET3") was designed, ordered from Life Technologies and sub-cloned into pET28a between the Ncol and Xhol restriction sites. In preparation of recombinant protein expression, BL21 (DE3) competent E. coli cells were transformed with Strep(ll)-tagged hpTET3 plasmid and selected on Luria-Bertani (LB) plates containing kanamycin (25 pg / mL final concentration). A single colony was picked and cultured overnight at 37 °C in 50 mL of LB broth with appropriate antibiotic at 37 °C at 180 rpm in an Innova S44i incubator shaker (Eppendorf). For protein expression, bacteria from the overnight small-scale culture were diluted 1000-fold in LB medium with appropriate antibiotic and cells were grown at 37 °C and 200 rpm until an optical density (ODeoo) of 0.3 was reached. Then cultures were cooled down to 16 °C and target protein expression was induced at an ODeoo between 0.5 and 0.6 by addition of isopropyl-[3-D-thiogalactopyranoside to a final concentration of 0.5 mM. Furthermore, cells were grown in 2YT medium. Protein expression was carried out for 18 h. Subsequently, cells were harvested by centrifugation at 8000 rpm for 10 min at 4 °C, rinsed in cold PBS, then pelleted again at 8000 rpm. The pellet was flash-frozen in liquid nitrogen and stored at -80 °C until further use.

[0202] For purification, frozen cell pellets were thawed on ice, resuspended in ice-cold lysis buffer (50 mM HEPES pH 6.8, 500 mM NaCI, 10% glycerol, 0.5 mM TCEP, 10 pM ZnCl2, 0.5 mg / mL lysozyme, EDTA-free protease inhibitor cocktail (Roche), 5 mM ATP and 5 mM MgCl2) and incubated on ice for 30 min. To reduce viscosity from chromosomal DNA and RNA, lysates were treated with Benzonase. Cells were lysed using a high-pressure homogenizer (EmulsiFlex C5) and the cell lysate was cleared by centrifugation at 40,000 g for 45 min at 4 °C. Cleared lysate was loaded on a StrepTrap XT column (Cytiva 5 mL column) for affinity purification. After sample application, the column was washed with 8 column volumes (CV) of wash buffer (50 mM HEPES pH 6.8, 500 mM NaCI, 10% glycerol and 0.5 mM TCEP) and bound protein was eluted with 10 CV buffer containing 50 mM HEPES (pH 6.8), 100 mM NaCI, 10% glycerol, 0.5 mM TCEP and 50 mM biotin. Fractions containing hpTET3 were pooled and concentrated with an Amicon ultra centrifugal filter device (30 000 MWCO, 15 mL). Then, the eluate was diluted 1 -fold with a buffer containing 50 mM HEPES pH 6.8, 100 mM NaCI, 10% glycerol, 10 mM ATP, 10 mM MgCI2and 0.5 mM TCEP and incubated for 30 min on ice. This procedure aims at the removal of contaminating tightly bound molecular chaperone Hsp70 (DnaK) - Hsp40 (DnaJ) - hpTET3 complexes. Subsequently, incorrectly folded hpTET3 aggregates were centrifuged for 30 min at 25,000 g at 4 °C and the supernatant was applied to a HiTrap Heparin HP column (Cytiva, 5 mL) pre-equilibrated in binding buffer (50 mM HEPES pH 6.8, 100 mM NaCI, 10% glycerol (v / v), 0.5 mM TCEP). The column was washed with 20 CV of binding buffer and hpTET3 was eluted with a linear salt gradient ranging from 0.1 to 1.5 M NaCI (0%-100% elution buffer containing 50 mM HEPES pH 6.8, 1.5 M NaCI, 10% glycerol (v / v), 0.5 mM TCEP) over 20 CV. Alternatively, purification of the protein is possible with a CaptoS column without the ATP step. hpTET3-containing fractions were buffer exchanged and concentrated to a final buffer containing 50 mM HEPES pH 6.8, 250 mM NaCI, 10% glycerol and 0.5 mM TCEP using Am icon centrifugal filters with a molecular weight cut-off of 30 kDa. Aliquots of pure hpTET3 were flash-frozen in liquid nitrogen and stored at -80 °C. Protein purity was confirmed by Coomassie-stained SDS-PAGE. Typical yields were 2-3 mg per liter of E. coli culture.

[0203] All column chromatography steps were carried out at 4 °C with pre-cooled buffers on an AKTA pure chromatography system and protein-containing fractions were kept on ice during whole purification procedure.

[0204] The purity of the protein according to analysis by SDS-PAGE is about 99% (Figure 4).

[0205] 2.2 hpHsTET3 and hpXtTET3

[0206] The codon optimized gene sequences (Azenta Life Sciences) encoding for the catalytic domains of TET3 from Homo sapiens (UniProt No. 043151 , aa 824-1726, A1180-1635, hpHsTET3) and Xenopus tropicalis (UniProt No. A0JP82, aa 953-1833, A1309-1742, hpXtTET3), with the low complexity insert deleted and replaced with a GS linker (SEQ ID NOs: 8-10; Figure 12) similar to the previously reported work on human TET2

[0025] , in frame with a N-terminal Strep-tag and Precission protease cleavage site, were cloned into pET28a using the NEB Hifi Assembly kit (New England Biolabs, Cat. No. E5520S). The identity of the sequences was confirmed by Sanger sequencing (Azenta Life Sciences). For protein expression the plasmids were transformed in Escherichia co / / T7 express (New England Biolabs, Cat. No. C2566H).

[0207] For protein expression, an overnight culture of cells containing the respect expression plasmids (LB medium supplemented with 50 pg / ml kanamycin) was diluted 1 :100 in 2YT medium (supplemented with 50 pg / ml kanamycin) and incubated at 37 °C, 180 rpm until an ODeoo of ~0.8 was reached. Cells were cooled down on wet ice for 20 min and protein expression was induced with 0.5 mM isopropyl-[3-D-1 - thiogalactopyranoside (IPTG) and incubated overnight at 18 °C. Cells were subsequently harvested by centrifugation (4000 x g; 20 min, RT) and the cell pellets stored at -20 °C.

[0208] The cell pellet obtained from a 2-liter expression culture was resuspended in ~20 mL lysis buffer (50 mM Hepes, 500 mM NaCI, 10% v / v glycerol, 10 pM ZnCl2 mM, 0.5 mM tris(2-chlorethyl)phosphat (TCEP), pH 6.8) supplemented with DNase I (AppliChem, Cat. No. A3778) and complete™ protease inhibitors (Roche, Cat. No. 1169749800). Cells were lysed by homogenization using an EmulsiFlexC5 (Avestin Inc.) and cleared lysate was obtained by centrifugation (20,000 x g; 20 min; 4 °C). Nucleic acids were removed from the lysate by adding dropwise polyethyleneimine (8% w / v; Merck Cat. No. 408727) to a final concentration of 0.5% w / v, followed by a second centrifugation step (20,000 x g; 20 min; 4 °C).

[0209] The cleared lysate was filtered and loaded onto a 1 ml StrepTrap XT column (IBA BioSciences, Cat. No. 2-5024-001 ) attached to an AktaGo FPLC (Cytiva) at 8 °C, the column washed with wash buffer (50 mM Hepes, 1 .5 M NaCI, 10% v / v glycerol, 0.5 mM TCEP, pH 6.8) and the protein finally eluted with elution buffer (50 mM Hepes, 100 mM NaCI, 0.5 mM TCEP, 10% v / v glycerol, 50 mM biotin, pH 7.2). protein-containing fractions where diluted 1 :10 in low-salt buffer (50 mM Hepes, pH 6.8, 0.5 mM TCEP, 10% v / v glycerol) and loaded onto a 1 mL HiTrap CaptoS column on an AktaGo FPLC (Cytiva) and eluted with a salt gradient (50 mM Hepes pH 6.8, 1.5 M NaCI, 0.5 mM TCEP, 10% v / v glycerol). Fractions containing pure protein were pooled and buffer exchanged to storage buffer (25 mM Hepes pH 6.8, 150 mM NaCI, 10 % glycerol, 2 mM TCEP, pH 6.8), concentrated to 1 mg / mL using centrifugal filter devices (Amicon, Merck) aliquoted, flash-frozen in liquid nitrogen and stored at -80 °C.

[0210] 3. Protein stability

[0211] Protein stability was analyzed using a Tycho™ NT.6 (Nanotemper Technologies) according to the manufactures protocol. Approximately 10 pL of a protein solution (0.3- 0.5 mg / mL in 50 mM Hepes, 150 mM NaCI, 10% v / v glycerol) was applied in capillary. Then the tryptophan (and tyrosine) fluorescence at 330 nm and 350 nm over a temperature gradient (35 °C - 95 °C) was determined. The 350 nm / 330 nm ratio is measure for a spectral shift in the fluorescence and the emission profile from which the inflection temperature can be derived.

[0212] 4. Sequence conservation mapping For mapping the sequence conservation on the surface of the Alphafold model of human TET3 ConSurf3was used with the following settings: Homologs were collected from the UNIREF90 database, with the HMMER search algorithm with an E-value cutoff of 0.0001 and homologues thresholds: hit cutoff is 97% (This is the maximal sequence identity between homologues); maximal number of final homologues =150; Maximal overlap between homologues is 10% (If overlap between two homologues exceeds 10%, the highest scoring homologue is chosen). Coverage is 60% (This is the minimal percentage of the query sequence covered by the homologue). Minimal sequence identity with the query sequence is 50%. The multiple sequence alignment was built using MAFFT and the conservation scores were calculated with the Bayesian method.

[0213] All structural figures were prepared with PyMol 2.4 (Schrodinger LLC).

[0214] Table 1. Sequence identity (id.) and similarity (sim.) of the catalytic domains of TET3 homologs investigated in this study, and the respect TET2 homologs, using EMBOSS Needle Pairwise Sequence Alignment (PSA) (EMBL-EBI https: / / www.ebi.ac.uk / jdispatcher / psa / emboss_needle). N-term and C-term correspond to the sequence preceding and following the low complexity insert (LCI). The overall sequence identity / similarity of the catalytic domains were calculated excluding the LCI., HsTET2 UniProt No. Q6N021 , MmTET2 UniProt No. Q4JK59, XtTET2 UniProt No. F6XM2

[0215] 3. Preparation of methylated lambda DNA Lambda DNA (Oxford Nanopore: EXP-CTL001 ) was amplified by whole genome amplification (WGA) using the Direct WGA kit (Jena Bioscience, cat. no. PCR-382S) in order to produce unmodified DNA. The WGA was conducted according to Direct WGA kit’s protocol - 1.0 ng lambda DNA was incubated with 12 pL reaction buffer, 1 mM dNTP Mix, 1 pL primer mix and 1 pL enzyme mix in a 20 pL reaction volume at 30 °C for overnight (16 h) and followed by a heat inactivation for 5 min at 65 °C.

[0216] Unmodified lambda DNA was sheared by Bioruptor® Pico (Diagenode) (2 pg lambda DNA, 2 times of 2 cycles at 5790”) to produce 1 kb size of fragmented DNA.

[0217] 1 pg of unmethylated, sheared DNA was methylated in vitro using 4 U of M.Sssl enzyme (NEB, cat. no. M0226S) in the presence of 160 pM S-adenosylmethionine (SAM) (NEB, cat. no. B9003S) in the provided reaction buffer (1x) (10 mM Tris-HCI pH 7.9, 50 mM NaCI, 10 mM MgCI2and 1 mM DTT) (NEB, cat. no. B7002S) in a 50 pL reaction volume for 90 min at 37 °C. To ensure complete CpG methylation, the reaction mixture was supplemented with additional 4 U of M.Sssl, 145 pM SAM and the reaction buffer (0.1x) in a 55 pL reaction volume. Then the reaction was incubated for an additional 90 min at 37 °C followed by a heat inactivation for 20 min at 65 °C. Methylated DNA was purified with 1.8x AMPure XP beads (Beckman Coulter, pro. no. A63882) according to the manufacturer’s protocol. DNA methylation was confirmed by quantitative UHPLC-QQQ-MS / MS.

[0218] 3. DNA oxidation with TET3 enzymes

[0219] All reactions were performed with 1 pg genomic DNA in a total volume of 50 pL at 37 °C for 1 h at 500 rpm. DNA concentrations were quantified using 1x dsDNA-HS (High Sensitivity) Qubit Assay-Kit.

[0220] 1 pg of human or mouse genomic DNA was incubated with 4 pM (10 pg) hpTET3 in buffer containing 50 mM HEPES, 50 mM NaCI, 1 mM a-ketoglutarate, 2 mM ascorbic acid, 1.2 mM ATP, 105 pM Fe(NH4)2(SO4)2 and 2.5 mM DTT at 37 °C for 1 h. The pH of the reaction was 7.4 after addition of all buffer components. Alternatively, since 25.7% of total cytosines are methylated (mdC) in M.Sssl treated lambda DNA, compared to approximately 4.5% of total cytosines are mdC in human gDNA, 1 pg of M.Sssl treated lambda DNA was incubated with 11 pM (30 pg) recombinant TET3 at 37 °C for 1 h with the same buffer condition. In a separate experiment, 3 pM synthetic 5mdC-containing 35mer DNA oligonucleotide (5’-CTATACCTCCTCAACTT-5mCGATCACCGTCTCCGGCG-3‘ (SEQ ID NO: 22); 3’-

[0221] GATATGGAGGAGTTGAAGCTAGTGGCAGAGGCCGC-5’ (SEQ ID NO: 23); Sigma- Aldrich) were incubated with 2 pM (5 pg) of the respective enzymes. After that, 0.8 U of Proteinase K (New England Biolabs GmBH, Cat. no. P8107S) and SDS to a final concentration of 0.05% w / v were added to the reaction mixture and incubated for 1 h at 50 °C for digestion of the hpTET3 enzymes, followed by heat-inactivation of Proteinase K (95 °C, 10 min). Subsequently, oxidized gDNA-samples were purified with 1.8x AMPure XP beads (Beckman Coulter, Cat. No. A63882) according to the manufacturer’s protocol and samples containing the synthetic dsDNA were purified using the Monarch® PCR&DNA Clean Up Kit (5 pg) (New England Biolabs GmBH, Cat. No. T1030L), according to the manufacturer’s protocol. In some cases, this was followed by DNA quantification by the QubitTM dsDNA HS Assay (InvitrogenTM, Cat. No. Q32854).

[0222] Oxidation efficiency was quantified by UHPLC-QQQ-MS / MS.

[0223] It is important to note that ascorbic acid and Fe(NH4)2(SO4)2 must be freshly prepared and Fe(ll) should be dissolved in water and added immediately before the reaction starts to minimize oxidation to Fe(lll).

[0224] The results of the oxidation are shown in Table 2 and Table 3.

[0225] Table 2: Oxidation of genomic DNA with hpTET3. Modified nucleosides normalized to dG are given as mean values with the respective ±SD of two independent biological replicates. Untreated gDNA hpTET3 oxidized gDNA

[0226] M.Sssl methylation efficiency of ADNA was calculated as follows: The A genome (48,502 bp) contains 24,182 cytosine / guanine bases and 6,226 CpG sites. A quantitative methylation would consequently result in 6,226 methylated and 17,956 unmethylated cytosines. These values were divided by the amount of total cytosine / guanine (24,182). The theoretically calculated value was compared to the obtained value after UHPLC-QQQ-MS analysis. Here, the absolute amounts of mdC and dC, as quantified by UHPLC-QQQ-MS, were once again divided by the absolute amount of dG. The results are shown in Tables 3a and 3b.

[0227] Table 3: hmdll and 8-oxo-dG per dN of various genomic DNA (Figure 6) before and after treatment with hpTET3 as quantified by UHPLC-QQQ-MS. Table 3a: Untreated gDNA

[0228] Table 3b: hpTET3 oxidized gDNA

[0229] Table 4: Nucleoside abundance of hpTET3-treated dsODN and HEK293T gDNA. The mean data ± SE of three technical replicates are normalized to the dT amount in the sample and is displayed in the power of x103for easier interpretation.

[0230] 4. Cell culture

[0231] All cell lines used were cultivated at 37 °C in water saturated, C02-enriched (5%) atmosphere.

[0232] HEK293T cells (CLS) and human hepatocellular carcinoma HepG2 cells (CLS) were grown in DMEM with high glucose content (Sigma-Aldrich D6546), supplemented with 10% (v / v) fetal bovine serum (FBS) (Life Technologies 10500-064 or PAN Biotech, Cat. No. P04-03590), 1 % (v / v) penicillin-streptomycin (Sigma-Aldrich P0781 ) and optionally 1 % (v / v) L-alanyl-L-glutamine (Sigma-Aldrich G8541 ). Cells were passaged every second day or twice a week at a ratio of 1 : 10 when reaching a confluence of 70- 80%.

[0233] MOLM-1 3 cells were grown in RPM1 1640 (Sigma-Aldrich R0883), containing 10% (v / v) FBS (Invitrogen 10500-064) and 1 % (v / v) L-alanyl-L-glutamine (Sigma-Aldrich G8541 ). The cells were routinely passaged in a ratio of 1 :6 to 1 :10 when a density of 2 x 106cells / mL was reached.

[0234] J1 mESCs were cultivated on 0.2% (w / v) gelatine-coated plates in DMEM (Sigma- Aldrich D6546), supplemented with 10% (v / v) Pansera ES-grade FBS (Pan Biotech), 1 x MEM-nonessential amino acids (NEAA, Sigma-Aldrich M71145), 2 mM L-alanyl-L- glutamine, 1x Penicillin-Streptomycin (Sigma-Aldrich AP078), 0.1 mM [3- mercaptoethanol, 103U / mL mouse recombinant LIF (mLIF, Sigma-Aldrich ESG1107), 1 .5 pM CGP 77675 (Sigma-Aldrich SML0314) and 3 pM CHIR 99021 (Axon Medchem) (a2iL conditions). mESCs were maintained in the naive state in a2iL medium and passaged every 2 - 3 d in a ratio of 1 :4 to 1 :8 when a confluency of 60 - 75% was reached. To shift cells from the hypomethylated naive state to a primed state with increased mdC and hmdC levels, cells were cultured in medium supplemented with FBS and LIF as described above but in the absence of GSK3a / p and Src kinase inhibitor. Cells were primed for 72 h in total before gDNA isolation.

[0235] Cells were tested for Mycoplasma contamination at least every 2 months.

[0236] 5. Isolation of gDNA

[0237] Cells were lysed directly in the plates with RLT buffer (Qiagen) supplemented with 0.01 equiv. of 2-Mercaptoethanol (14.3 mM final concentration), antioxidants 3,5-di-tert- butyl-4-hydroxytoluene (BHT, 200 pM) and deferoxamine mesylate salt (Desferal, 200 pM). To further homogenize the lysate and shear the gDNA, samples were subjected to bead milling using a Qiagen TissueLyser for 30s at 30 Hz. After cell lysis, gDNA isolation was performed as previously described in Traube et al

[0021] , Isolated gDNA was subjected to nucleoside digest and UHPLC-QQQ-MS / MS measurement before and after TET3 oxidation.

[0238] In some cases, genomic DNA isolation was done from pelleted cells using Monarch® gDNA Purification Kit (New England Biolabs GmbH, Cat. No. T301 OS) according to the manufacturer’s protocol.

[0239] 6. DNA digestion

[0240] The purified DNA products were digested to nucleosides with the Nucleoside Digestion Mix from New England Biolabs GmbH (Cat. No. M0649S) in a total volume of 50 pL using 1 pL of enzyme and 5 pL of 10x reaction buffer at 37 °C for 2h. Samples were filtered by using an AcroPrep Advance 96-well Supor filter plate, 0.2 pm (Pall Life Sciences, Cytiva, Cat. No. 8019) and subjected to UHPLC-QQQ-MS / MS. 7. Analysis of conversion rates using UHPLC-QQQ-MS

[0241] Absolute quantification of modified nucleosides was performed with a previously published method

[0021] , For the exact quantification of nucleosides using the stable isotope dilution technique, an Agilent 1290 Infinity II equipped with a variable wavelength detector (VWD) combined with an Agilent Technologies G6490 Triple Quad LC / MS system with electrospray ionization (ESI-MS, Agilent Jetstream) was used. Chromatography was performed by an InfinityLab Poroshell 120 SB-C18 column (2.7 pm, 2.1 x 150 mm; Agilent Technologies, cat. no. 683775-902) at 35 °C and a flowrate of 0.35 mL / min using water supplemented with 0.0075% (vol / vol) formic acid (FA) as solvent A and acetonitrile (MeCN) supplemented with 0.0075% (vol / vol) FA as solvent B. The gradient started at 100% solvent A, followed by an increase to 3.5% solvent B over 4 min (0 min - 4 min). From 4 min to 7 min solvent B was increased to 5% and from 7.0 min to 7.5 min, solvent B was increased further to 80%, maintained at 80% for 2.0 min before returning to 100% solvent A in 0.5 min and a 3.0 min re-equilibration period. In some cases, the autosampler was cooled to 4 °C and the injection volume was set to 10 pL for all samples. The operating parameters were: positive-ion mode, cell accelerator voltage of 5 V, N2 gas temperature of 120 °C, N2 gas flow of 11 L / min, sheath gas temperature of 280 °C with a flow of 11 L / min, capillary voltage of 3000 V in positive ion mode, nozzle voltage set to 0 V or 500 V, nebulizer at 60 psi, high- pressure RF at 150 V and low-pressure RF at 60 V. The instrument was operated in dynamic MRM mode. The fragmentor voltage was 380 V for all compounds, while other compound-dependent parameters are summarized in Table 5 together with the retention times and mass transitions of unlabeled and isotope-labeled nucleosides. MS1 resolution was set to "Wide" and the MS2 resolution to "Unit. Each sample was co-injected with 1 pL of 0.5 pM stable isotope labeled internal standard (ISTD) mix containing the following isotope standards:15Ns-13Cio-dA,13Cg-dC,15Ns-13Cio-dG, [15N2-13Cw]-dT, D3-m5dC,15N2, D2-hm5dC,15N2-f5dC,15N2-ca5dC,15Ns-8oxodG and [D2]-hmdU. The sample data were analyzed by Agilent’s Quantitative MassHunter Software (v B07.01 ) using the built-in calibration function.

[0242] The results are shown in Table 5 and Table 6. Table 5: Compound-dependent LC-MS / MS-parameters. Rt: retention time CE: collision energy; CAV: collision cell accelerator voltage. Table 6: Abundance of modified cytidine nucleosides in HEK293T gDNA used for the evaluation of the performance of the hpTET3 enzymes. From the two biological replicates three technical replicates were analyzed. gDNA: isolated genomic DNA was digested to nucleoside level and the present nucleosides were quantified. Control: nucleoside abundance in gDNA following the incubation in the oxidation buffer in the absence of hpTET3 enzymes.

[0243] 8. Sequencing library preparation

[0244] Libraries for the Sequel He were prepared according to PacBio's SMRTbell express template prep kit 2.0 (PN: 100-938-900). Briefly, for DNA-repair and A-tailing, 300 ng - 1 pg DNA (1 kb) in 46 pL nuclease free water, 8 pL repair buffer, 4 pL end repair mix and 2 pL DNA repair mix were mixed by pipetting and incubated at 30°C for 30 min. The reaction was inactivated by incubating at 65 °C for 5min. Next, 4 pL SMRTbell barcoded adapter (Barcoded overhang adapter kit 8B, PN: 101 -628-500), 30 pL ligation mix and 1 pL ligation enhancer were mixed by pipetting. 31 pL of the ligation mix was directly added to the A-tailed DNA. The reaction was incubated for 30 min at 20 °C and subsequently purified using SMRTbell cleanup beads. Non-SMRTbell-DNA was depleted by nuclease treatment. The concentration of the final library was determined using the Qubit dsDNA-HS-Assay-Kit (ThermoFisher: cat. no. Q32851 ).

[0245] 9. Cohort section

[0246] Three distinct samples were chosen from Lambda virus DNA (dcm-dam-). M.Sssl were used with recombinant TET3 enzyme oxidation to modify all CpGs in two samples. One sample remained unmodified, serving as a control, while the others were intentionally modified.

[0247] 10. Data Provision

[0248] Initiating with HiFi reads from ccsmeth tools, the inventors subsequently aligned these reads to the reference genome. To provide data for model training, ccsmeth extracted features from the sequencing data, extracts a meaningful feature for forward and reverse strand. Post-model training, the inventors utilized the ccsmeth tools, call modification, and call_modification_frequency, for detecting modification frequencies in our three samples.

[0249] 11. Training

[0250] Ccsmeth algorithm was trained for unmodified control and 5mdC-modified samples, as well as, for unmodified control and cadC-modified samples. Random initialization and training on 75% of the datasets were conducted, with 75% of each of samples used for optimizing parameters until loss reached a minimum for inference. Subsequently, performance was assessed on the remaining 25% (testing set). The inventors used a learning rate of 1x1 O’3and stopped training after no improvement in test loss could be seen for 4 epochs, allowing a maximum of 20 epochs.

[0251] 12. Results and Discussion

[0252] The fundament for the invention is newly developed class of recombinant TET3 enzymes from mouse, human and western claw frog, sharing 84-95% sequence identity and 92-96% similarity. They can be overexpressed in prokaryotic or eukaryotic cells, preferably E. coli, and oxidize mdC in the genome to cadC with over 97-99% yield.

[0253] The inventors implemented the new TET-based technology using SMRT sequencing in which a polymerase base pairs a nucleotide within a template with a fluorescently labelled incoming triphosphate. A detector measures the fluorescent signal from the triphosphate bound in the active site of the polymerase base pairing with the templating base before the fluorescence label is cleaved off by the phosphodiester bond formation process. Because the to be sequenced DNA fragment is embedded in a circular sequencing structure (Figure 8 c), sequencing involves the polymerase to move multiple times along the circular DNA structure so that each base (including the cadC base) is read multiple times. This provides a set of data points for each base that allows averaging. SMRT-sequencing detects next to the fluorescence signal, which identifies the incoming base also the time the polymerase needs to form the phosphodiester bond (PW-value) and the time between each incorporation events (IPD-value) so that multiple parameters are measured for the to be sequenced base including cadC, which we form from mdC by TET-induced oxidation.

[0254] A highly shortened (by ~73%; 52 kDa, respectively) mouse recombinant TET3 enzyme (hpMmTET3) consisting of only 465 amino acids, human recombinant TET3 enzyme (hpHsTET3) consisting of only 467 amino acids and western claw frog recombinant TET3 enzyme (hpXtTET3) consisting of only 467 amino acids were designed. In these hpTET3 enzymes, the inventors replaced the low-complexity regions within the catalytic domains (cd) with glycine-serine linkers (Figure 2 a, b). The residual proteins hpHsTET3, hpMmTET3 and hpXtTET3 start at the N-termini with a Cys-rich domain, followed by the catalytically competent double-stranded 0-helix (DSBH) domain. To facilitate purification all proteins were equipped with an N-terminal Strep-tag II (SEQ ID NO: 11 ). The so generated recombinant hpTET3 enzymes could be overexpressed in E. coli. The inventors purified the hpTET3 enzymes (Figure 2 d) first by affinity chromatography over StrepTrap XT material and secondly, a Heparin or CaptoS column was used to remove chaperone contaminations (Hsp40 and Hsp70). The result of the two-step purification protocol was in all three cases proteins with a purity of >95% in yields of =2-3 mg / L of E. coli culture. Proteins with a purity of >99% were available with a yield of «1 mg / L culture. For a sequence alignment of all three hpTET enzymes and comparison with human TET2 see Figures 12+13 and Table 1.

[0255] In order to study the oxidation capabilities, the inventors treated a synthetic 35mer double stranded oligonucleotide containing a single 5mdC within a CpG context (3 pmol) with one of the three hpTET3 constructs (5 pg) for 1 h at 37 °C. The protein was subsequently digested with Proteinase K, the DNA was isolated followed by total digest with a mixture of commercially available enzymes (see supporting information online) to the nucleoside level. The inventors finally used the isotope dilution HPLC-coupled mass spectrometry method described in

[0021] in order to perform exact quantification of mdC, hmdC, fdC and cadC (Figure 9 c). All three hpTET3 enzymes oxidized 5mdC with yields of >98%. The best protein for the conversion of 5mdC to 5cadC were the mouse and the xenopus hpTET3 enzymes, which converted 98.4% and 98.6% of the mdCs to higher oxidized compounds, respectively. Importantly, both enzymes generated 95.8% and 96.2% 5cadC (Figures 9 d, e). In a next step, the catalytic capabilities of the new recombinant TET3 enzymes in genomic DNA were investigated. To this end the inventors first digested human or mouse genomic DNA isolated from HEK293T cells, with a mixture of enzymes to the single nucleosides level in a control experiment. In order to quantify all nucleosides present in the genome, particularly potentially present oxidized nucleosides such as 5- formyl-dC (fdC), 5-hydroxymethyl-dU (hmll) and 8-oxodG, the inventors performed quantitative triple quadruple mass spectrometry (UHPLC-QQQ-MS) with a full set of internal stable isotope labelled standards for dA, dT, dG, dC, mdC, hmdC, fdC and cadC and hmdll as well as 8-oxodG (Figure 9 a, c) using the method described in

[0021] , the contents of which are herein incorporated by reference. It was found that about 4% of all Cs are present at mdCs in this genomic material (Table 6). The inventors detected other oxidized nucleosides at levels of 0.025% (hmdC), 0.1 % (fdC) and 0.01 % (cadC) relative to the total amount dC. These nucleosides and in particular fdC are likely oxidized lesions that form mostly during DNA isolation and handling. Some hmdC has certainly an epigenetic background.

[0256] The inventors then treated in a second experiment the genomic DNA before the digest with hpTET3 (see Experimental section above) and repeated the digestion and quantification experiment using again the full set of isotope standards. This allowed to obtain highly accurate quantitative data. As depicted in Figure 9 a, it was found that upon oxidation, the signal for mdC completely vanishes, while the signal for cadC got very intense. The inventors repeated the study with various genomes (Figure 9 b) and noted that in all cases the mdC (and also 5-hydroxymethyl-dC) signal completely disappeared upon oxidation with hpTET3, to generate in all cases a new and strong cadC signal. In all cases residual mdC was not detected. Instead, only cadC was detected at levels of 4.23% in HEK293T gDNA (see Experimental section above) as expected for a situation in which all the mdC are oxidized to cadC. The quantitative analysis of the oxidation reaction revealed a yield of mdC to cadC oxidation within genomic DNA of 99.96%.

[0257] Importantly, further analysis of the hpTET3-oxidized genomic DNA for the content of hmdll and 8oxodG, which are potential unwanted oxidative side products, did first of all not yield a significant increase for 8oxodG and secondly only a tiny increase for the hmdU signal was detected. This increase was with only 48 hmdU’s formed per whole lambda genome (48502 bp) small. These data demonstrate that the mdC to cadC oxidation with hpTET3 is not only highly efficient, but also very specific (Figure 6).

[0258] The inventors next treated the HEK293T gDNA (1 pg) with 10 pg of the different hpTET3 enzymes (1 h, 37 °C) and again isolated and digested the DNA in a third experiment. The analysis using the isotope dilution method

[0021] showed that with all three hpTET3 enzymes >99.5% of the mdC were converted (Figure 9 e).

[0259] Again, cadC was by far the dominant reaction product, which formed in yields between 97-98%. The best enzyme proved to be hpXtTET3 with cadC formed in 98%, followed by the mouse and human variant with 97.2% and 97.3%, respectively (Figure 9 e), albeit hpXtTET3 appears to be slightly less stable compared to hpMmTET3 and hpHmTET3 (Figure 14).

[0260] Oxidation of DNA, whether it is enzymatically or chemically undertaken, often generates oxidative lesions such as 8-oxodesoxyguanine (8-oxodG) and 5- hydroxymethlyuridine (5hmdll). A frequently overlooked problem is that many of these generated oxidized side products lead to sequencing errors. Particularly 8-oxodG is decoded by most polymerases both as a dG and dT nucleoside, since it can pair with dA in its syn-conformation. In order to investigate and quantify the amount of oxidation byproducts, the inventors performed both with the 35mer double stranded oligonucleotide and with genomic DNA a deep mass spectrometry-based analysis of the oxidation by-products using synthetic isotope standards depicted in Figures 6 c and d. The obtained levels were normalized against dG (hmdll) or dT (8-oxodG). In the experiment it was detected that the levels of 8-oxodG and hmdll increased after treatment of the DNA with the enzymes. For genomic DNA an increase of the 8-oxodG level by a factor of «2.5, from 0.07 8-oxodG / 1000 dT to 0.2 8-oxodG / 1000 dT was observed (Figure 6 c). The hmdU levels increased by a factor of «20-45 from 0.2 hmdU / 1000 dG to between 4-9 hmdU / 1000 dG.

[0261] Similar increases were detected for the 35mer double strand (Figure 6 d). This result is very surprising because the dG-nucleosides are typically those that are oxidized the easiest and considered to be the prime targets for oxidative damage. The 8-oxodG levels, however, were only increased by a factor of about 2.5, while the hmdU levels rise 10-times stronger by a factor of about 20. This indicates that the oxidation is likely to be enzyme driven. It seems that the hpTET3 enzymes once in a while, bind a dT instead of an mdC for oxidation to hmdll. Importantly, the conversion of dT to hmdll, does not change the coding potential and hence does not interfere with sequencing except for the above-mentioned TAPS and TAB-seq methods. Although this oxidation just adds 4-9 hmdU / 1000 dG is a rather minor side fraction, it nevertheless needs to be taken into account. Thus, the higher oxidative efficiency of the hpTET3 enzymes comes with this small, unavoidable caveat. More important, however, is the observation that the 8-oxodG level is only marginally increased.

[0262] For SMRT sequencing, three model genomes were prepared from lambda phage DNA (dam-, dem-). The first genome (LMD-dC) contained no mdC. In the second genome (LMD-mdC) all CpGs sequences were methylated enzymatically with the CpG-specific methyltransferase M.Sssl. For the third genome (LMD-cadC) the LMD-mdC genome was oxidized with hpTET3 to convert all mdCs to cadCs. In order to prove that these manipulations were successful, next the three genomes were digested to the nucleoside level and the nucleoside compositions were analyzed using UHPLC-QQQ- MS. The obtained data (Figure 7) proved the high methylation efficiency of M.Sssl (99.99%) and again the high oxidation efficacy to cadC using hpTET3 (99.96%).

[0263] Following library preparation along the SMRT sequencing protocol, and sequencing on a Sequel He platform, the inventors obtained 385942 reads for the unmodified lambda genome, 433815 reads for the 5mdC-containing genome and 456646 reads for the cadC containing genome. Next, alignment feature extraction was performed (IPD & PW values) using ccsmeth (https: / / github.com / PengNi / ccsmeth) (Figure 10 a). Regarding the IPD values (Figure 10 b) the inventors saw a large difference in the 21 - Kmer between the dC, mdC and cadC situations. mdC and cadC showed, compared to dC a different signal pattern in close vicinity of the xdC position (Kmer-positions 8- 19). While the signal patterns for mdC and cadC are similar, the normalized time values are strongly increased for cadC. Interesting are also the PW-pattern differences (Figure 10 b) between dC, mdC and cadC. Most significant is that we see for cadC a strong time increase at the Kmer-position 18, which is 7 positions away (downstream) from the cadC position. This data analysis shows how complex the footprint differences are between the dC, mdC and cadC situations, particularly outside of CpG dyads. The PW and IPD difference are manifested not only at the nucleotide itself but in addition several nucleotides away up- or downstream.

[0264] Next, it was explored whether the complex but strong kinetic PW und IPD data for cadC could be used to train an Al-based convolutional neuronal network (CNN). Following the training pipeline of ccsmeth, a mdC-model was trained based on the LMD-dC and LMD-5mdC data sets, as well as a cadC-model based on the LMD-dC and LMD-cadC data sets (Figure 10 a). The key parameter obtained from the Al-model are shown in Figure 10 c.

[0265] It was discovered that the cadC kinetic sequencing data in combination with the trained algorithm provides a cadC-model that exceeds the performance of the canonical ccsmeth and 5mC-LMD model in all aspects. For 5mdC the CNN required 181 training rounds and it reaches a model accuracy of 0.945. The model procession reached a value of 0.962 and the recall (number of describable CpG dyads) was finally 0.956. For cadC, the model provides more accurate values. The model needed only 16 training steps to obtain an accuracy of 0.987, a precision of 0.987, and a recall of 0.988 (Figure 10 c)

[0266] In a next step, the new models were tested (Figure 10 d) by applying them to the individual data sets obtained for LMD-dC, LMD-mdC and LMD-cadC. For the detection of mdC the inventors compared their models LMD-mdC and LMD-cadC with the standard ccsmeth-Model. Based on the UHPLC-QQQ-MS measurements above, it was known that the modified genetic material underlying the LMD-mdC and LMD-cadC models contain 99.x% mdC and 99.x cadC in the CpG dyads. In the LMD-mdC model, the methylation frequency of a CpG dyad should consequently be almost 1 , while it should be close to 0 in the non-methylated LMD-dC DNA.

[0267] In case of the mdC model, it was found that only 78.9% of all CpGs of the unmodified LMD_dC (coverage >= 5) display a methylation frequency of 0-10% and not more than 77.7% of the measured CpGs from the LMD_mdC show a methylation frequency of 90-100% (Figure 7). The numbers are even lower in case of the canonical ccsmeth model. Here, only 64.81 % of the LMD-dC-CpGs show a frequency of 0-10% and 69.26% of all LMD-mdC-CpGs are within a range of 90-100% (Figure 7). With the new LMD-cadC model, the inventors detected for the cadC frequency in CpG dyads a sharp signal (blue) close to 100%. More precisely, 94.68% of all CpGs from the LMD-dC sample display a modification frequency of 0-10% and 94.53% of all CpGs from the cadC sample display a modification frequency of 90-100% (Figure 7). These data show that the new model is perfectly able to predict all the cadCs in the CpG dyads. This in turn shows that the oxidation of mdC to cadC (a base which provides strongly different PW und IPD values) with the new hpTET3 enzyme in combination with the Al-derived cadC-model allows highly accurate and sensitive sequencing of mdC. Importantly, for this sequencing method bisulfite treatment is not required.

[0268] In sum, all three recombinant TET3 enzymes enable the highly efficient oxidation of all mdCs with yields >97% in the genome to 5cadCs. As side products small amounts of 8oxodG and hmdll are formed, at levels that can in particular for hmdll not be ignored in the context of TAB-seq and TAPS, which rely on the readout of 5mdC as dll or DHU, respectively. For other sequencing protocols, this side oxidation plays only a minor role (if at all) because the oxidation product hmdU codes like the starting material dT. Important is the finding that the amount of 8-oxodG does not increase significantly. Given that TET enzymes utilize Fe2+and an a-KG to generate a highly reactive Fe(4)=O species for the nucleobase oxidation, the small amounts of oxidation side products with the highly shortened proteins is a surprise. These results, together with the high 5cadC yields, will help facilitate hpTET3-based sequencing. Further, SMRT sequencing of hpTET3 treated DNA showed that cadC provides highly characteristic IPD and PW values at and in close vicinity to the cadC position, which allow the precise localization of cadC in all types of sequence contexts. This strong kinetic footprint allows to train an Al-based CNN model, so that finally SMRT sequencing of mdC via cadC could be achieved with unprecedented accuracy. The strength of the SMRT technology is that it allows long read sequencing, which gives highly accurate data also for repetitive elements genome. This technology may open a new avenue for methylation analysis and early cancer diagnostics. 13. Sequences

[0269] Natural TET3 enzymes

[0270] > Mouse (Mus musculus) TET3 enzyme (short isoform; UniProt No. A0A5K1WP6;

[0271] SEQ ID NO: 1 ):

[0272] MDSGPVYHGDSRQLSTSGAPVNGAREPAGPGLLGAAGPWRVDQKPDWEAASGP

[0273] THAARLEDAHDLVAFSAVAEAVSSYGALSTRLYETFNREMSREAGSNGRGPRPESC

[0274] SEGSEDLDTLQTALALARHGMKPPNCTCDGPECPDFLEWLEGKIKSMAMEGGQGR

[0275] PRLPGALPPSEAGLPAPSTRPPLLSSEVPQVPPLEGLPLSQSALSIAKEKNISLQTAIA

[0276] IEALTQLSSALPQPSHSTSQASCPLPEALSPSAPFRSPQSYLRAPSWPVVPPEEHPS

[0277] FAPDSPAFPPATPRPEFSEAWGTDTPPATPRNSWPVPRPSPDPMAELEQLLGSASD

[0278] YIQSVFKRPEALPTKPKVKVEAPSSSPAPVPSPISQREAPLLSSEPDTHQKAQTALQ

[0279] QHLHHKRNLFLEQAQDASFPTSTEPQAPGWWAPPGSPAPRPPDKPPKEKKKKPPT

[0280] PAGGPVGAEKTTPGIKTSVRKPIQIKKSRSRDMQPLFLPVRQIVLEGLKPQASEGQA

[0281] PLPAQLSVPPPASQGAASQSCATPLTPEPSLALFAPSPSGDSLLPPTQEMRSPSPM

[0282] VALQSGSTGGPLPPADDKLEELIRQFEAEFGDSFGLPGPPSVPIQEPENQSTCLPAP

[0283] ESPFATRSPKKIKIESSGAVTVLSTTCFHSEEGGQEATPTKAENPLTPTLSGFLESPL

[0284] KYLDTPTKSLLDTPAKKAQSEFPTCDCVEQIVEKDEGPYYTHLGSGPTVASIRELME

[0285] DRYGEKGKAIRIEKVIYTGKEGKSSRGCPIAKVWIRRHTLEEKLLCLVRHRAGHHCQ

[0286] NAVIVILILAWEGIPRSLGDTLYQELTDTLRKYGNPTSRRCGLNDDRTCACQGKDPNT

[0287] CGASFSFGCSWSMYFNGCKYARSKTPRKFRLTGDNPKEEEVLRNSFQDLATEVAP

[0288] LYKRLAPQAYQNQVTNEDVAIDCRLGLKEGRPFSGVTACMDFCAHAHKDQHNLYN

[0289] GCTWCTLTKEDNRCVGQIPEDEQLHVLPLYKMASTDEFGSEENQNAKVSSGAIQV

[0290] LTAFPREVRRLPEPAKSCRQRQLEARKAAAEKKKLQKEKLSTPEKIKQEALELAGVT

[0291] TDPGLSLKGGLSQQSLKPSLKVEPQNHFSSFKYSGNAWESYSVLGSCRPSDPYS

[0292] MSSVYSYHSRYAQPGLASVNGFHSKYTLPSFGYYGFPSSNPVFPSQFLGPSAWGH

[0293] GGSGGSFEKKPDLHALHNSLNPAYGGAEFAELPGQAVATDNHHPIPHHQQPAYPGP

[0294] KEYLLPKVPQLHPASRDPSPFAQSSSCYNRSIKQEPIDPLTQAESIPRDSAKMSRTPL

[0295] PEASQNGGPSHLWGQYSGGPSMSPKRTNSVGGNWGVFPPGESPTIVPDKLNSFG

[0296] ASCLTPSHFPESQWGLFTGEGQQSAPHAGARLRGKPWSPCKFGNGTSALTGPSLT

[0297] EKPWGMGTGDFNPALKGGPGFQDKLWNPVKVEEGRIPTPGANPLDKAWQAFGMP

[0298] LSSNEKLFGALKSEEKLWDPFSLEEGTAEEPPSKGWKEEKSGPTVEEDEEELWSD

[0299] SEHNFLDENIGGVAVAPAHCSILIECARRELHATTPLKKPNRCHPTRISLVFYQHKNLN QPNHGLALWEAKMKQLAERARQRQEEAARLGLGQQEAKLYGKKRKWGGAMVAE PQ H KE KKGAI PTRQALAM PTDSAVTVSSYAYTKVTG PYS RWI

[0300] > Mouse (Mus musculus) TET3 enzyme (long isoform; UniProt No. Q8BG87; SEQ ID NO: 2):

[0301] MSQFQVPLAVQPDLSGLYDFPQGQVMVGGFQGPGLPMAGSETQLRGGGDGRKK RKRCGTCDPCRRLENCGSCTSCTNRRTHQICKLRKCEVLKKKAGLLKEVEINAREG

[0302] TGPWAQGATVKTGSELSPVDGPVPGQMDSGPVYHGDSRQLSTSGAPVNGAREPA GPGLLGAAGPWRVDQKPDWEAASGPTHAARLEDAHDLVAFSAVAEAVSSYGALST RLYETFNREMSREAGSNGRGPRPESCSEGSEDLDTLQTALALARHGMKPPNCTCD GPECPDFLEWLEGKIKSMAMEGGQGRPRLPGALPPSEAGLPAPSTRPPLLSSEVP QVPPLEGLPLSQSALSIAKEKNISLQTAIAIEALTQLSSALPQPSHSTSQASCPLPEAL SPSAPFRSPQSYLRAPSWPWPPEEHPSFAPDSPAFPPATPRPEFSEAWGTDTPPA TPRNSWPVPRPSPDPMAELEQLLGSASDYIQSVFKRPEALPTKPKVKVEAPSSSPA PVPSPISQREAPLLSSEPDTHQKAQTALQQHLHHKRNLFLEQAQDASFPTSTEPQA PGWWAPPGSPAPRPPDKPPKEKKKKPPTPAGGPVGAEKTTPGIKTSVRKPIQIKKS RSRDMQPLFLPVRQIVLEGLKPQASEGQAPLPAQLSVPPPASQGAASQSCATPLTP EPSLALFAPSPSGDSLLPPTQEMRSPSPMVALQSGSTGGPLPPADDKLEELIRQFEA EFGDSFGLPGPPSVPIQEPENQSTCLPAPESPFATRSPKKIKIESSGAVTVLSTTCFH SEEGGQEATPTKAENPLTPTLSGFLESPLKYLDTPTKSLLDTPAKKAQSEFPTCDCV EQIVEKDEGPYYTHLGSGPTVASIRELMEDRYGEKGKAIRIEKVIYTGKEGKSSRGC PIAKWVIRRHTLEEKLLCLVRHRAGHHCQNAVIVILILAWEGIPRSLGDTLYQELTDTL RKYGNPTSRRCGLNDDRTCACQGKDPNTCGASFSFGCSWSMYFNGCKYARSKTP RKFRLTGDNPKEEEVLRNSFQDLATEVAPLYKRLAPQAYQNQVTNEDVAIDCRLGLK EGRPFSGVTACMDFCAHAHKDQHNLYNGCTWCTLTKEDNRCVGQIPEDEQLHVL PLYKMASTDEFGSEENQNAKVSSGAIQVLTAFPREVRRLPEPAKSCRQRQLEARKA AAEKKKLQKEKLSTPEKIKQEALELAGVTTDPGLSLKGGLSQQSLKPSLKVEPQNHF SSFKYSGNAWESYSVLGSCRPSDPYSMSSVYSYHSRYAQPGLASVNGFHSKYTL PSFGYYGFPSSNPVFPSQFLGPSAWGHGGSGGSFEKKPDLHALHNSLNPAYGGAE FAELPGQAVATDNHHPIPHHQQPAYPGPKEYLLPKVPQLHPASRDPSPFAQSSSCY NRSIKQEPIDPLTQAESIPRDSAKMSRTPLPEASQNGGPSHLWGQYSGGPSMSPKR

[0303] TNSVGGNWGVFPPGESPTIVPDKLNSFGASCLTPSHFPESQWGLFTGEGQQSAPH

[0304] AGARLRGKPWSPCKFGNGTSALTGPSLTEKPWGMGTGDFNPALKGGPGFQDKLW NPVKVEEGRIPTPGANPLDKAWQAFGMPLSSNEKLFGALKSEEKLWDPFSLEEGTA E E P P S KG VVKE E KS G PTVE E DEEELWSDSEHNFLDENIGGVAVAPAHCSILIECARR

[0305] ELHATTPLKKPNRCHPTRISLVFYQHKNLNQPNHGLALWEAKMKQLAERARQRQEE

[0306] AARLGLGQQEAKLYGKKRKWGGAMVAEPQHKEKKGAIPTRQALAMPTDSAVTVSS YAYTKVTGPYSRWI

[0307] > Human (Homo sapiens) TET3 enzyme (UniProt No. 043151 ; SEQ ID NO: 3):

[0308] MSQFQVPLAVQPDLPGLYDFPQRQVMVGSFPGSGLSMAGSESQLRGGGDGRKKR

[0309] KRCGTCEPCRRLENCGACTSCTNRRTHQICKLRKCEVLKKKVGLLKEVEIKAGEGA

[0310] GPWGQGAAVKTGSELSPVDGPVPGQMDSGPVYHGDSRQLSASGVPVNGAREPA

[0311] GPSLLGTGGPWRVDQKPDWEAAPGPAHTARLEDAHDLVAFSAVAEAVSSYGALST

[0312] RLYETFNREMSREAGNNSRGPRPGPEGCSAGSEDLDTLQTALALARHGMKPPNCN

[0313] CDGPECPDYLEWLEGKIKSWMEGGEERPRLPGPLPPGEAGLPAPSTRPLLSSEVP

[0314] QISPQEGLPLSQSALSIAKEKNISLQTAIAIEALTQLSSALPQPSHSTPQASCPLPEAL

[0315] SPPAPFRSPQSYLRAPSWPVVPPEEHSSFAPDSSAFPPATPRTEFPEAWGTDTPPA

[0316] TPRSSWPMPRPSPDPMAELEQLLGSASDYIQSVFKRPEALPTKPKVKVEAPSSSPA

[0317] PAPSPVLQREAPTPSSEPDTHQKAQTALQQHLHHKRSLFLEQVHDTSFPAPSEPSA

[0318] PGWWPPPSSPVPRLPDRPPKEKKKKLPTPAGGPVGTEKAAPGIKPSVRKPIQIKKS

[0319] RPREAQPLFPPVRQIVLEGLRSPASQEVQAHPPAPLPASQGSAVPLPPEPSLALFAP

[0320] SPSRDSLLPPTQEMRSPSPMTALQPGSTGPLPPADDKLEELIRQFEAEFGDSFGLP

[0321] GPPSVPIQDPENQQTCLPAPESPFATRSPKQIKIESSGAVTVLSTTCFHSEEGGQEA

[0322] TPTKAENPLTPTLSGFLESPLKYLDTPTKSLLDTPAKRAQAEFPTCDCVEQIVEKDEG

[0323] PYYTHLGSGPTVASIRELMEERYGEKGKAIRIEKVIYTGKEGKSSRGCPIAKWVIRRH

[0324] TLEEKLLCLVRHRAGHHCQNAVIVILILAWEGIPRSLGDTLYQELTDTLRKYGNPTSR

[0325] RCGLNDDRTCACQGKDPNTCGASFSFGCSWSMYFNGCKYARSKTPRKFRLAGDN

[0326] PKEEEVLRKSFQDLATEVAPLYKRLAPQAYQNQVTNEEIAIDCRLGLKEGRPFAGVT

[0327] ACMDFCAHAHKDQHNLYNGCTWCTLTKEDNRCVGKIPEDEQLHVLPLYKMANTDE

[0328] FGSEENQNAKVGSGAIQVLTAFPREVRRLPEPAKSCRQRQLEARKAAAEKKKIQKE

[0329] KLSTPEKIKQEALELAGITSDPGLSLKGGLSQQGLKPSLKVEPQNHFSSFKYSGNAV

[0330] VESYSVLGNCRPSDPYSMNSVYSYHSYYAQPSLTSVNGFHSKYALPSFSYYGFPSS

[0331] NPVFPSQFLGPGAWGHSGSSGSFEKKPDLHALHNSLSPAYGGAEFAELPSQAVPT

[0332] DAHHPTPHHQQPAYPGPKEYLLPKAPLLHSVSRDPSPFAQSSNCYNRSIKQEPVDP

[0333] LTQAEPVPRDAGKMGKTPLSEVSQNGGPSHLWGQYSGGPSMSPKRTNGVGGSW

[0334] GVFSSGESPAIVPDKLSSFGASCLAPSHFTDGQWGLFPGEGQQAASHSGGRLRGK

[0335] PWSPCKFGNSTSALAGPSLTEKPWALGAGDFNSALKGSPGFQDKLWNPMKGEEG RIPAAGASQLDRAWQSFGLPLGSSEKLFGALKSEEKLWDPFSLEEGPAEEPPSKGA

[0336] VKEEKGGGGAEEEEEELWSDSEHNFLDENIGGVAVAPAHGSILIECARRELHATTPL

[0337] KKPNRCHPTRISLVFYQHKNLNQPNHGLALWEAKMKQLAERARARQEEAARLGLG

[0338] QQEAKLYGKKRKWGGTVVAEPQQKEKKGVVPTRQALAVPTDSAVTVSSYAYTKVT GPYSRWI

[0339] > Frog (Xenopus tropicalis) TET3 enzyme (UniProt No. A0JP82; SEQ ID NO: 4):

[0340] MDTQPAPVPHVLPQDVYEFPDDQESLGRLRVSEMPAELNGGGGGGSAAAFAMELP

[0341] EQSNKKRKRCGVCVPCLRKEPCGACYNCVNRSTSHQICKMRKCEQLKKKRVVPM

[0342] KGVENCSESILVDGPKTDQMEAGPVNHVQEGRLKQECDSTLPSKGCEDLANQLLM

[0343] EANSWLSNTAAPQDPCNKLNWDKPTIPNHAANNNSNLEDAKNLVAFSAVAEAMSTY

[0344] GMPASGTPSSVSLQLYEKFNYETNRDNSGHLEGNAPSCPEDLNTLKAALALAKHGV

[0345] KPPNCNCDGPECPDYLEWLENKIKSTVKGSQESPFPNLGQVSKELVQKQYPKEQV

[0346] LNLENKNSTCPSGNLPFSQNALSLAKEKNISLQTAIAIEALTQLSSALPQTNNECPNA

[0347] PSQPLINPHDQLTHFPSAKGNQLPMLPVARNELFQNQQSQLYTGKNALPVPQSPRQ

[0348] TSWEQNKKSSYQEGQYIPENLSHSSSVLPSDASTPQKPEFLQQWVQNADLLKSPS

[0349] DPMTGLKQLLGNTDEYIKSVFKGPEALPNKKNVKPKHTIKSIKKESTEFLKMSPDQQ

[0350] LSQLLQTNEFHRNTQAALQQHLHHKRNLFVDPNAMEACTQEQQNWWVPSSQQAP

[0351] VSKTTEKPVKERKKRRQSPSQKQVEPKPKPQRKQVQIKKPKVKEGSAVFMPVSQIS

[0352] LDTFRRVEKEENQGKEMDAENSLPNNVQTELLESQSLQLTGSQANPDDRKTVNTQ

[0353] EMCNENQSNIGKANNFALCVNRANSFVAKDQCPTPSTHDTSSSSGQGDSANQHTN

[0354] VSDVPGQNDLSCLDDKLEDLIRQFEAEFGEDFSLPGSAVPSQNGEGPPKQTPSGD

[0355] PQFKLPFPSQLLPPENSTKPATHSNPALSNNPVSREVSNNLDSLFSSKSPKQIKIESS

[0356] GAITWSTTCFYSEENQHLDGTPTKSDLPFNPTLSGFLDSPLKYLTSPTKSLIDTPAK

[0357] KAQAEFPTCDCVEQINEKDEGPYYTHLGSGPTVASIRELMEERFGQKGDAIRIEKVIY

[0358] TGKEGKSSRGCPIAKWVIRRQSEDEKLMCLVRQRAGHHCENAVIIILIMAWEGIPRSL

[0359] GDSLYNDITETITKYGNPTSRRCGLNDDRTCACQGKDPNTCGASFSFGCSWSMYF

[0360] NGCKYARSKTPRKFRLIGENPKEEDGLKDNFQNLATKVAPVYKMLAPQAYQNQVNN

[0361] EDIAIDCRLGLKEGRPFSGVTACMDFCAHAHKDQHNLYNGCTWCTLTKEDNRMIG

[0362] RVAEDEQLHVLPLYKVSTTDEFGSEEGQLEKIKKGGIHVLSSFPREVRKLSEPAKSC

[0363] RQRQLEAKKAAAEKKKLQKEKLVSPDKTKQEPSDKKTCQQNPGVPQQQTKPCIKV

[0364] EPSNHYNNFKYNGNGWESYSVLGSCRPSDPYSMNSVYSYHSFYAQPNLPSVNGF

[0365] HSKYALPPFGFYGFPNNPWPNQFMNYGTSDARNSGWMNNCFEKKPELQSLADG

[0366] MNQSYGSELSEQSFRRSSEVPHHYSLQNPSSQKSVNVPHRTTPAPVETTPYSNLP CYNKVIKKEPGSDPLVDSFQRANSVHSHSPGVNHSLQASDLPISYKANGALSSSGR TNAESPCSMFMPNDKNGLEKKDYFGVHSNAPGLKDKQWPPYGTDVSVRQHDSLD SQSPGKVWSSCKLSDSSAALPSSASTQDKNWNGRQVSLNQGMKESALFQEKLWN SVAASDRCSATPSDRSSITPCSELQDKNWGSFPNPTVNSLKTDSSQNHWDPYSLD

[0367] DNMDDGQSKSVKEED DEEIWSDSEHNFLDENIGGVAVAPGHGSILIECARRELHATT PLKKPNRCHPTRISLVFYQHKNLNQPNHGLALWEAKMKQLAERARAREEEAAKLG I KQEVKSLGKKRKWGGAATTETPPVEKKDYTPTRQAATILTDSATTSFSYAYTKVTGP YSRFI

[0368] Extension sequences

[0369] >C-terminal extension sequence of domain (a) (SEQ ID NO: 5) IQKEK

[0370] >C-terminal extension sequence of domain (a) (SEQ ID NO: 6)

[0371] LQKEK

[0372] >C-terminal extension sequence of domain (b) (SEQ ID NO: 7) AARLG

[0373] Synthetic sequences connecting domain (SEQ ID NO: 8)

[0374] GGGGSGGGGSGGGGS connecting domain (SEQ ID NO: 9)

[0375] GGGGSGGGGSGGGGE connecting domain (SEQ ID NO: 10)

[0376] GGGGSGGGGSGGGGD

[0377] >Strep-tag II (SEQ ID NO: 11 )

[0378] MASWSHPQFEK >TEV protease cleavage site (SEQ ID NO: 12)

[0379] SGGGGGENLYFQG

[0380] >PreScission protease cleavage site (SEQ ID NO: 13)

[0381] SGGGGGGALEVLFQGP

[0382] Recombinant hpTET3 enzyme domains

[0383] >domain (a) of Mus musculus, amino acids 696-1048 of UniProt No. A0A5K1WP6

[0384] (SEQ ID NO: 14)

[0385] SEFPTCDCVEQIVEKDEGPYYTHLGSGPTVASIRELMEDRYGEKGKAIRIEKVIYTGK

[0386] EGKSSRGCPIAKWVIRRHTLEEKLLCLVRHRAGHHCQNAVIVILILAWEGIPRSLGDT

[0387] LYQELTDTLRKYGNPTSRRCGLNDDRTCACQGKDPNTCGASFSFGCSWSMYFNG

[0388] CKYARSKTPRKFRLTGDNPKEEEVLRNSFQDLATEVAPLYKRLAPQAYQNQVTNED VAIDCRLGLKEGRPFSGVTACMDFCAHAHKDQHNLYNGCTWCTLTKEDNRCVGQI PEDEQLHVLPLYKMASTDEFGSEENQNAKVSSGAIQVLTAFPREVRRLPEPAKSCR QRQLEARKAAAEKKK

[0389] >domain (a) of Mus musculus, amino acids 831-1183 of UniProt No. Q8BG87 (SEQ ID

[0390] NO: 15)

[0391] SEFPTCDCVEQIVEKDEGPYYTHLGSGPTVASIRELMEDRYGEKGKAIRIEKVIYTGK

[0392] EGKSSRGCPIAKWVIRRHTLEEKLLCLVRHRAGHHCQNAVIVILILAWEGIPRSLGDT

[0393] LYQELTDTLRKYGNPTSRRCGLNDDRTCACQGKDPNTCGASFSFGCSWSMYFNG

[0394] CKYARSKTPRKFRLTGDNPKEEEVLRNSFQDLATEVAPLYKRLAPQAYQNQVTNED VAIDCRLGLKEGRPFSGVTACMDFCAHAHKDQHNLYNGCTWCTLTKEDNRCVGQI PEDEQLHVLPLYKMASTDEFGSEENQNAKVSSGAIQVLTAFPREVRRLPEPAKSCR

[0395] QRQLEARKAAAEKKK

[0396] >domain (a) of Homo sapiens, amino acids 824-1175 of UniProt No. 043151 (SEQ ID

[0397] NO: 16)

[0398] EFPTCDCVEQIVEKDEGPYYTHLGSGPTVASIRELMEERYGEKGKAIRIEKVIYTGKE

[0399] GKSSRGCPIAKWVIRRHTLEEKLLCLVRHRAGHHCQNAVIVILILAWEGIPRSLGDTL

[0400] YQELTDTLRKYGNPTSRRCGLNDDRTCACQGKDPNTCGASFSFGCSWSMYFNGC KYARSKTPRKFRLAGDNPKEEEVLRKSFQDLATEVAPLYKRLAPQAYQNQVTNEEIA IDCRLGLKEGRPFAGVAACMDFCAHAHKDQHNLYNGCTWCTLTKEDNRCVGKIPE DEQLHVLPLYKMANTDEFGSEENQNAKVGSGAIQVLTAFPREVRRLPEPA KSCRQRQLEARKAAAEKKK

[0401] >domain (a) of Xenopus tropicalis, amino acids 953-1304 of UniProt No. A0JP82 (SEQ

[0402] ID NO: 17)

[0403] EFPTCDCVEQINEKDEGPYYTHLGSGPTVASIRELMEERFGQKGDAIRIEKVIYTGKE GKSSRGCPIAKWVIRRQSEDEKLMCLVRQRAGHHCENAVIIILIMAWEGIPRSLGDSL YNDITETITKYGNPTSRRCGLNDDRTCACQGKDPNTCGASFSFGCSWSMYFNGCK YARSKTPRKFRLIGENPKEEDGLKDNFQNLATKVAPVYKMLAPQAYQNQVNNEDIAI

[0404] DCRLGLKEGRPFSGVTACMDFCAHAHKDQHNLYNGCTWCTLTKEDNRMIGRVAE DEQLHVLPLYKVSTTDEFGSEEGQLEKIKKGGIHVLSSFPREVRKLSEPAKSCRQRQ LEAKKAAAEKKK

[0405] >domain (b) of Mus musculus, amino acids 1509-1599 UniProt No. A0A5K1WP6 (SEQ ID NO: 18)

[0406] EELWSDSEHNFLDENIGGVAVAPAHCSILIECARRELHATTPLKKPNRCHPTRISLVFY QHKNLNQPNHGLALWEAKMKQLAERARQRQEE

[0407] >domain (b) of Mus musculus, amino acids 1644-1734 of UniProt No. Q8BG87 (SEQ ID NO: 19)

[0408] EELWSDSEHNFLDENIGGVAVAPAHCSILIECARRELHATTPLKKPNRCHPTRISLVFY QHKNLNQPNHGLALWEAKMKQLAERARQRQEE

[0409] >domain (b) of Homo sapiens, amino acids 1635-1726 of UniProt No. 043151 (SEQ ID NO: 20)

[0410] EEELWSDSEHNFLDENIGGVAVAPAHGSILIECARRELHATTPLKKPNRCHPTRISLVF YQHKNLNQPNHGLALWEAKMKQLAERARARQEE

[0411] >domain (b) of Xenopus tropicalis, amino acids 1742-1833 of UniProt No. A0JP82 (SEQ ID NO: 21 )

[0412] DEEIWSDSEHNFLDENIGGVAVAPGHGSILIECARRELHATTPLKKPNRCHPTRISLV FYQ HKNLNQPNHG LALWE AKM KQ LAE RARAR E E E

[0413] >synthetic 5mdC-containing 35mer DNA oligonucleotides 5’-CTATACCTCCTCAACTT-5mCGATCACCGTCTCCGGCG-3‘ (SEQ ID NO: 22)

[0414] 3’-GATATGGAGGAGTTGAAGCTAGTGGCAGAGGCCGC-5’ (SEQ ID NO: 23)

[0415] References

[0416]

[0001] P. A. Jones, Nat. Rev. Genet. 2012, 13, 484-492.

[0417] [2] A. Papanicolau-Sengos, K. Aidape, Annu. Rev. Pathol.: Meeh. Dis. 2022, 17, 295-321.

[0418] [3] S. Li, T. 0. Tollefsbol, Methods (Amsterdam, Neth.) 2021, 187, 28-43.

[0419] [4] Y. M. D. Lo, D. S. C. Han, P. Jiang, R. W. K. Chiu, Science 2021 , 372, eaaw3616.

[0420] [5] C. Liu, X. Cui, B. S. Zhao, P. Narkhede, Y. Gao, J. Liu, X. Dou, Q. Dai, L. S. Zhang, C. He. J. Am. Chem. Soc. 2020, 142, 4539-4543.

[0421] [6] R. Vaisvila, V. K. C. Ponnaluri, Z. Sun, B. W. Langhorst, L. Saleh, S. Guan, N. Dai, M. A. Campbell, B. S. Sexton, K. Marks, M. Samaranayake, J. C. Samuelson, H. E. Church, E. Tamanaha, I. R. Correa, Jr., S. Pradhan, E. T. Dimalanta, T. C. Evans, Jr., L. Williams, T. B. Davis, Genome Res. 2021 , 31, 1280-1289.

[0422] [7] B. Searle, M. Muller, T. Carell, A. Kellett, Angew. Chem. Int. Ed. 2023, 62, e202215704.

[0423] [8] J. Bonet, M. Chen, M. Dabad, S. Heath, A. Gonzalez-Perez, N. Lopez-Bigas, J. Lagergren, Bioinformatics 2021 , 38, 1235-1243.

[0424] [9] O. Y. O. Tse, P. Jiang, S. H. Cheng, W. Peng, H. Shang, J. Wong, S. L. Chan, L. C. Y. Poon, T. Y. Leung, K. C. A. Chan, R. W. K. Chiu, Y. M. D. Lo, Proc. Natl. Acad. Sci. U. S. A. 2021 , 118, e2019768118.

[0425]

[0010] T. A. Clark, X. Lu, K. Luong, Q. Dai, M. Boitano, S. W. Turner, C. He, J. Korlach, BMC Biol. 2013, 11, 4.

[0426]

[0011] M. J. Booth, T. W. B. Ost, D. Beraldi, N. M. Bell, M. R. Branco, W. Reik and S. Balasubramanian, Nat. Protocols, 2013, 8, 1841 -1851.

[0427]

[0012] H. Sahin, R. Salehi, S. Islam, M. Muller, P. Giehr and T. Carell, Angew. Chem. Int. Ed., 2024, e202418500.

[0428]

[0013] R. Vaisvila, V. K. C. Ponnaluri, Z. Sun, B. W. Langhorst, L. Saleh, S. Guan, N. Dai, M. A. Campbell, B. S. Sexton, K. Marks, M. Samaranayake, J. C. Samuelson, H. E. Church, E. Tamanaha, I. R. Correa, S. Pradhan, E. T. Dimalanta, T. C. Evans, L. Williams and T. B. Davis, Genome Res., 2021 , 31, 1280-1289.

[0429]

[0014] M. Frommer, L. E. McDonald, D. S. Millar, C. M. Collis, F. Watt, G. W. Grigg, P. L. Molloy and C. L. Paul, Proc. Natl. Acad. Sci. U. S. A., 1992, 89, 1827-1831.

[0430]

[0015] T. Wang, J. M. Fowler, L. Liu, C. E. Loo, M. Luo, E. K. Schutsky, K. N. Berrios, J. E. DeNizio, A. Dvorak, N. Downey, S. Montermoso, B. Y. Pingul, M. Nasrallah, W. S. Gosal, H. Wu and R. M. Kohli, Nat. Chem. Biol., 2023, 19, 1004-1012.

[0431]

[0016] Y. Liu, P. Siejka-Zielihska, G. Velikova, Y. Bi, F. Yuan, M. Tomkova, C. Bai, L. Chen, B. Schuster-Bockler and C.-X. Song, Nat. Biotechn., 2019, 37, 424-429.

[0432]

[0017] M. Yu, Gary C. Hon, Keith E. Szulwach, C.-X. Song, L. Zhang, A. Kim, X. Li, Q. Dai, Y. Shen, B. Park, J.-H. Min, P. Jin, B. Ren and C. He, Cell, 2012, 149, 1368-1380.

[0433]

[0018] Y. Liu, Z. Hu, J. Cheng, P. Siejka-Zielihska, J. Chen, M. Inoue, A. A. Ahmed and C.-X. Song, Nat. Comm., 2021, 12, 618.

[0434]

[0019] E. K. Schutsky, J. E. DeNizio, P. Hu, M. Y. Liu, C. S. Nabel, E. B. Fabyanic, Y. Hwang, F. D. Bushman, H. Wu and R. M. Kohli, Nat. Biotechn., 2018, 36, 1083- 1090.

[0435]

[0020] M. A. Carpenter, M. Li, A. Rathore, L. Lackey, E. K. Law, A. M. Land, B. Leonard, S. M. Shandilya, M. F. Bohn, C. A. Schiffer, W. L. Brown and R. S. Harris, J. Biol. Chem., 2012, 287, 34801 -34808.

[0436]

[0021] F. R. Traube, S. Schiffers, K. Iwan, S. Kellner, F. Spada, M. Muller, T. Carell, Nat. Protoc. 2019, 14, 283-312.

[0437]

[0022] Y. Liu, P. Siejka-Zielinska, G. Velikova, Y. Bi, F. Yuan, M. Tomkova, C. Bai, L. Chen, B. Schuster-Bockler, C. X. Song, Nat. Biotechnol. 2019, 37, 424-429.

[0438]

[0023] M. Yu, D. Han, G. C. Hon, C. He, Methods Mol. Biol. 2018, 1708, 645-663.

[0439]

[0024] M. Yu, G. C. Hon, K. E. Szulwach, C. X. Song, P. Jin, B. Ren, C. He, Nat. Protoc.

[0440] 2012, 7, 2159-2170.

[0441]

[0025] L. Hu, J. Lu, J. Cheng, Q. Rao, Z. Li, H. Hou, Z. Lou, L. Zhang, W Li, W Gong,

[0442] M. Liu, C. Sun, X. Yin, J. Li, X. Tan, P. Wang, Y. Wang, D. Fang, Q. Cui, P. Yang,

[0443] C. He, H. Jiang, C. Luo and Y. Xu, Nature, 2015, 527, 118-122.

Claims

Claims1 . A recombinant Ten Eleven Translocation 3 (TET3) enzyme catalyzing the oxidation of 5-methyl-2'-deoxycytidine (mdC) and / or 5-methylcytidine (mC), comprising: a) a first TET3 domain comprising an amino acid sequence having amino acids 697-1048 of UniProt No. A0A5K1WP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 85% over the whole length thereof; b) a second TET3 domain comprising an amino acid sequence having amino acids 1509-1599 of UniProt No. A0A5K1WP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 85% over the whole length thereof, and c) a connecting domain located between domain (a) and domain (b) comprising an amino acid sequence of at least 5 amino acids which is heterologous to a TET3 enzyme, wherein the recombinant TET3 enzyme does not contain amino acid portions of a TET3 enzyme having a length of 20 amino acids or more outside domain (a) and domain (b).

2. The recombinant TET3 enzyme of claim 1 , wherein domain (a) further comprises C-terminal extension of up to 5 amino acids, particularly with an amino acid sequence selected from IQKEK (SEQ ID NO: 5), LQKEK (SEQ ID NO. 6) or any partial sequence thereof, and / or wherein domain (b) further comprises C-terminal extension of up to 5 amino acids, particularly with an amino acid sequence selected from AARLG (SEQ ID NO: 7) or any partial sequence thereof.

3. The recombinant TET3 enzyme of any one of claims 1 -2, wherein domain (a) comprises:(i) an amino acid sequence having amino acids 696-1048 or amino acids 697- 1048 of UniProt No. A0A5K1WP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 90% over the whole length thereof;(ii) an amino acid sequence having amino acids 831 -1183 of UniProt No. Q8BG87 (SEQ ID NO: 2) or an amino acid sequence having an identity of at least 90% over the whole length thereof;(iii) an amino acid sequence having amino acids 824-1175 of UniProt No. 043151 (SEQ ID NO: 3) or an amino acid sequence having an identity of at least 90% over the whole length thereof; or(iv)an amino acid sequence having amino acids 953-1304 of UniProt No. A0JP82 (SEQ ID NO: 4) or an amino acid sequence having an identity of at least 90% over the whole length thereof; and / or wherein domain (b) comprises:(i) an amino acid sequence having amino acids 1509-1599 of UniProt No. A0A5K1 WP6 (SEQ ID NO: 1 ) or an amino acid sequence having an identity of at least 90% over the whole length thereof;(ii) an amino acid sequence having amino acids 1644-1734 of UniProt No. Q8BG87 (SEQ ID NO: 2) or an amino acid sequence having an identity of at least 90% over the whole length thereof;(iii) an amino acid sequence having amino acids 1635-1726 of UniProt No. 043151 (SEQ ID NO: 3) or an amino acid sequence having an identity of at least 90% over the whole length thereof; or(iv)an amino acid sequence having amino acids 1742-1833 of UniProt No. A0JP82 (SEQ ID NO: 4) or an amino acid sequence having an identity of at least 90% over the whole length thereof.

4. The recombinant TET3 enzyme of any one of claims 1 -3, wherein the connecting domain has a length of 5-100, preferably 5-30, more preferably 10-20, and most preferably about 15 amino acids; and / or consists of amino acids selected from G, S, A, T, E, N, D, and / or K; and / or has the amino acid sequence[(Gn)Xm]r, wherein X is in each occurrence independently selected from S, E or D, n is 1 -5, particularly 2-4, m is 0-5, particularly 1 -3,r is 1 -5, particularly 2-4.

5. The recombinant TET3 enzyme of any one of claims 1 -4, which does not contain amino acid portions of a TET3 enzyme having a length of 10 amino acids or more, and particularly of 6 amino acids or more outside domain (a) and domain (b).

6. The recombinant TET3 enzyme of any one of claims 1 -5, wherein the amino acid T corresponding to amino acid T at position 940 of UniProt No. A0A5K1VVP6, amino acid T at position 1075 of UniProt No. Q8BG87, amino acid T at position 1067 of UniProt No. 043151 , and amino acid T at position 1196 is substituted with another amino acid, particularly with an amino acid selected from A, G, K and / or V, more particularly with A, and / or K. wherein the amino acid Y corresponding to amino acid Y at position 1567 of UniProt No. A0A5K1WP6, amino acid Y at position 1702 of UniProt No. Q8BG87, amino acid Y at position 1694 of UniProt No. 043151 or amino acid Y at position 1901 of UniProt No. A0JP82 is substituted with another amino acid, particularly with F or W.

7. The recombinant TET3 enzyme of any one of claims 1 -6, which catalyzes oxidation of 5-methyl-2'-deoxycytidine (mdC) to 5- hydroxymethyl-2' -deoxycytidine (hmdC), 5-formyl-2'-deoxycytidine (fdC) and / or 5- carboxy-2'-deoxycytidine (cadC), and which optionally catalyzes oxidation of 5- methylcytdine (mC) to 5-hydroxymethylcytidine (hmC), 5-formylcytidine (fC), and / or 5-carboxycytidine (caC).

8. A nucleic acid molecule encoding the TET3 enzyme of any one of claims 1 -7, optionally in operative linkage to an expression control sequence, or a vector comprising said nucleic acid molecule.

9. A host cell transformed or transfected with a nucleic acid molecule or a vector of claim 8, which is preferably an E. coli cell.

10. Use of a recombinant TET3 enzyme of any one of claims 1 -7 for the sequencing of a nucleic acid substrate, particularly for the sequencing of a nucleic acid substrate comprising mdC and / or mC nucleotides.11 . The use of claim 10, wherein the sequencing procedure is a single molecule-sequencing procedure.

12. A method for sequencing a nucleic acid substrate comprising mdC and / or mC nucleotides, wherein the method comprises the steps:(A) contacting the nucleic acid substrate with a recombinant TET3 enzyme of any one of claims 1 -7 under conditions, where mdC and / or mC nucleotides in the nucleic acid substrate are oxidized,(B) optionally isolating the oxidized nucleic acid substrate, and(C) determining the amount and / or position of mdC and / or mC nucleotides in the nucleic acid substrate.

13. The method of claim 12, wherein the oxidation in step (A) takes place.(i) under conditions, wherein mdC and / or mC nucleotides are selectively oxidized to cadC and / or caC nucleotides, wherein preferably the oxidation in step (A) is performed at an ion strength corresponding to 50-120 mM sodium chloride, more preferably corresponding to 60-100 mM sodium chloride;(ii) under conditions, wherein mdC and / or mC nucleotides are selectively oxidized to 5-hmdC, 5-hmC, 5-fdC and / or 5-fC nucleotides, wherein preferably the oxidation in step (A) is performed at an ion strength corresponding to 130-160 mM sodium chloride, more preferably corresponding to 140-150 mM sodium chloride; or(iii) under conditions, wherein mdC and / or mC are selectively oxidized to 5-fdC and / or 5-fC nucleotides, wherein preferably the oxidation in step (A) is performed at an ion strength corresponding to 180-250 mM sodium chloride, more preferably corresponding to 190-220 mM sodium chloride, even more preferably 200-210 mM.

14. A reagent for sequencing a nucleic acid substrate comprising mdC and / or mC nucleotides comprising:(A) a recombinant TET3 enzyme of any one of claims 1 -7 and(B) an oxidizing agent.

5. The reagent of claim 14, wherein the oxidizing agent is selected from the group consisting of a hypervalent iodine reagent, a nitroxyl radical-generating agent, Mn02, chromate, or any combination thereof.