Method for screening a molecule for binding to a protein of interest
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- UNIVERSITY OF BASEL
- Filing Date
- 2024-06-11
- Publication Date
- 2026-04-22
AI Technical Summary
Current methods for screening molecules that bind to target proteins, especially weak binders, are inefficient due to sensitivity to experimental conditions and kinetic issues in affinity selection protocols, limiting the detection of binding molecules in drug discovery.
A fusion protein comprising a DNA polymerase with terminal transferase activity and a protein of interest is used to record binding information directly into DNA-encoded libraries, allowing for the identification of binders through nucleic acid moiety extension and analysis, enhancing the retrieval of micromolar binders.
This approach significantly improves the detection of binding molecules, with nearly 90% of binders identified compared to 38% using state-of-the-art methods, providing a comprehensive picture of binding events and overcoming previous limitations.
Smart Images

Figure IMGF000032_0001 
Figure IMGF000033_0001 
Figure IMGF000033_0002
Abstract
Description
[0001] Method for screening a molecule for binding to a protein of interest
[0002] The field of the invention
[0003] The present invention relates to a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof, a method for screening a molecule for binding to a protein of interest (POI) or a fragment therof using such a fusion protein and a method for screening a DNA-encoded library of molecules for binding to a protein of interest (POI) or a fragment therof.
[0004] Background of the invention
[0005] Identification of molecules that bind to target proteins of interest is a formidable challenge. Technologies that facilitate the isolation of binding molecules have profound implications for pharmaceutical research because most drug development programs rely on the ability to isolate small organic compounds that bind to a given protein. With an aging population and an increased understanding of the mechanisms of disease at a molecular level, biomedical scientists are facing the demand for more and better drugs. Additionally, elucidation of the biological function of proteins will, in many cases, require access to specific ligands (an approach that is often termed 'Chemical Genetics' (Strausberg, R.L. and Schreiber, S. L., Science 300 (2003), 294-295). Techniques for the general, fast, inexpensive isolation of small, organic, binding molecules are lacking at present. Screening compound libraries for binding or inhibitory activity against target proteins is an essential part of early drug discovery. High- throughput screening (HTS) facilities, however, come with a large infrastructural burden. DNA encoded libraries (DELs) offer a radically different approach to initial hit-identification. DEL technology uses DNA tags to track the synthetic history of individual members in a split-and- pool combinatorial synthesis scheme. Since each compound’s identity is encoded in its covalently connected DNA, compounds can be pooled and then selected against target proteins in an affinity selection. After selection, the retained molecules are washed from beads and their DNA tags are sequenced with next generation sequencing (NGS). DNA sequencing is the final data readout, and the hope is that NGS read-counts give a somewhat faithful representation of binding affinity. Unfortunately, while this elegant concept works well for potent binders - information on weak binders is often below the detection threshold. Thus there is a need to provide effective tools and methods which allow identifying binders, in particular weak binders of proteins of interest.
[0006] Summary of the invention
[0007] The present invention relates to a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof, a method for screening a molecule for binding to a protein of interest (POI) or a fragment therof using such a fusion protein and a method for screening a DNA-encoded library of molecules for binding to a protein of interest (POI) or a fragment therof.
[0008] The inventors of the present invention have developed a new approach to DEL selection which uses DNA polymerase with terminal transferase activity to directly record binding information into DEL DNA. With this new approach the inventors were able to overcome many of the challenges of classical selection approaches for DELs as e.g. problems with the affinity selection protocol itself, which is normally different for every target and is highly sensitive to experimental conditions and kinetic issues (e.g. selecting for slow koff under the wash conditions) when equilibrium binding information would give a more comprehensive picture of binding event. The present invention provides the concept and demonstrated proof-of-principle on a known target with a purpose-built 105-member DEL. Surprisingly, the present invention has been proved to be significantly better at retrieving micromolar binders than state-of-the art methods. Nearly 90 % of binders were found using the method of the present invetion that were contained in a library compared to only 38 % with the state-of-the-art approach.
[0009] In a first aspect the present invention relates to a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof.
[0010] In a further aspect the present invention relates to a method for screening a molecule for binding to a protein of interest (POI) or a fragment therof, the method comprising the steps of: a) providing i) an incubation medium comprising a deoxynucleotide triphosphate (dNTP), ii) a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof and iii) a conjugate compound comprising a molecule and a nucleic acid moiety; and incubating the fusion protein and the conjugate compound in the incubation medium; b) separating conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended from conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is not extended; c) amplifying and analyzing the extended nucleic acid moiety of the conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended obtained in step b); d) correlating the extended nucleic acid moiety analyzed in step c) with the corresponding molecule of the conjugate compound used in step a), wherein the molecules correlated in step d) are selected as binder of the POI.
[0011] In a further aspect the present invention relates to a method for screening a DNA-encoded library of molecules for binding to a protein of interest (POI) or a fragment therof, the method comprising the steps of: a) providing i) an incubation medium comprising a deoxynucleotide triphosphate (dNTP), ii) a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof and iii) a DNA-encoded library of chemical molecules comprising multiple instances of one sole molecule of the library, each instance being covalently linked to a nucleic acid moiety to form conjugate compounds comprising a molecule and a nucleic acid moiety; and incubating the fusion protein and the conjugate compounds; b) separating conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended from conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is not extended; c) amplifying and analyzing the extended nucleic acid moiety of the conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended obtained in step b); d) correlating the extended nucleic acid moiety analyzed in step c) with the corresponding multiple instances of one sole chemical molecule of the library used in step a), wherein the multiple instances of one sole chemical molecule of the library so correlated in step d) are selected as binder of the POI. Brief description of the figures
[0012] Figure 1. (A) Depicted are the chemical structures of Conj_ID_No_4 (AAZ), Conj_ID_No_5 (CTZ) and Conj_ID_No_6 (SAA). (B) The proximity-induced extension of the DNA strand of the small-molecule DNA conjugates reveals that the mean extension leangth depends on the affinity of the small-molecule to the CAII part of the fusion protein.
[0013] Figure 2. Concentrations of the small-molecule DNA conjugates Conj_ID_No_4 (AAZ), Conj_ID_No_5 (CTZ) and Conj_ID_No_6 (SAA) relative to the concentration of the nonbinding Conj_ID_NO_8 (Amine).
[0014] Figure 3. Scatter plot of significantly enriched library members in an affinity enrichment against CAII on Ni-NTA beads. Position on the x-axis is determined by the number of the compound within diversity point 2 of the library; position on the y-axis is determined by the number of the compound within diversity point 1 of the library. Dots are scaled according to their log-Fold enrichment value. (A) single-stranded DNA-encoded library. (B) doublestranded DNA-encoded library.
[0015] Figure 4. Scatter plot of significantly enriched library members of a single-stranded DNA encoded library in a DELSTAR enrichment against CAII-SUMO-TdT. Position on the x-axis is determined by the number of the compound within diversity point 2 of the library; position on the y-axis is determined by the number of the compound within diversity point 1 of the library. Dots are scaled according to their log-Fold enrichment value. (A) Standard conditions (B) 0.5 pmol of CAII-SUMO-TdT instead of 1.0 pmol.
[0016] Figure 5. Scatter plot of significantly enriched library members of a double-stranded DNA encoded library in a DELSTAR enrichment against CAII-SUMO-TdT. Position on the x-axis is determined by the number of the compound within diversity point 2 of the library; position on the y-axis is determined by the number of the compound within diversity point 1 of the library. Dots are scaled according to their log-Fold enrichment value. (A) Standard conditions (B) 0.5 pmol of CAII-SUMO-TdT instead of 1.0 pmol.
[0017] Figure 6. (A) Activity test of Calmodulin-TdT in comparison to SUMO-TdT. Depicted is a ladder (AcuteBand, LubioScience), a 50mer piece of DNA (SEQ ID NO: 13) before incubation, incubation of this DNA with SUMO-TdT (SEQ ID NO:7) after 30 minutes at 37 °C at 50 nM concentration, and a 50mer DNA incubated with Calmodulin-TdT (SEQ ID NO:29) ad the same conditions. The extension with SUMO-TdT is slightly stronger, extending the strand by at least 50 nucleotides. Calmodulin extends the strand by at least 40 nucleotides. (B) An incubation of Calmodulin-TdT with DBCO-DNA or Calmodulin binding peptide- conjugated DNA under standard conditions. Lanes show that the proximity-induced elongation is stronger with the DNA-Calmodulin binding peptide conjugate (SEQ ID NO:33 and SEQ ID NO: 13) compared to background elongation on DBCO-DNA (SEQ ID NO: 13). (C) The SDS- PAGE shows SpyTag-CAII (Lane 2), SpyCatcher- SUMO-TdT (Lane 3) and the conjugatefusion of it (Lane 4). The conversion to the full conjugate is not complete. (D) Proximity- induced extension assay with standard conditions using SpyTag-CAII- SpyCatcher-TdT show formation of proximity-induced product with CAII-binding small molecule conjugates AAZ (Conjugate 1), CTZ (Conjugate 2) and SAA (3, see Table VII). A non-binding DBCO-DNA (SEQ ID NO 16) does not show extension after the pull-down with dT25 beads.
[0018] Figure 7. (A) Model system emulating the POI-small molecule interaction by using DNA- DNA base pairing. The DNA-SNAP-TdT adduct is incubated with a piece of DNA tuned to a specific dissociation constant simulated by introducing mismatches in the sequence. (B) The larger the dissociation constant, the smaller is the tail added to the DNA. If SNAP-TdT is used directly instead of DNA-SNAP-TdT, the non-proximity induced extension is minor and slightly smaller than with the weakest interaction.
[0019] Detailed description of the invention
[0020] As outlined above, the present invention provides a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof, a method for screening a molecule for binding to a protein of interest (POI) or a fragment therof using such a fusion protein and a method for screening a DNA-encoded library of molecules for binding to a protein of interest (POI) or a fragment therof.
[0021] Thus, in a first aspect the present invention provides a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof. For the purposes of interpreting this specification, the following definitions will apply and whenever appropriate, terms used in the singular will also include the plural and vice versa. It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
[0022] Features, integers, characteristics, compounds described in conjunction with a particular aspect, embodiment or example of the invention are to be understood to be applicable to any other aspect, embodiment or example described herein unless incompatible therewith. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The invention is not restricted to the details of any foregoing embodiments.
[0023] The term “comprise” and variations thereof, such as, “comprises” and “comprising” is generally used in the sense of include, that is, as “including, but not limited to” , that is to say permitting the presence of one or more features or components.
[0024] The singular forms “a,” “an,” and “the” include plural referents unless the content clearly dictates otherwise.
[0025] The term "about" refers to a range of values ± 10% of a specified value. For example, the phrase "about 200" includes ± 10% of 200, or from 180 to 220.
[0026] The term “DNA-encoded library of chemical molecules” or abbreviated as “DEL” as used hererin refers to a library of chemical molecules which are covalently bound to nucleic acid moieties like short nucleic acids that serve as identification bar codes and in some cases also direct and control the chemical synthesis. The technique enables the mass creation via split- and-pool synthesis, and interrogation of libraries via affinity selection, typically on an immobilized protein target. Double stranded (ds) DNA-encoded library of chemical molecules and single stranded (ss) DNA-encoded library of chemical molecules can be used in the present invention, preferably a single stranded (ss) DNA-encoded library of chemical molecules is used. The term “DNA polymerase with terminal transferase activity or a fragment thereof’ as used hererin refers to a DNA polymerase or a fragment thereof which catalyzes the addition of deoxynucleotides to the 3' hydroxyl terminus of DNA molecules. A preferred DNA polymerase with terminal transferase activity or a fragment thereof is a terminal deoxynucleotidyl transferase (TDT) or a fragment thereof. A fragment of a DNA polymerase with terminal transferase activity contains usually between 100 and 1000 amino acids, preferably between 200 and 600 amino acids, more preferably between 300 and 500 amino acids, even more preferably between 350 and 400 amino acids, even more preferably about 379 amino acids.
[0027] The term “terminal deoxynucleotidyl transferase (TDT) or a fragment thereof’ also known as DNA nucleotidylexotransferase (DNTT) or terminal transferase as used herein refers to a DNA polymerase that catalyses the addition of nucleotides to the 3' terminus of a DNA molecule. Unlike most DNA polymerases, it does not require a template. The preferred substrate of this enzyme is a 3'-overhang, but it can also add nucleotides to blunt or recessed 3' ends. . A fragment of a TDT contains usually between 100 and 1000 amino acids, preferably between 200 and 600 amino acids, more preferably between 300 and 500 amino acids, even more preferably between 350 and 400 amino acids, even more preferably about 379 amino acids.
[0028] The term “protein of interest (POI) or a fragment therof ’ as used hererin refers to any protein, preferably proteins with human pharmaceutical activity i.e. a therapeutic protein or therapeutic enzymes. Examples of proteins of interest with human pharmaceutical activity are cytokines, growth factors, hormones antibodies or vaccines. The term “fragment therof’ and the term “functionally active fragments” are equivalenty used herein. The term "protein of interest (POI) or a fragment thereof' includes naturally occurring proteins of interest (POI) and also includes artificially engineered proteins of interest (POI). Artificially proteins of interest (POI) are e.g. variants or functionally active fragments of the POI. By “variants or functionally active fragments thereof’ in relation to the POI of the present invention is meant that the fragment or variant (such as an analogue, derivative or mutant) is capable of exercising the same physiological function as the POI. Such variants include naturally occurring allelic variants and non-naturally occurring variants. Additions, deletions, substitutions and derivatizations of one or more of the amino acids are contemplated so long as the modifications do not result in loss of functional activity of the fragment or variant. Preferably the functionally active fragment or variant has at least about 80% sequence identity more preferably at least about 90% sequence identity, even more preferably at least about 95% sequence identity, most preferably at least about 98% sequence identity to the relevant part of the POI. A fragment of a POI as defined herein does have the same functional properties as the POI from which it is derived. A POI contains usually between 100 and 1000 amino acids, preferably between 200 and 800 amino acids, more preferably between 300 and 700 amino acids, even more preferably between 300 and 600 amino acids. A fragment of a POI contains usually between 25 and 500 amino acids, preferably between 50 and 400 amino acids, more preferably between 100 and 300 amino acids, even more preferably between 100 and 200 amino acids.
[0029] The term “linker between the DNA polymerase with terminal transferase activity or a fragment therof and the protein of interest (POI) or a fragment therof ’ as used hererin refers to linker e.g. a linker between 1 and 200, preferably between 5 and 150, more preferably between 10 and 150 amino acids or a linker comprising a chemical entity. The linker is usually introduced in between the amino terminal end of the DNA polymerase with terminal transferase activity and the carboxyl terminal end of the protein of interest (POI).
[0030] The term “conjugate compound comprising a binder of the POI or a fragment thereof and a nucleic acid moiety” as used hererin refers to compound wherein a binder of the POI or a fragment thereof is coupled, preferably covalently coupled to a nucleic acid moiety e.g. by coupling a DBCO modified DNA with an azido-functionalized binder of the POI or a fragment thereof.
[0031] The term “binder of the POI” as used hererin refers to molecule e.g. a chemical entity or a peptide which is capable to bind to the POI, preferably a molecule e.g. a chemical entity or a peptide which is capable to bind to the POI so that the binding causes a modification of the activity and / or chemical structure of the POI. A preferred binder of the POI is a chemical entity.
[0032] The term “nucleic acid moiety” as used hererin refers to a unit, or building block, on nucleic acid level including coding and non-coding nucleic acid moieties. Preferably the nucleic acid moiety is DNA, more preferably ssDNA or dsDNA, even more preferably ssDNA. More preferably the nucleic acid moiety is a non-coding nucleic acid moiety, even more preferably the nucleic acid moiety is a non-coding DNA, in particular a non-coding ssDNA or dsDNA more particular a non-coding ssDNA. The nucleic acid moiety of the conjugate compound usually comprises between 5 and 1500, 10 and 100, 10 and 500 or 5 and 200 nucleic bases, preferably between 5 and 1500 nucleic bases. The nucleic acid moiety of the conjugate compound when interacting with the DNA polymerase of the fusion protein, usually provides for a poly-A tailing of the binder of the POI or a fragment thereof.
[0033] The term "tag" as used herein, may encompass affinity tags, solubilization tags, chromatography tags and epitope tags. Affinity tags (also used as purification tags) are appended / fused to proteins so that they allow purification of the tagged molecule from their crude biological source using an affinity purification techniques. These include glutathione-S- transferase (GST), biotin, modified biotin and poly(His) tag. The poly(His) tag is a widely-used tag; it binds to metal-containing matrices. Solubilization tags are used, especially for recombinant proteins expressed in chaperone- deficient species such as E. coli, to assist in the proper folding in proteins and keep them from precipitating. These include thioredoxin (TRX) and poly(NANP).
[0034] In some embodiments the fusion protein comprises a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof, with the proviso that the fusion protein is not a fusion protein comprising the terminal deoxynucleotidyl transferase (TDT) and Cas9.
[0035] In some embodiments the fusion protein comprises a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof, with the proviso that the fusion protein is not a fusion protein comprising the terminal deoxynucleotidyl transferase (TDT) and a SUMO protein or a fragment thereof.
[0036] In some embodiments the fusion protein comprises a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof, with the proviso that the fusion protein is not a fusion protein comprising the terminal deoxynucleotidyl transferase (TDT) and a SUMO protein or a fragment thereof, and with the proviso that the fusion protein is not a fusion protein comprising the terminal deoxynucleotidyl transferase (TDT) and a Glutathione S-transferase or a fragment thereof.
[0037] In some embodiments the fusion protein comprises a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof, with the proviso that the fusion protein is not a fusion protein comprising the terminal deoxynucleotidyl transferase (TDT) and Cas9, with the proviso that the fusion protein is not a fusion protein comprising the terminal deoxynucleotidyl transferase (TDT) and a SUMO protein or a fragment thereof, and with the proviso that the fusion protein is not a fusion protein comprising the terminal deoxynucleotidyl transferase (TDT) and a Glutathione S- transferase or a fragment thereof.
[0038] In some embodiments the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof comprises a fragment of a DNA polymerase with terminal transferase activity and a protein of interest (POI) or a fragment therof, preferably a fragment of the terminal deoxynucleotidyl transferase (TDT) and a protein of interest (POI) or a fragment therof. The fragment of the DNA polymerase with terminal transferase activity usually comprises the C-terminal POLXc domain and more specifically, the 8kDa-domain, the finger domain, the palm domain, and the thumb domain. The fragment of the DNA polymerase with terminal transferase activity is preferably a fragment of the terminal deoxynucleotidyl transferase (TDT) with 43 kDa.
[0039] In some embodiments the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof comprises a DNA polymerase with terminal transferase activity or a fragment of a DNA polymerase with terminal transferase activity, preferably a fragment of a DNA polymerase with terminal transferase activity, and two or more copies of the protein of interest (POI) or of a fragment thereof, preferably a terminal deoxynucleotidyl transferase (TDT) or a fragment thereof, preferably a fragment of the terminal deoxynucleotidyl transferase (TDT) and two or more copies of the protein of interest (POI) or of a fragment therof. In some embodiments the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof comprises two or more copies of a DNA polymerase with terminal transferase activity or of a fragment of a DNA polymerase with terminal transferase activity, preferably two or more copies of a fragment of a DNA polymerase with terminal transferase activity and a protein of interest (POI) or a fragment thereof, preferably two or more copies of a terminal deoxynucleotidyl transferase (TDT) or of a fragment thereof, preferably two or more copies of a fragment of a terminal deoxynucleotidyl transferase (TDT) and a protein of interest (POI) or a fragment therof.
[0040] In some embodiments the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof comprises two or more copies of a DNA polymerase with terminal transferase activity or of a fragment of a DNA polymerase with terminal transferase activity, preferably two or more copies of a fragment of a DNA polymerase with terminal transferase activity and two or more copies of the protein of interest (POI) or a fragment thereof, preferably two or more copies of a terminal deoxynucleotidyl transferase (TDT) or of a fragment thereof, preferably two or more copies of a fragment of a terminal deoxynucleotidyl transferase (TDT) and two or more copies of the protein of interest (POI) or of a fragment thereof.
[0041] Two or more copies of the protein of interest (POI) or of a fragment thereof are preferably two copies of the protein of interest (POI) or of a fragment thereof, three copies of the protein of interest (POI) or of a fragment thereof, four copies of the protein of interest (POI) or of a fragment thereof, or five copies of the protein of interest (POI) or of a fragment thereof, more preferably two to five copies of the protein of interest (POI) or of a fragment thereof. Copies of the protein of interest or of a fragment therof usually comprise multiples of the identical protein sequence of the protein of interest or of a fragment therof.
[0042] Two or more copies of the DNA polymerase with terminal transferase activity or of a fragment of a DNA polymerase with terminal transferase activity are preferably two copies of the DNA polymerase with terminal transferase activity or of a fragment of a DNA polymerase with terminal transferase activity, three copies of the DNA polymerase with terminal transferase activity or of a fragment of a DNA polymerase with terminal transferase activity, four copies of the DNA polymerase with terminal transferase activity or of a fragment of a DNA polymerase with terminal transferase activity, five copies of the DNA polymerase with terminal transferase activity or of a fragment of a DNA polymerase with terminal transferase activity, more preferably two to five copies of the DNA polymerase with terminal transferase activity or of a fragment of a DNA polymerase with terminal transferase activity. Copies of the DNA polymerase with terminal transferase activity or of a fragment therof usually comprise multiples of the identical nucleic acid sequence of the DNA polymerase with terminal transferase activity or of a fragment therof.
[0043] In some embodiments the fusion protein comprises a linker between the DNA polymerase with terminal transferase activity or a fragment therof and the protein of interest (POI) or a fragment therof. Preferably, the linker comprises between 1 and 300 amino acids, preferably between 10 and 150 amino acids, more preferably between 30 and 150 amino acids, even more preferably between 40 and 150 amino acids or the linker comprises or is a chemical entity. In one embodiment the linker comprises a SUMO protein or a fragment thereof e.g. a SUMO tag and comprises between 1 and 300 amino acids, preferably between 10 and 150 amino acids and is even more preferably the amino acid sequence as shown in SEQ ID NO: 6. In a further embodiment the linker comprises SEQ ID NO: 44. In a further embodiment the linker comprises a SpyTag connected to a SpyCatcher [8], wherein the linker comprises preferably the amino acid sequence shown in SEQ ID NO: 45 connected to the amino acid sequence shown in SEQ ID NO: 46, wherein preferably the amino acid sequences of SEQ ID NO: 45 and SEQ ID NO: 46 are connected via an isopeptide bond between the first D from the amino terminal end of SEQ ID NO: 45 and and the second K from the amino terminal end of SEQ ID NO: 46. In a preferred embodiment the linker is selected from the amino acid sequence as shown in SEQ ID NO: 6, the amino acid sequence as shown in SEQ ID NO: 44, the amino acid sequence as shown in SEQ ID NO: 45, the amino acid sequence as shown in SEQ ID NO: 45, and combinations thereof. In some embodiments, the linker comprises a chemical entity. Thus in some embodiments the fusion protein comprises a linker between the DNA polymerase with terminal transferase activity or a fragment therof and the protein of interest (POI) or a fragment therof, wherein the linker is a chemical entity. Chemical entities which can be used as linkers in the present invention are usually common bioconjugation motifs known to those skilled in the art and are preferably selected from the group consisting of maleimide-benzyl guanine, maleimide-hexylchloride, activated ester-benzyl guanine, activated ester-hexyl chloride conjugates preferentially with a short carbon or PEG linker.
[0044] In a preferred embodiment the fusion protein comprises one or more, preferably two to five, linkers between the DNA polymerase with terminal transferase activity or a fragment therof and the protein of interest (POI) or a fragment therof, wherein the linker is a chemical entity, and wherein the one or more, preferably two to five, linkers are fused each to an amino acid of the protein of interest (POI) or of a fragment therof different from the N-terminal and / or C- terminal amino acid of the protein of interest (POI) or of a fragment therof, on one part of the linker, and wherein the same one or more linkers are fused each to a DNA polymerase with terminal transferase activity or to a fragment therof on another part of the linker.
[0045] In some embodiments the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof comprises a tag. The tag can be located at the N-terminal or at the C-terminal site of the fusion protein and is preferably located at the N-terminal of the fusion protein.
[0046] In some embodiments the DNA polymerase with terminal transferase activity or a fragment therof is a terminal deoxynucleotidyl transferase (TDT) or a fragment thereof, preferably is a terminal deoxynucleotidyl transferase (TDT) or a fragment thereof comprised by the amino acid sequence as shown in SEQ ID NO: 5.
[0047] In some embodiments the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof comprises a DNA polymerase with terminal transferase activity or a fragment thereof, preferably a terminal deoxynucleotidyl transferase (TDT) or a fragment thereof, preferably a terminal deoxynucleotidyl transferase (TDT) or a fragment thereof comprised by the amino acid sequence as shown in SEQ ID NO: 5 and calmodulin or a fragment thereof, preferably calmodulin or a fragment thereof comprised by the amino acid sequence as shown in SEQ ID NO: 43. In a peferred embodiment the fusion protein comprises a DNA polymerase with terminal transferase activity or a fragment thereof and calmodulin or a fragment thereof. In particular the fusion protein comprises a terminal deoxynucleotidyl transferase (TDT) or a fragment thereof and calmodulin or a fragment thereof as shown in SEQ ID NO: 29.
[0048] In some embodiments the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof comprises a sequence selected from the group as shown in SEQ ID NOs: 7-12, and comprises preferably a sequence selected from the group as shown in SEQ ID NOs: 8, 9, 11 and 12 in particular as shown in SEQ ID NOs: 11 or 12, more preferably a sequence selected from the group as shown in SEQ ID NOs: 8, 9, 11, 12, 29, 31 and 32, in particular as shown in SEQ ID NOs: 11, 12, 29 or 31, even more preferably a sequence selected from the group as shown in SEQ ID NOs: 8, 9, 11, 12 and 29, in particular as shown in SEQ ID NOs: 11, 12 or 29.
[0049] In some embodiments the POI or a fragment therof is selected from the group consisting of an enzyme or a fragment thereof and a protein or a fragment thereof targeting or involved in cellular proliferation, preferably a protein or a fragment thereof targeting or involved in cellular proliferation. More preferably, the POI or a fragment therof is a protein or a fragment thereof targeting or involved in cellular proliferation selected from the group consisting of IL2, RIPK1, MAPK14, PARP1, SIRT3, PI3Kbeta, PI3Kgamma, PI3Kdelta, BRCA1, BRCA2, BRAT1, INTS9, and INTS11, a fragment thereof or variants thereof. Even more preferably the POI is selected from the group consisting of carbonic anhydrase II (CAII), Rad52, Rad51, KRAS- G12C, PI3Kalpha, PALB2, and DCAF15 a fragment thereof or variants thereof.
[0050] In a further aspect the present invention provides a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof, wherein the fusion protein is bound to a conjugate compound comprising a binder of the POI or a fragment thereof and a nucleic acid moiety. By “fusion protein is bound to a conjugate compound comprising a binder of the POI or a fragment thereof and a nucleic acid moiety” is meant herein that the fusion protein and the conjugate compound can be connected by non-covalent binding or covalent binding. Non-covalent binding includes p-p (aromatic) interactions, van der Waals interactions, H-bonding interactions, and ionic interactions. Preferably, the fusion protein and the conjugate compound are bound by non-covalent binding. The fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof are as described above.
[0051] In some embodiments the conjugate compound comprising a binder of the POI or a fragment thereof and a nucleic acid moiety is a conjugate wherein the binder is a molecule which is covalently coupled to the nucleic acid moiety. In some embodiments the binder of the POI or a fragment thereof is a chemical entity or a peptide, preferably a chemical entity. In some embodiments the binder of the POI or a fragment thereof is a peptide, selected from the group consisting of streptag, streptavidin, streptactin, SnoopTag, SnoopCachter, and a calmodulin- binding peptide. Preferably the peptide is a calmodulin-binding peptide, more preferably the calmodulin-binding peptide as shown in SEQ ID NO 33. In some embodiments the nucleic acid moiety of the conjugate is ssDNA or dsDNA, preferably ssDNA. In some embodiments the conjugate comprises the DNA as shown in SEQ ID NOs 13-15 or 16-18 and a chemical entity.
[0052] Thus in a further aspect the present invention provides a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein or a fragment therof, wherein the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and the protein or a fragment therof, is bound to a further fusion protein comprising a peptide and a POI or a fragment thereof, wherein the petide is capable to bind to the protein or a fragment therof of the fusion protein. The protein or a fragment therof is usually selected from a protein or a fragment therof, which has a corresponding binding peptide such as a protein or a fragment therof being one part of known binding systems such as streptag / streptavidin or streptactin, SnoopTag / SnoopCachter, calmodulin / calmodulin-binding peptide, wherein the petide capable to bind to the protein or a fragment therof of the fusion protein being the other part of these binding systems. In one embodiment the fusion protein comprises a DNA polymerase with terminal transferase activity or a fragment thereof and calmodulin or a fragment thereof, wherein the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and calmodulin or a fragment thereof is bound to a further fusion protein comprising a calmodulin-binding peptide and a POI or a fragment thereof. By “fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and calmodulin or a fragment therof is bound to a further fusion protein comprising a calmodulin-binding peptide and a POI or a fragment thereof’ is meant herein that the (first) fusion protein and the further (second) fusion protein can be connected by non-covalent binding or covalent binding. Non-covalent binding includes p-p (aromatic) interactions, van der Waals interactions, H-bonding interactions, and ionic interactions. Preferably, the (first) fusion protein and the further (seconds) fusion protein are bound by non- covalent binding. DNA polymerase with terminal transferase activity or a fragment thereof, protein of interest (POI) or a fragment therof, calmodulin or a fragment therof and calmodulin-binding peptide are as described above.
[0053] In a further aspect the present invention provides a method for screening a molecule for binding to a protein of interest (POI) or a fragment therof, the method comprising the steps of: a) providing i) an incubation medium comprising a deoxynucleotide triphosphate (dNTP), ii) a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof and iii) a conjugate compound comprising a molecule and a nucleic acid moiety; and incubating the fusion protein and the conjugate compound in the incubation medium; b) separating conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended from conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is not extended; c) amplifying and analyzing the extended nucleic acid moiety of the conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended obtained in step b); d) correlating the extended nucleic acid moiety analyzed in step c) with the corresponding molecule of the conjugate compound used in step a), wherein the molecules correlated in step d) are selected as binder of the POI.
[0054] The fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof and the conjugate compound comprising a molecule and a nucleic acid moiety to be used for the method are as described above.
[0055] The incubation usually comprises dATP, dCTP dGTP and / or dTTP, preferably the incubation medium comprises dATP. The fusion protein and the conjugate compound is usually incubated in the incubation medium for 10 to 60 min., preferably incubated at 37 °C for 30 min. In some embodiments in step a) the DNA polymerase with terminal transferase activity or a fragment thereof is inactivated prior to performing step b), preferably inactivated prior to performing step b) by heat treatment, more preferably by heating the incubate of step a) for 10 min at 98 °C.
[0056] In some embodiments step b) comprises adding a solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound to the incubation medium and separating the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound hybridized to the extended nucleic acid moiety of the conjugate compound from the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound not hybridized to the extended nucleic acid moiety of the conjugate compound.
[0057] In some embodiments step b) comprises adding polydeoxynucleotide beads comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound to the incubation medium and separating polydeoxynucleotide beads comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound hybridized to the extended nucleic acid moiety of the conjugate compound from poly deoxynucleotide beads comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound not hybridized to the extended nucleic acid moiety of the conjugate compound.
[0058] In some embodiments the conjugate compounds with the extended nucleic acid moiety are eluted from the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound or the polydeoxynucleotide beads comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound.
[0059] In some embodiments a washing step is performed after the conjugate compounds with the extended nucleic acid moiety are eluted from the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound or the poly deoxynucleotide beads comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound.
[0060] In some embodiments in the conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended, the nucleic acid moiety is extended by between 1 and 200 nucleotides, preferably by between 10 and 100 nucleotides, more preferably by between 20 and 50 nucleotides.
[0061] In some embodiments the extended nucleic acid moiety is amplified and sequenced in step c).
[0062] Correlation of the extended nucleic acid moiety analyzed in step c) with the corresponding molecule of the conjugate compound used in step a) is ususally performed by measuring the concentration of the individual strands, preferably by measuring the concentration of the specific strand by quantitative PCR with selective primers or by counting the occurence of the barcode after sequencing and comparing it to a result with no protein of interest.
[0063] In a further aspect the present invention provides a method for screening a DNA-encoded library of molecules for binding to a protein of interest (POI) or a fragment therof, the method comprising the steps of: a) providing i) an incubation medium comprising a deoxynucleotide triphosphate (dNTP), ii) a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof and iii) a DNA-encoded library of chemical molecules comprising multiple instances of one sole molecule of the library, each instance being covalently linked to a nucleic acid moiety to form conjugate compounds comprising a molecule and a nucleic acid moiety, and incubating the fusion protein and the conjugate compounds; b) separating conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended from conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is not extended; c) amplifying and analyzing the extended nucleic acid moiety of the conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended obtained in step b); d) correlating the extended nucleic acid moiety analyzed in step c) with the corresponding multiple instances of one sole chemical molecule of the library used in step a), wherein the multiple instances of one sole chemical molecule of the library so correlated in step d) are selected as binder of the POI.
[0064] The DNA-encoded library of chemical molecules can be a double stranded (ds) DNA-encoded library of chemical molecules and single stranded (ss) DNA-encoded library of chemical molecules and is preferably a single stranded (ss) DNA-encoded library of chemical molecules.
[0065] The incubation usually comprises dATP, dACTP dGTP and / or dTTP, preferably the incubation medium comprises (dATP). The fusion protein and the conjugate compound is usually incubated in the incubation medium for 10 to 60 min, preferably incubated at 37°C for 30 min.
[0066] Fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof and conjugate compound comprising a molecule and a nucleic acid moiety to be used for the method are as described above.
[0067] In some embodiments in step a) the DNA polymerase with terminal transferase activity or a fragment thereof is inactivated prior to performing step b), preferably inactivated prior to performing step b) by heat treatment, more preferably by heating the incubate of step a) for 10 min at 98° C.
[0068] In some embodiments step b) comprises adding a solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound to the incubation medium and separating the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound hybridized to the extended nucleic acid moiety of the conjugate compound from the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound not hybridized to the extended nucleic acid moiety of the conjugate compound. In some embodiments step b) comprises adding polydeoxynucleotide beads comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound to the incubation medium and separating polydeoxynucleotide beads comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound hybridized to the extended nucleic acid moiety of the conjugate compound from poly deoxynucleotide beads comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound not hybridized to the extended nucleic acid moiety of the conjugate compound .
[0069] In some embodiments the conjugate compounds with the extended nucleic acid moiety are eluted from the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound or the polydeoxynucleotide beads comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound.
[0070] In some embodiments a washing step is performed after the conjugate compounds with the extended nucleic acid moiety are eluted from the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound or the polydeoxynucleotide beads comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound.
[0071] In some embodiments in the conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended, the nucleic acid moiety is extended by between 1 and 200 nucleotides, preferably by between 10 and 100 nucleotides, more preferably by between 20 and 50 nucleotides.
[0072] In some embodiments the extended nucleic acid moiety is amplified and sequenced in step c).
[0073] Correlation of the extended nucleic acid moiety analyzed in step c) with the corresponding molecule of the conjugate compound used in step a) is ususally performed by is ususally performed by measuring the concentration of the individual strands, preferably by measuring the concentration of the specific strand by quantitative PCR with selective primers or by counting the occurence of the barcode after sequencing and comparing it to a result with no protein of interest.
[0074] The libraries as used in the above method of the invention provide a repertoire of chemical diversity such that each chemical moiety is linked to a DNA moiety that facilitates identification of the chemical moiety.
[0075] By the present screening methods, it is possible to identify optimised chemical structures that participate in binding interactions with a protein of interest by drawing upon a repertoire of structures randomly formed by the association of diverse building blocks without the necessity of either synthesising them one at a time or knowing their interactions in advance.
[0076] The incubation of the library with the other components provided can be in the form of a heterogeneous or homogeneous admixture. Thus, the members of the library can be in the solid phase with the fusion protein present in the liquid phase. Alternatively, the fusion protein can be in the solid phase with the members of the library present in the liquid phase. Still further, both the library members and the fusion protein can be in the liquid phase.
[0077] Binding conditions are those conditions compatible with the known natural binding function of the POI. Those compatible conditions are buffer, pH and temperature conditions that maintain the biological activity of the POI and the DNA polymerase with terminal transferase activity of a fragment thereof, thereby maintaining the ability of the molecule to participate in its preselected binding interaction. Typically, those conditions include an aqueous, physiologic solution of pH and ionic strength normally associated with the POI.
[0078] The use of solid supports is generally well known in the art. Useful solid support matrices are well known in the art and include cross- linked dextran such as that available under the tradename SEPHADEX from Pharmacia Fine Chemicals (Piscataway, N. J.); agarose, borosilicate, polystyrene or latex beads about 1 micron to about 5 millimeters in diameter, polyvinyl chloride, polystyrene, cross-linked polyacrylamide, nitrocellulose or nylon-based webs such as sheets, strips, paddles, plates microtiter plate wells and the like insoluble matrices. Examples
[0079] Example 1:
[0080] A) Materials and Methods
[0081] Bacterial strains and growth conditions. For plasmid purification and cloning, E. coli DH5a were grown on LB agar planes and in LB broth at 37 °C. Ampicillin was used at a concentration of 100 pg / mL to select for the expression vectors. For protein expression, E. coli BL21(DE3) was used instead at the same conditions and were induced at a OD600 between 0.5 and 0.6 with 0.5 mM Isopropyl P-D-l -thiogalactopyranoside (IPTG).
[0082] Construction of plasmids. For all expressions, the pET19b vector backbone has been used. The sequence of the thermostable terminal di deoxynucleotidyl transferase used (TdTevo hereafter, SEQ ID NO: 5) was adapted from Barthel et al.[l] The sequence of carbonic anhydrase II (CAII hereafter) was a gift from Ryan Mehl (RRID:Addgene_l 05665). The linker used between TdT and the protein of interest is given in SEQ ID NO: 6. The nucleotide sequences inserted into the backbone for all fusion proteins including the first nucleotide of the Ncol restriction size and including the last nucleotide of the BamHI restriction site are given in Table I and the corresponding plasmids and fusions in Table II and the corresponding protein sequences in Table III
[0083] Table I.
[0084] SEQ ID NO: 1
[0085] CCATGGGCCATCATCATCATCATCATCTCGAGTCAGTGAGCATGAGCGATTCCGAA GTGAACCAGGAAGCCAAACCCGAAGTCAAACCGGAGGTGAAACCAGAAACCCAC ATTAATTTGAAGGTGTCGGACGGCTCATCAGAAATTTTTTTCAAAATTAAAAAAAC CACCCCGTTACGTAGGCTGATGGAAGCGTTCGCGAAACGCCAAGGGAAGGAAATG GATAGTCTCCGGTTTTTATATGATGGCATTCGCATTCAAGCGGATCAAACGCCAGA AGATTTAGACATGGAAGATAATGATATAATTGAGGCGCATCGCGAACAGACTAGT AATTCGAGCTCGAACAACAACAACAATAACAATAACAACAACCTCGGGATCGAGG GAAGGATTTCACACATGTCTATGGGCGGCCGCGATATCGTCGACGGCTCCGAATT CTCACCGAGTCCCGTTCCGGGTAGCCAGAACGTCCCTGCTCCGGCCGTGAAGAAG ATCTCGCAGTACGCCTGCCAACGGCGGACCACTCTTAATAATTATAACCAACTGTT TACAGATGCGCTGGAAATCTTAGCTGAAAACGCTGAGTTTCGTGAGAACGAGGGA CGTTGCCTTGCTTTCATGCGTGCGGCATCAGTTCTGAAGAGTTTACCTTTCCCTAT AACGAGTATGAAAGACCTGGAGGGCTTACCCTGCTTAGGGGACAAAGTTAAGCGT
[0086] ATTATAGAAGAAATACTTGAGGACGGTGAGAGTTCAGAAGCGAAGGCGGTACTGA
[0087] ATGACGAGCGTTACAAATCGTTCAAGCTGTTTACATCGGTCTTCGGTGTGGGATTG
[0088] AAGACTGCCGAAAAGTGGTACAGAATGGGCTTTAGAACCTTGAGCAAGATACAGA
[0089] GCGACAAGTCGCTGCGTCTTACGCAAATGCAAAAGGCAGGTTTCTTGTACTACGA
[0090] GGACCTGGTCTCCTGTGTTAATAGACCTGAGGCCGAGGCGGTCTCCATGCTTGTAA
[0091] AAGAAGCAGTTGTTACGTTTTTGCCGGGTGCTTTGGTGACCCTTACCGGCGGCTTC
[0092] CGTAGAGGGAAGATGACTGGGCATGATGTGGATTTCCTGATTACCTCCCCGGAGG
[0093] CCGGTGAGGACGAAGAGCAGCAGTTACTTCATAAAGTCACAGATTTCTGGAAACA
[0094] GCAAGGGCTTTTACTTTATTGCGACATCTTGGAGTCTACATTTGAAAAGTTTAAAC
[0095] AACCATCGCGCAAAGTTGATGCTCTTGACCACTTTCAGAAGTGCTTCCTGATATTG
[0096] AAGTTGGATCATGGTCGGGTGCACTCAGAGAAGTCTGGCCAGCAGGAGGGCAAAG
[0097] GCTGGAAGGCAATACGGGTTGATCTTGTGATGTGTCCCTATGACCGTAGAGCCTTT
[0098] GCATTGCTTGGGTGGACGGGTTCCCGCCAATTTGAAAGAGATTTAAGACGTTATG
[0099] CCACCCACGAAAGAAAAATGATGCTGGATAACCACGCACTGTATGACCGCACAAA
[0100] GCGGGTGTTTCTTGAGGCGGAGTCAGAAGAAGAGATCTTTGCGCACTTAGGTTTA
[0101] GACTACATCGAGCCGTGGGAACGTAACGCCTAAGGATCC
[0102] SEQ ID NO: 2
[0103] CCATGGGCCATCATCATCATCATCATGAAGTTAAACCGGAAACCCACATCAACCTG
[0104] AAAGTTTCTGACGGTTCTTCTGAAATCTTCTTCAAAATCAAAAAAACCACCCCGCT
[0105] GCGTCGTCTGATGGAAGCTTTCGCTAAACGTCAGGGTAAAGAAATGGACTCTCTG
[0106] CGTTTCCTGTACGACGGTATCCGTATCCAGGCTGACCAGACCCCGGAAGACCTGG
[0107] ACATGGAAGACAACGACATCATCGAAGCTCACCGTGAACAGCTGGCTGAGAATCT
[0108] TTATTTTCAGGGCCATATGGCGCATCATTGGGGTTACGGTAAACACAACGGTCCGG
[0109] AGCATTGGCACAAAGATTTTCCAATTGCGAAGGGCGAACGTCAAAGCCCGGTTGA
[0110] CATTGATACGCACACGGCAAAGTACGACCCGAGCCTGAAACCGCTGAGCGTTTCC
[0111] TATGACCAGGCTACGAGCCTGCGTATCCTGAACAATGGCCACACCTTCAACGTGG
[0112] AGTTTGATGATTCCCAAGATAAGGCGGTTCTGAAAGGTGGTCCGTTGGATGGCAC
[0113] CTACCGCCTGATCCAATTTCACTTTCACTGGGGTAGCCACGACGGTCAGGGCAGCG
[0114] AGCATACCGTGGACAAAAAGAAGTATGCAGCCGAACTGCACCTGGTGCATTGGAA
[0115] CACGAAGTACGGCGACTTCGGTAAAGCGGTCCAGCAACCGGACGGTCTGGCTGTT CTGGGTATTTTCCTGAAGGTCGGCAGCGCGAACCCGGGTCTGCAGAAAGTGGTTG
[0116] ACGTGTTGGACTCTATCAAGACCAAAGGCAAGAGCGCGGACTTCACCAATTTCGA
[0117] TCCGCGTGGTCTGCTGCCGGAGAGCCTGGATTACTGGACTTATCCGGGCAGCCTG
[0118] ACCACCCCGCCATTGCTGGAGTGCGTGACCTGGATCGTCTTGAAAGAACCGATCA
[0119] GCGTTAGCTCTGAACAGGTCAGCAAGTTCCGCAAGCTGAATTTCAATGGTGAGGG
[0120] CGAGCCGGAAGAACCGATGGTCGATAATTGGCGTCCTACCCAACCGCTGAAAAAC
[0121] CGCCAGATTAAAGCATCCTTTAAGCTCGAGTCAGTGAGCATGAGCGATTCCGAAG
[0122] TGAACCAGGAAGCCAAACCCGAAGTCAAACCGGAGGTGAAACCAGAAACCCACAT
[0123] TAATTTGAAGGTGTCGGACGGCTCATCAGAAATTTTTTTCAAAATTAAAAAAACCA
[0124] CCCCGTTACGTAGGCTGATGGAAGCGTTCGCGAAACGCCAAGGGAAGGAAATGGA
[0125] TAGTCTCCGGTTTTTATATGATGGCATTCGCATTCAAGCGGATCAAACGCCAGAAG
[0126] ATTTAGACATGGAAGATAATGATATAATTGAGGCGCATCGCGAACAGACTAGTAA
[0127] TTCGAGCTCGAACAACAACAACAATAACAATAACAACAACCTCGGGATCGAGGGA
[0128] AGGATTTCACACATGTCTATGGGCGGCCGCGATATCGTCGACGGCTCCGAATTCTC
[0129] ACCGAGTCCCGTTCCGGGTAGCCAGAACGTCCCTGCTCCGGCCGTGAAGAAGATC
[0130] TCGCAGTACGCCTGCCAACGGCGGACCACTCTTAATAATTATAACCAACTGTTTAC
[0131] AGATGCGCTGGAAATCTTAGCTGAAAACGCTGAGTTTCGTGAGAACGAGGGACGT
[0132] TGCCTTGCTTTCATGCGTGCGGCATCAGTTCTGAAGAGTTTACCTTTCCCTATAAC
[0133] GAGTATGAAAGACCTGGAGGGCTTACCCTGCTTAGGGGACAAAGTTAAGCGTATT
[0134] ATAGAAGAAATACTTGAGGACGGTGAGAGTTCAGAAGCGAAGGCGGTACTGAATG
[0135] ACGAGCGTTACAAATCGTTCAAGCTGTTTACATCGGTCTTCGGTGTGGGATTGAAG
[0136] ACTGCCGAAAAGTGGTACAGAATGGGCTTTAGAACCTTGAGCAAGATACAGAGCG
[0137] ACAAGTCGCTGCGTCTTACGCAAATGCAAAAGGCAGGTTTCTTGTACTACGAGGA
[0138] CCTGGTCTCCTGTGTTAATAGACCTGAGGCCGAGGCGGTCTCCATGCTTGTAAAAG
[0139] AAGCAGTTGTTACGTTTTTGCCGGGTGCTTTGGTGACCCTTACCGGCGGCTTCCGT
[0140] AGAGGGAAGATGACTGGGCATGATGTGGATTTCCTGATTACCTCCCCGGAGGCCG
[0141] GTGAGGACGAAGAGCAGCAGTTACTTCATAAAGTCACAGATTTCTGGAAACAGCA
[0142] AGGGCTTTTACTTTATTGCGACATCTTGGAGTCTACATTTGAAAAGTTTAAACAAC
[0143] CATCGCGCAAAGTTGATGCTCTTGACCACTTTCAGAAGTGCTTCCTGATATTGAAG
[0144] TTGGATCATGGTCGGGTGCACTCAGAGAAGTCTGGCCAGCAGGAGGGCAAAGGCT
[0145] GGAAGGCAATACGGGTTGATCTTGTGATGTGTCCCTATGACCGTAGAGCCTTTGCA
[0146] TTGCTTGGGTGGACGGGTTCCCGCCAATTTGAAAGAGATTTAAGACGTTATGCCAC CCACGAAAGAAAAATGATGCTGGATAACCACGCACTGTATGACCGCACAAAGCGG
[0147] GTGTTTCTTGAGGCGGAGTCAGAAGAAGAGATCTTTGCGCACTTAGGTTTAGACT
[0148] ACATCGAGCCGTGGGAACGTAACGCCTAATAAGGATCC
[0149] SEQ ID NO: 3
[0150] CCATGGGCCATCATCATCATCATCATGAAGTTAAACCGGAAACCCACATCAACCTG
[0151] AAAGTTTCTGACGGTTCTTCTGAAATCTTCTTCAAAATCAAAAAAACCACCCCGCT
[0152] GCGTCGTCTGATGGAAGCTTTCGCTAAACGTCAGGGTAAAGAAATGGACTCTCTG
[0153] CGTTTCCTGTACGACGGTATCCGTATCCAGGCTGACCAGACCCCGGAAGACCTGG
[0154] ACATGGAAGACAACGACATCATCGAAGCTCACCGTGAACAGCTGGCTGAGAATCT
[0155] TTATTTTCAGGGCCATATGATGGACAAAGATTGCGAAATGAAACGTACCACCCTG
[0156] GATAGCCCGCTGGGCAAACTGGAACTGAGCGGCTGCGAACAGGGCCTGCATGAAA
[0157] TTAAACTGCTGGGTAAAGGCACCAGCGCGGCCGATGCGGTTGAAGTTCCGGCCCC
[0158] GGCCGCCGTGCTGGGTGGTCCGGAACCGCTGATGCAGGCGACCGCGTGGCTGAAC
[0159] GCGTATTTTCATCAGCCGGAAGCGATTGAAGAATTTCCGGTTCCGGCGCTGCATCA
[0160] TCCGGTGTTTCAGCAGGAGAGCTTTACCCGTCAGGTGCTGTGGAAACTGCTGAAA
[0161] GTGGTTAAATTTGGCGAAGTGATTAGCTATCAGCAGCTGGCGGCCCTGGCGGGTA
[0162] ATCCGGCGGCCACCGCCGCCGTTAAAACCGCGCTGAGCGGTAACCCGGTGCCGAT
[0163] TCTGATTCCGTGCCATCGTGTGGTTAGCTCTAGCGGTGCGGTTGGCGGTTATGAAG
[0164] GTGGTCTGGCGGTGAAAGAGTGGCTGCTGGCCCATGAAGGTCATCGTCTGGGTAA
[0165] ACCGGGTCTGGGACTCGAGTCAGTGAGCATGAGCGATTCCGAAGTGAACCAGGAA
[0166] GCCAAACCCGAAGTCAAACCGGAGGTGAAACCAGAAACCCACATTAATTTGAAGG
[0167] TGTCGGACGGCTCATCAGAAATTTTTTTCAAAATTAAAAAAACCACCCCGTTACGT
[0168] AGGCTGATGGAAGCGTTCGCGAAACGCCAAGGGAAGGAAATGGATAGTCTCCGGT
[0169] TTTTATATGATGGCATTCGCATTCAAGCGGATCAAACGCCAGAAGATTTAGACATG
[0170] GAAGATAATGATATAATTGAGGCGCATCGCGAACAGACTAGTAATTCGAGCTCGA
[0171] ACAACAACAACAATAACAATAACAACAACCTCGGGATCGAGGGAAGGATTTCACA
[0172] CATGTCTATGGGCGGCCGCGATATCGTCGACGGCTCCGAATTCTCACCGAGTCCCG
[0173] TTCCGGGTAGCCAGAACGTCCCTGCTCCGGCCGTGAAGAAGATCTCGCAGTACGC
[0174] CTGCCAACGGCGGACCACTCTTAATAATTATAACCAACTGTTTACAGATGCGCTGG
[0175] AAATCTTAGCTGAAAACGCTGAGTTTCGTGAGAACGAGGGACGTTGCCTTGCTTTC
[0176] ATGCGTGCGGCATCAGTTCTGAAGAGTTTACCTTTCCCTATAACGAGTATGAAAGA CCTGGAGGGCTTACCCTGCTTAGGGGACAAAGTTAAGCGTATTATAGAAGAAATA
[0177] CTTGAGGACGGTGAGAGTTCAGAAGCGAAGGCGGTACTGAATGACGAGCGTTACA
[0178] AATCGTTCAAGCTGTTTACATCGGTCTTCGGTGTGGGATTGAAGACTGCCGAAAAG
[0179] TGGTACAGAATGGGCTTTAGAACCTTGAGCAAGATACAGAGCGACAAGTCGCTGC
[0180] GTCTTACGCAAATGCAAAAGGCAGGTTTCTTGTACTACGAGGACCTGGTCTCCTGT
[0181] GTTAATAGACCTGAGGCCGAGGCGGTCTCCATGCTTGTAAAAGAAGCAGTTGTTA
[0182] CGTTTTTGCCGGGTGCTTTGGTGACCCTTACCGGCGGCTTCCGTAGAGGGAAGATG
[0183] ACTGGGCATGATGTGGATTTCCTGATTACCTCCCCGGAGGCCGGTGAGGACGAAG
[0184] AGCAGCAGTTACTTCATAAAGTCACAGATTTCTGGAAACAGCAAGGGCTTTTACTT
[0185] TATTGCGACATCTTGGAGTCTACATTTGAAAAGTTTAAACAACCATCGCGCAAAGT
[0186] TGATGCTCTTGACCACTTTCAGAAGTGCTTCCTGATATTGAAGTTGGATCATGGTC
[0187] GGGTGCACTCAGAGAAGTCTGGCCAGCAGGAGGGCAAAGGCTGGAAGGCAATAC
[0188] GGGTTGATCTTGTGATGTGTCCCTATGACCGTAGAGCCTTTGCATTGCTTGGGTGG
[0189] ACGGGTTCCCGCCAATTTGAAAGAGATTTAAGACGTTATGCCACCCACGAAAGAA
[0190] AAATGATGCTGGATAACCACGCACTGTATGACCGCACAAAGCGGGTGTTTCTTGA
[0191] GGCGGAGTCAGAAGAAGAGATCTTTGCGCACTTAGGTTTAGACTACATCGAGCCG
[0192] TGGGAACGTAACGCCTAATAAGGATCC
[0193] SEQ ID NO: 4
[0194] CCATGGGCCATCATCATCATCATCATGCTGAGAATCTTTATTTTCAGGGCCATATG
[0195] GCGCATCATTGGGGTTACGGTAAACACAACGGTCCGGAGCATTGGCACAAAGATT
[0196] TTCCAATTGCGAAGGGCGAACGTCAAAGCCCGGTTGACATTGATACGCACACGGC
[0197] AAAGTACGACCCGAGCCTGAAACCGCTGAGCGTTTCCTATGACCAGGCTACGAGC
[0198] CTGCGTATCCTGAACAATGGCCACACCTTCAACGTGGAGTTTGATGATTCCCAAGA
[0199] TAAGGCGGTTCTGAAAGGTGGTCCGTTGGATGGCACCTACCGCCTGATCCAATTTC
[0200] ACTTTCACTGGGGTAGCCACGACGGTCAGGGCAGCGAGCATACCGTGGACAAAAA
[0201] GAAGTATGCAGCCGAACTGCACCTGGTGCATTGGAACACGAAGTACGGCGACTTC
[0202] GGTAAAGCGGTCCAGCAACCGGACGGTCTGGCTGTTCTGGGTATTTTCCTGAAGG
[0203] TCGGCAGCGCGAACCCGGGTCTGCAGAAAGTGGTTGACGTGTTGGACTCTATCAA
[0204] GACCAAAGGCAAGAGCGCGGACTTCACCAATTTCGATCCGCGTGGTCTGCTGCCG
[0205] GAGAGCCTGGATTACTGGACTTATCCGGGCAGCCTGACCACCCCGCCATTGCTGG
[0206] AGTGCGTGACCTGGATCGTCTTGAAAGAACCGATCAGCGTTAGCTCTGAACAGGT CAGCAAGTTCCGCAAGCTGAATTTCAATGGTGAGGGCGAGCCGGAAGAACCGATG
[0207] GTCGATAATTGGCGTCCTACCCAACCGCTGAAAAACCGCCAGATTAAAGCATCCTT
[0208] TAAGTAACTCGAGGATCC
[0209] Table II.
[0210] Table III.
[0211] SEQ ID NO: 5
[0212] SPSPVPGSQNVPAPAVKKISQYACQRRTTLNNYNQLFTDALEILAENAEFRENEGRCL
[0213] AFMRAASVLKSLPFPITSMKDLEGLPCLGDKVKRIIEEILEDGESSEAKAVLNDERYKS
[0214] FKLFTSVFGVGLKTAEKWYRMGFRTLSKIQSDKSLRLTQMQKAGFLYYEDLVSCVNR
[0215] PEAEAVSMLVKEAWTFLPGALVTLTGGFRRGKMTGHDVDFLITSPEAGEDEEQQLL
[0216] HKVTDFWKQQGLLLYCDILESTFEKFKQPSRKVDALDHFQKCFLILKLDHGRVHSEKS
[0217] GQQEGKGWKAIRVDLVMCPYDRRAFALLGWTGSRQFERDLRRYATHERKMMLDNH
[0218] ALYDRTKRVFLEAESEEEIFAHLGLDYIEPWERNA
[0219] SEQ ID NO: 6
[0220] LESVSMSDSEVNQEAKPEVKPEVKPETHINLKVSDGSSEIFFKIKKTTPLRRLMEAFAK
[0221] RQGKEMDSLRFLYDGIRIQADQTPEDLDMEDNDIIEAHREQTSNSSSNNNNNNNNNN
[0222] LGIEGRISHMSMGGRDIVDGSEF
[0223] SEQ ID NO: 7
[0224] MGHHHHHHLESVSMSDSEVNQEAKPEVKPEVKPETHINLKVSDGSSEIFFKIKKTTPL
[0225] RRLMEAFAKRQGKEMDSLRFLYDGIRIQADQTPEDLDMEDNDIIEAHREQTSNSSSNN NNNNNNNNLGIEGRISHMSMGGRDIVDGSEFSPSPVPGSQNVPAPAVKKISQYACQRR TTLNNYNQLFTDALEILAENAEFRENEGRCLAFMRAASVLKSLPFPITSMKDLEGLPCL GDKVKRIIEEILEDGESSEAKAVLNDERYKSFKLFTSVFGVGLKTAEKWYRMGFRTLS KIQSDKSLRLTQMQKAGFLYYEDLVSCVNRPEAEAVSMLVKEAVVTFLPGALVTLTG
[0226] GFRRGKMTGHDVDFLITSPEAGEDEEQQLLHKVTDFWKQQGLLLYCDILESTFEKFKQ PSRI<VDALDHFQI<CFLILI<LDHGRVHSEI<SGQQEGI<GWI<AIRVDLVMCPYDRRAFAL LGWTGSRQFERDLRRYATHERKMMLDNHALYDRTKRVFLEAESEEEIFAHLGLDYIE PWERNA
[0227] SEQ ID NO: 8
[0228] MGHHHHHHEVKPETHINLKVSDGSSEIFFKIKKTTPLRRLMEAFAKRQGKEMDSLRFL YDGIRIQADQTPEDLDMEDNDIIEAHREQLAENLYFQGHMAHHWGYGKHNGPEHWH KDFPIAKGERQSPVDIDTHTAKYDPSLKPLSVSYDQATSLRILNNGHTFNVEFDDSQDK AVLKGGPLDGTYRLIQFHFHWGSHDGQGSEHTVDKKKYAAELHLVHWNTKYGDFG
[0229] KAVQQPDGLAVLGIFLKVGSANPGLQKVVDVLDSIKTKGKSADFTNFDPRGLLPESLD YWTYPGSLTTPPLLECVTWIVLKEPISVSSEQVSKFRKLNFNGEGEPEEPMVDNWRPT QPLKNRQIKASFKLESVSMSDSEVNQEAKPEVKPEVKPETHINLKVSDGSSEIFFKIKKT TPLRRLMEAFAKRQGKEMDSLRFLYDGIRIQADQTPEDLDMEDNDIIEAHREQTSNSS
[0230] SNNNNNNNNNNLGIEGRISHMSMGGRDIVDGSEFSPSPVPGSQNVPAPAVKKISQYAC
[0231] QRRTTLNNYNQLFTDALEILAENAEFRENEGRCLAFMRAASVLKSLPFPITSMKDLEG
[0232] LPCLGDKVKRIIEEILEDGESSEAKAVLNDERYKSFKLFTSVFGVGLKTAEKWYRMGF
[0233] RTLSKIQSDKSLRLTQMQKAGFLYYEDLVSCVNRPEAEAVSMLVKEAWTFLPGALV TLTGGFRRGKMTGHDVDFLITSPEAGEDEEQQLLHKVTDFWKQQGLLLYCDILESTFE KFKQPSRKVDALDHFQKCFLILKLDHGRVHSEKSGQQEGKGWKAIRVDLVMCPYDR RAFALLGWTGSRQFERDLRRYATHERKMMLDNHALYDRTKRVFLEAESEEEIFAHLG
[0234] LDYIEPWERNA
[0235] SEQ ID NO: 9
[0236] MGHHHHHHEVKPETHINLKVSDGSSEIFFKIKKTTPLRRLMEAFAKRQGKEMDSLRFL YDGIRIQADQTPEDLDMEDNDIIEAHREQLAENLYFQGHMMDKDCEMKRTTLDSPLG KLELSGCEQGLHEIKLLGKGTSAADAVEVPAPAAVLGGPEPLMQATAWLNAYFHQPE AIEEFPVPALHHPVFQQESFTRQVLWKLLKVVKFGEVISYQQLAALAGNPAATAAVKT
[0237] ALSGNPVPILIPCHRVVSSSGAVGGYEGGLAVKEWLLAHEGHRLGKPGLGLESVSMS DSEVNQEAKPEVKPEVKPETHINLKVSDGSSEIFFKIKKTTPLRRLMEAFAKRQGKEM DSLRFLYDGIRIQADQTPEDLDMEDNDIIEAHREQTSNSSSNNNNNNNNNNLGIEGRIS HMSMGGRDIVDGSEFSPSPVPGSQNVPAPAVKKISQYACQRRTTLNNYNQLFTDALEI LAENAEFRENEGRCLAFMRAASVLKSLPFPITSMKDLEGLPCLGDKVKRIIEEILEDGES SEAKAVLNDERYKSFKLFTSVFGVGLKTAEKWYRMGFRTLSKIQSDKSLRLTQMQKA GFLYYEDLVSCVNRPEAEAVSMLVKEAVVTFLPGALVTLTGGFRRGKMTGHDVDFLI TSPEAGEDEEQQLLHKVTDFWKQQGLLLYCDILESTFEKFKQPSRKVDALDHFQKCFL ILKLDHGRVHSEKSGQQEGKGWKAIRVDLVMCPYDRRAFALLGWTGSRQFERDLRR YATHERKMMLDNHALYDRTKRVFLEAESEEEIFAHLGLDYIEPWERNA
[0238] SEQ ID NO: 10
[0239] MGHHHHHHAENLYFQGHMAHHWGYGI<HNGPEHWHI<DFPIAI<GERQSPVDIDTHTA KYDPSLKPLSVSYDQATSLRILNNGHTFNVEFDDSQDKAVLKGGPLDGTYRLIQFHFH WGSHDGQGSEHTVDKKKYAAELHLVHWNTKYGDFGKAVQQPDGLAVLGIFLKVGS ANPGLQKVVDVLDSIKTKGKSADFTNFDPRGLLPESLDYWTYPGSLTTPPLLECVTWI VLKEPISVSSEQVSKFRKLNFNGEGEPEEPMVDNWRPTQPLKNRQIKASFK
[0240] Expression of proteins. Proteins were expressed in E. coli BL21 (DE3) containing respective protein sequences as summarized in Table II. The bacteria were grown in 400 mL LB medium supplemented with Ampicillin (100 pg / mL) and 400 pM ZnCl2for the CAII plasmids (Seq_ID_2 and Seq_ID_4) at 37 °C until an OD600 of 0.5-0.6 was reached. The culture was induced with IPTG (0.5 mM) and the protein was expressed at 250 rpm for 18 h at 18 °C. The cells were harvested, resuspended in lysis buffer containing Tris-HCl (100 mM) at pH 8.0, Sucrose (lOOg / L), Glycerol (0.1 L / L), sodium chloride (1 M) and 1 tablet of cOmplete EDTA- free protease inhibitor cocktail (Roche), and lysed by sonification (Hielscher UP200St, 5 cycles (10 s on max. power, 150 s off)) at 4 °C. The supernatant was collected after centrifugation at max speed (Centrifuge 5418 R, Eppendorf) and the protein purified by His / Ni-beads (ROTI®Garose) and the bound protein eluted in fractions. Fractions containing the desired protein were pooled and concentrated by a centrifugal concentrator (Sartorius Vivaspin 500 or Vivaspin 2) with a suitable MW cutoff of 50’000 MWCO and a PES membrane, except for His-CAII (Seq ID lO), where 10’000 MWCO have been used instead, to obtain the desired protein in storage buffer containing 200 mM KH2PO4, 100 mM NaCl at pH 6.5. In the case of His-CAII (Seq ID lO) and His-SUMO-TdTEvo (Seq_ID_7), the protein was aliquoted and snap-frozen at -80 °C.
[0241] In the case of His- SUMO-CAD- SUMO-TdTEvo (Seq_ID_8) and His-SUMO-SNAP-SUMO- TdTEvo (Seq_ID_9), the protein was first incubated with a TEV protease according to the manufacturer’s procedure (New England Biolabs) at 4 °C and the reaction mixture was purified with His / Ni-beads (Rotigarose) and the flow-through was concentrated instead and the buffer exchanged to the storage buffer again. The proteins were then aliquoted and snap-frozen to be stored at - 80 °C.
[0242] Protein sequences are summarized in Table IV. Cleave products for TEV are summarized in Table V.
[0243] Table IV.
[0244] SEQ ID NO: 11
[0245] GHMAHHWGYGKHNGPEHWHKDFPIAKGERQSPVDIDTHTAKYDPSLKPLSVSYDQA TSLRILNNGHTFNVEFDDSQDKAVLKGGPLDGTYRLIQFHFHWGSHDGQGSEHTVDK KKYAAELHLVHWNTKYGDFGKAVQQPDGLAVLGIFLKVGSANPGLQKVVDVLDSIK TKGKSADFTNFDPRGLLPESLDYWTYPGSLTTPPLLECVTWIVLKEPISVSSEQVSKFR KLNFNGEGEPEEPMVDNWRPTQPLKNRQIKASFKLESVSMSDSEVNQEAKPEVKPEV KPETHINLKVSDGSSEIFFKIKKTTPLRRLMEAFAKRQGKEMDSLRFLYDGIRIQADQT PEDLDMEDNDDEAHREQTSNS S SNNNNNNNNNNLGIEGRI SHMSMGGRDI VDGSEF S PSPVPGSQNVPAPAVKKISQYACQRRTTLNNYNQLFTDALEILAENAEFRENEGRCLA FMRAASVLKSLPFPITSMKDLEGLPCLGDKVKRIIEEILEDGESSEAKA VENDER YKSF KLFTSVFGVGLKTAEKWYRMGFRTLSKIQSDKSLRLTQMQKAGFLYYEDLVSCVNRP
[0246] EAEAVSMLVKEAVVTFLPGALVTLTGGFRRGKMTGHDVDFLITSPEAGEDEEQQLLH KVTDFWKQQGLLLYCDILESTFEKFKQPSRKVDALDHFQKCFLILKLDHGRVHSEKSG QQEGI<GWI<AIRVDLVMCPYDRRAFALLGWTGSRQFERDLRRYATHERI<MMLDNHA LYDRTKRVFLEAESEEEIFAHLGLDYIEPWERNA
[0247] SEQ ID NO: 12
[0248] GHMMDKDCEMKRTTLDSPLGKLELSGCEQGLHEIKLLGKGTSAADAVEVPAPAAVL
[0249] GGPEPLMQATAWLNAYFHQPEAIEEFPVPALHHPVFQQESFTRQVLWKLLKVVKFGE
[0250] VISYQQLAALAGNPAATAAVKTALSGNPVPILIPCHRVVSSSGAVGGYEGGLAVKEW LLAHEGHRLGKPGLGLESVSMSDSEVNQEAKPEVKPEVKPETHINLKVSDGSSEIFFKI
[0251] KKTTPLRRLMEAFAKRQGKEMDSLRFLYDGIRIQADQTPEDLDMEDNDIIEAHREQTS
[0252] NSSSNNNNNNNNNNLGIEGRISHMSMGGRDIVDGSEFSPSPVPGSQNVPAPAVKKISQ
[0253] YACQRRTTLNNYNQLFTDALEILAENAEFRENEGRCLAFMRAASVLKSLPFPITSMKD
[0254] LEGLPCLGDKVKRIIEEILEDGESSEAKAVLNDERYKSFKLFTSVFGVGLKTAEKWYR
[0255] MGFRTLSKIQSDKSLRLTQMQKAGFLYYEDLVSCVNRPEAEAVSMLVKEAVVTFLPG
[0256] ALVTLTGGFRRGKMTGHDVDFLITSPEAGEDEEQQLLHKVTDFWKQQGLLLYCDILE
[0257] STFEKFKQPSRKVDALDHFQKCFLILKLDHGRVHSEKSGQQEGKGWKAIRVDLVMCP
[0258] YDRRAFALLGWTGSRQFERDLRRYATHERKMMLDNHALYDRTKRVFLEAESEEEIFA
[0259] HLGLDYIEPWERNA
[0260] Table V.
[0261] Deprotection of acetazolamide 1
[0262] 1 2
[0263] Procedure according to More et al. [2]
[0264] To a suspension of azetazolamide (1, 1.00 g, 4.00 mmol) in EtOH (20 mL), cone. HC1 (5 mL) was added and the mixture was refluxed for 6 h. The solvent was evaporated, sat. NaHCO3added and the mixture extracted with EtOAc. The combined organic layers were washed with brine, dried over Na2SO4, filtrated, and concentrated to yield sulfonamide 2 (0.35 g, 50 %) as a white solid.
[0265] ‘H-NMR (400 MHz, DMSO-d6): 8 (ppm) 8.04 (s, 2 H), 7.82 (s, 2 H). Synthesis of acetazolamide carboxylic acid 4
[0266] 2 3 4
[0267] Procedure according to More et al. [2]
[0268] To a solution of sulfonamide 2 (50 mg, 277 pmol, 1.0 eq.) in DMF (1.5 mL), anhydride 3 (27.7 mg, 277 pmol, 1.0 eq.) was added and the mixture was heated to 100 °C for 16 h. The solvent was evaporated to obtain the carboxylic acid 4 (65 mg, 91 %) as a white solid.
[0269] ‘H-NMR (400 MHz, DMSO-d6): 5 (ppm) 8.31 (s, 1 H), 2.80 (t, HH = 6.9 Hz, 2 H), 2.61 (t, HH = 6.8 Hz, 2 H).
[0270] Synthesis of acetazolamide azide 6
[0271] Carboxylic acid 4 (65 mg, 232 pmol, 1.0 eq.), azido linker 5 (45.5 mg, 50 mL, 208 pmol, 0.9 eq.) and EDC HC1 (45 mg, 232 pmol, 1.0 eq.) were dissolved in DMF (1.0 mL) and stirred at room temperature for 18 h. The solvent was removed and the crude purified by preparative HPLC to yield the product 6 (68 mg, 68 %) as white solid.
[0272] ‘H-NMR (400 MHz, DMSO-d6): 8 (ppm) 13.01 (s, 1 H), 8.31 (s, 2 H), 7.99 (t, HH = 5.6 Hz, 1 H), 3.62-3.57 (m, 2 H), 3.57-3.48 (m, 10 H), 3.18 (q, HH = 5.9 Hz, 2 H), 2.73 (t, HH = 6.8 Hz, 2 H).
[0273] 13C-NMR (101 MHZ, DMSO-d6): 6 (ppm) 172.3, 171.2, 164.7, 161.6, 70.3, 70.2, 70.2, 70.0, 69.7, 69.6, 50.5, 39.1, 30.7, 29.8.
[0274] Synthesis of azido-linked chlorothiazide 9 Synthesis of compound chlorothiazide precursor 8
[0275] Modified procedure according to Katzl and Ruis.[3]
[0276] Sulfonamide 7 (2.00 g, 7.00 mmol, 1.0 eq.) and anhydride 3 (6.00 g, 60.2 mmol, 8.6 eq.) were heated to 195 °C for 60 min. After the melt was cooled down, it was suspended in warm H2O (55 °C) and the precipitate filtered off. It was dried on hv to yield intermediate 8 (1.3 g, 53 %) as a white solid.
[0277] ‘H-NMR (500 MHz, DMSO-d6): 8 (ppm) 8.97 (s, 1 H), 8.41 (s, 1 H), 8.05 (s, 2 H), 3.11- 3.08 (m, 2 H), 2.91-2.89 (m, 2 H).
[0278] HRMS (ESI, m / z): CI0H8C1N3O5S2H+: 349.9594 (calc.), 381.9753 (found+MeOH).
[0279] Synthesis of chlorothiazide azide 9
[0280] 8 5 9
[0281] Intermediate 8 (15 mg, 42.9 pmol, 1.0 eq.), azido linker 5 (14 mg, 12.8 pL, 64.3 pmol, 1.5 eq.) and DIPEA (16.6 mg, 22.4 pL, 129 pmol, 3 eq.) were dissolved in DMF (500 pL) and stirred at room temperature for 18 h. The crude mixture was directly purified by preparative HPLC to yield the product 9 (14 mg, 58 %) as a white solid.
[0282] ‘H-NMR (500 MHz, DMSO-d6): 6 (ppm) 12.43 (s, 1 H), 8.25 (s, 1 H), 8.05 (t, HH = 5.6 Hz, 1 H), 7.88 (s, 2 H), 7.48 (s, 1 H), 3.60-3.58 (m, 2 H), 3.57-3.48 (m, 8 H), 3.41-3.38 (m, 4 H), 3.20 (q, HH = 5.7 Hz, 2 H), 2.83 (t, HH = 7.2 Hz, 2 H), 2.55 (t, HH = 7.2 Hz, 2 H).
[0283] 13C-NMR (126 MHz, DMSO-d6): 6 (ppm) 170.4, 160.9, 138.4, 138.3, 134.4, 125.0, 119.9, 119.0, 69.8, 69.8, 69.7, 69.6, 69.2, 69.1, 50.0, 38.7, 30.8, 30.4.
[0284] HRMS (ESI, m / z): CI8H25C1N7O8S2: 566.0973 (calc.), 566.0900 (found). Synthesis of azido-linked sulfanilamide 12
[0285] Synthesis of sulfanilamide carboxylic acid 11
[0286] Procedure according to Akocak et al. [4]
[0287] Phenyl sulfonamide 10 (5.00 g, 29.9 mmol, 1.0 eq.) and anhydride 3 (3.20 g, 32.0 mmol, 1.1 eq.) were mixed in MeCN (50 mL) and refluxed for 16 h. The reaction was allowed to cool to room temperature and the precipitate was filtered off to yield the product 11 (7.26 g, 93 %) as a white solid.
[0288] ‘H-NMR (400 MHz, DMSO-d6): 6 (ppm) 12.15 (s, 1 H), 10.30 (s, 1 H), 7.73 (s, 4 H), 7.23 (s, 2 H), 2.63-2.57 (m, 2 H), 2.56-2.52 (m, 2 H).
[0289] Synthesis of sulfanilamide azide 12
[0290] Intermediate 11 (63 mg, 232 pmol, 1.0 eq.), azidolinker 5 (45.5 mg, 50 mL, 208 pmol, 0.9 eq.) and EDC HC1 (45 mg, 232 pmol, 1.0 eq.) were dissolved in DMF (1.0 mL) and stirred at room temperature for 18 h. The solvent was removed and the crude purified by preparative HPLC to yield the product 12 (49 mg, 50 %) as white solid.
[0291] ‘H-NMR (400 MHz, DMSO-d6): 8 (ppm) 10.28 (s, 1 H), 7.93 (s, 1 H), 7.74 (s, 4 H), 7.23 (s, 2 H), 3.67-3.46 (m, 10 H), 3.44-3.31 (m, 4 H), 3.19 (q,3JHH = 5.9 Hz, 2 H), 2.58 (t,3JHH = 7.0 Hz, 2 H), 2.42 (t,3JHH = 7.0 Hz, 2 H).
[0292] 13C-NMR (101 MHz, DMSO-d6): 6 (ppm) 171.2, 171.1, 142.3, 138.0, 126.7, 118.4, 69.8, 69.8, 69.7, 69.6, 69.3, 69.1, 50.0, 38.6, 31.7, 30.0. Synthesis of small molecule nucleic acid conjugates. DNA with a 5’ DBCO modification (30 pL, 100 pM in water, 1.0 eq, Microsynth) was incubated with an azi do-functionalized small molecule (3 pL, 10 mM in DMSO, 10 eq) and MOPS buffer (2 pL, 50 mM MOPS, 500 mM NaCl, pH 8.2) for 20 h at room temperature. Afterwards, 3.5 pL (10 vol-%) of a sodium acetate buffer (3.0 M NaOAc, pH 5.2) and 105 pL ethanol were added and the mixture was incubated on ice for 2 h. The mixture was then centrifuged (4 °C, max speed, Centrifuge 5418 R, Eppendorf) and the supernatant was discarded. The pellet was washed twice with 150 pL ethanol containing 30 vol-% water. The supernatant was discarded and the pellet air-dried by leaving their vial open for 20 min in a ventilated fume hood to minimize dust. The pellet was then redissolved in water. The sequences of the nucleic acids are summarized in Table VI and the small molecule-nucleic acid conjugates in Table VII.
[0293] Table VI.
[0294] SEQ ID NO: 13
[0295] [DBCO] GGA GCT TGT ATA TGC TAA CAC TTA CGT AAT TCA CAC ACG TCC
[0296] SEQ ID NO: 14
[0297] [DBCO] GGA GCT TGT ATA TGC TCA CTG GTT ACC AAT TCA CAC ACG TCC
[0298] SEQ ID NO: 15
[0299] [DBCO] GGA GCT TGT ATA TGC TCA GTA AGT GAA AAT TCA CAC ACG TCC
[0300] Table VII.
[0301] Extension of DNA conjugates. The DNA conjugates (20 pL, 50 pM in water; Conj_No_l, Conj_No_2 or Conj_No_3, 5’ DBCO-modified SEQ ID NO: 16 (Microsynth) or the 5’ aminohexyl-modified SEQ ID NO: 17 (Microsynth)) were mixed with the splint SEQ ID NO: 10 (20 pL, 100 pM in water, 2.0 eq) and the 5’ phoshorylated extender sequence SEQ ID NO: 11 (30 pL, 100 pM in water, 3.0 eq, N annotate degenerate nucleotides) in ligation buffer (10 pL, 10X, New England Biolabs, without ATP) and water (13.7 pL) were heated to 75 °C and let cool down to room temperature over 1 h. Then, ATP (10 pL, 10 mM in water, 100 eq) was added as well as T4 DNA Ligase (1.25 pL, 500 U, 1.2 U / pmol, New England Biolabs) and incubated for 3 h at room temperature. Then, the solvent was removed and the residue redissolved in water and separated on HPLC (Separation on a Agilent 1260 Infinity II equipped with a Concise RiboSep RNA column (PS / DVB-C18, non-porous, 7.8 x 50 mm) with a gradient from 100 mM tri ethylammonium acetate in water, pH 7.2, to acetonitrile with 1 mL / min flow rate at a temperature of 60 °C (gradient: 5-8 % for 1 min to 10.8 % in 9 min to 25 % in 7 min to 99 % in 1 min for 4 min to 5 % in 0.1 min)). The lyophylized products were dissolved in 50 pL water. The sequences for the starting materials are summarized in Table VIII and VIII. The yields and and products are summarized in Table IX.
[0302] Table VIII.
[0303] SEQ ID NO: 16
[0304] [DBCO] GGA GCT TGT ATA TGC TCC AGG AGC TAT AAT TCA CAC ACG TCC
[0305] SEQ ID NO: 17
[0306] [Amine] GGA GCT TGT ATA TGC TAA C AA CGG TTG AAT TCA CAC ACG TCC
[0307] SEQ ID NO: 18
[0308] CGA ATG ATG CGG ACG TGT GT
[0309] SEQ ID NO: 19 pGCA TCA TTC GGT NNN NNN NNN NNN NNN NNN NNT AGA TCG GAA GAG CGT CGT GT
[0310] Table IX.
[0311] Proximity-induced extension of a small-molecule DNA conjugate with a TdT fusion protein. One of Conj_ID_No_4, Conj_ID_No_5, Conj_ID_No_6, or Conj_ID_No_8 (4 pL, 1 pM, 4 pmol) was mixed with sheared salmon-sperm DNA (8 pL, 1 mg / mL), dATP (8 pL, 10 mM) and CAII-SUMO-TdTEvo (Protein SEQ ID NO: 7, 4 pL, 1 pM), TdT Buffer (80 pL, 10X, 500 mM KOAc, 200 mM TrisOAc, 100 mM MgOAc2, pH 7.9) and filled up to a volume of 800 pL. The mixture was incubated at 37 °C for 30 min and subsequently deactivated for 10 min at 98 °C.
[0312] Poly-dT25 pull-down. Poly-dT2s magnetic beads (5 pL per pmol of initial compound) were washed with pull-down buffer (20 pL, 20 mM Tris pH 7.5, 1 mM EDTA, 0.01% vol-% Tween 20) on a magnetic rack and resuspended in the same buffer (10 pL). The deactivated sample of the proximity-induced extension was supplemented with Tween-20 (8 pL, 1 vol-%) and the resuspended beads were added. The sample was incubated for 30 min at room temperature and the beads were pulled down on a magnetic rack and washes three times with the pull-down buffer (100 pL). For elution, the beads were suspended in 40 pL pull-down buffer and heated to 98 °C for 10 min and immediatly placed on the magnetic rack where the supernatant was retrieved.
[0313] UREA-PAGE analysis of proximity-extended small molecule-DNA conjugates. A 1.5 mm gel was prepared by mixing acrylamide / bisacrylamide (1.25 mL, 40% 19: 1, Roth), with urea (4.2 g, 7 M final concentration), TBE (1 mL, 10X, I M Tris, 0.9 M boric acid, 0.01 M EDTA). The volume was adjusted to 10 mL, and ammonium persulfate (10 pL, 25%) and 7V- tetramethyl ethylenediamine (10 pL) were added. 3 / 4thof the elute from the step before were mixed with formamide loading dye (10 pL, 2X, 95 vol-% formamide, 5 mM EDTA, 25 mg / L bromophenol blue, pH 8.0) and heat-denatured at 98 °C for 5 min before loading on the prepared gel. The gel was run in TBE buffer (IX) in the mini-PROTEAN system (BioRad) and run at constant 120 V until the blue band reached the bottom of the gel. Bands were visualized by incubating the gel with SYBR gold (Thermo Fisher) according to the manufacturer’s instructions. qPCR analysis of proximity-extended small molecule-DNA conjugates. l / 4thof the elute was diluted 1 / 1000 and 4 pL of that dilution was transferred to a 0.2 mL PCR tube and the specific primers (1 pL of the primer pair, 5 pM each) to a for each conjugate were added (see Table X for a list of sequences and Table XI for the assignment to each conjugate) as well as 5 pL of the PowerUP SYBR Green Master Mix (ThermoFisher). qPCR was measured on QuantStudio 1 with their standard protocol (Program: 50 °C (2 min) - 95 °C (2 min) - (95 °C (15 s) - 57 °C (15 s) - 72 °C (30 s)) 40 - 95 °C (15 s) - 60 °C (1 min) - 95 °C (15 s)). The ACT values were converted to concentrations by measuring a standard curve from 1 nM to 10 fM of each conjugate. qPCR was performed thrice for each sample.
[0314] Table X.
[0315] SEQ ID NO: 20
[0316] CAC GAC GCT CTT CCG
[0317] SEQ ID NO: 21
[0318] GCT TGT ATA TGC TAA CAC TTA CGT
[0319] SEQ ID NO: 22
[0320] GCT TGT ATA TGC TCA GTA AGT GAA
[0321] SEQ ID NO: 23
[0322] TGT ATA TGC TCA CTG GTT ACC
[0323] SEQ ID NO: 24
[0324] CTT GTA TAT GCT CCA GGA GCT AT
[0325] SEQ ID NO: 25
[0326] TTG TAT ATG CTA ACA ACG GTT G
[0327] Table XI.
[0328] Proximity-induced extension of a single-stranded DNA encoded library with a TdT fusion protein. A single-stranded DNA encoded library which was synthesised according to published procedures [5], with a intact, free 3’ hydroxy end (1.0 pL, 1.0 pM, l.O pmol, or 3.1e6 molecules per library member) was added to a mixture consisting of sheared salmonsperm DNA (2 pL, 1 mg / mL), dATP (2 pL, 10 mM) and TdT Buffer (20 pL, 10X, 500 mM KOAc, 200 mM TrisO Ac, 100 mM MgOAc2, pH 7.9) and the volume was adjusted to 199 pL. Then, TdT fusion (SEQ ID NO: 3, SEQ ID NO: 11, or SEQ ID NO: 12; 1 pL, 1 pM) was added and the mixture was incubated at 37 °C for 30 min and subsequently deactivated for 10 min at 98 °C. The pull-down with dT25beads was performed as described above. For double-stranded, blunt-ended DNA encoded libraries synthesised according to published methods [5] with a free 3’ hydroxy group and a free 5’ hydroxy group has been used, and the incubation time was prolonged to 1 h instead.
[0329] Affinity-enrichment of DNA encoded libraries. Nickel-NTA beads (20 pL, HisPur Ni-NTA magnetic beads, ThermoFisher) were washed three times with selection buffer (120 pL, Dulbecco’s phosphate buffered saline with 0.01 vol-% Tween-20, 10 pg / mL sssDNA, 10 mM imidazole) and the supernatant was removed. The beads were then incubated with protein (100 pL, 37 pM, Prot SEQ ID NO: 10) or without protein for 30 min at 4 °C. To this volume, a DNA encoded library (1 pL, 1 pM, 1 pmol, single or double- stranded) was added as well as selection buffer (19 pL) and incubated for 30 min at 25 °C. Subsequently, the beads were washed five times with selection buffer (200 pL). The supernatant was removed, and the beads were resuspended in pull-down buffer (50 pL, 20 mM Tris-HCl pH 7.5, 1 mM EDTA, 0.01 vol-% Tween-20) and incubated for 10 min at 98 °C.
[0330] NGS library preparation. Based on the concentration of the samples after their selection which was determined by qPCR (1 pL sample, 2 pL water, 1 pL forward primer SEQ ID NO: 25, 1 pL reverse primer SEQ ID NO: 26 (5 pM each)), 2 pL of each selection were used in a 50 pL PCR reaction with 10-22 cycles to stay within the linear amplification range (Phusion DNA polymerase, according to manufactures protocol (NEB), 98 °C for 30 sec, then for the determined amount of cycles 98 °C (20 s), 69 °C (20 s), 72 °C (20 s), ending with 72 °C (300 s) and 12 °C (inf)). The PCR was cleaned-up with a PCR clean-up kit (Merchary & Nagel) according to their standard PCR cleanup protocol, except that the NTI buffer was diluted with 5 volumes of water to adjust the cut-off and that the washing was done twice. The cleaned-up PCR amplicons were eluted twice in 25 pL. 2 pL of these amplicons were used in a second PCR using indexed sequencing primers for the Illumina NGS platform (NEB, NebNext sets 1 and 2) using 1 pL of the forward primer and 1 pL of the reverse primer. Phusion DNA polymerase was used in a 100 pL reaction (98 °C (30 s), [98 °C (20 s), 69 °C (20 s), 72 °C (20 s)] x 15, 72 °C (300 s), 12 °C (inf)). For each sample, a individual combination of indices were used (manufacturer’s protocol). All individual PCR were pooled for equal concentration (concentration was measured by visualizing the bands on a Urea-PAGE gel with SYBR gold staining, but other methods are usable, too) and send for Illumina sequencing.
[0331] Data evaluation. Next generation sequencing data were analyzed with DECL-Gen [5c] and individual scripts. In short, for each read, the unique molecule identifier (UMI) and the barcode encoding the compound (codon-combination) were counted. For each UMI, only the most frequent codon-combination was counted. The counts of all compounds of a selection with a protein of interest where then compared to the counts of all compounds of a selection with the matrix (SUMO-TdT in the case for TdT fusion proteins as described above, empty beads in the case for free protein in the case of traditional affinity selections) according to the formula log2((x) / (y)), where x is average count (normalized to 1 million reads) of the sample and y the average counts (normalized to 1 million reads) in the control selection. The resulting number called log-Fold enrichment was shifted to have a median of 1, and only compounds that hat a log-Fold larger than 3 times the standard deviation of the log-Fold were selected as potential binders.
[0332] Table XII.
[0333] SEQ ID NO: 26
[0334] TGA CTG GAG TTC AGA CGT GTG CTC TTC CGA TCT GGA GCT TGT ATA TGC T
[0335] SEQ ID NO: 27
[0336] CAC TCT TTC CCT ACA CGA CGC TCT TCC GAT CT
[0337] Further fusion protein examples. DNA sequences (Table XIII, SEQ ID NO:28, 30 and 42) or fragments thereof were bought (AddGene or GeneScript) and cloned into a pET19b vector by restriction and ligation. The sequence of calmodulin was derived from AddGene (182500) and encodes for human CALM1 (P0DP23, SEQ ID NO: 43). The sequence of SpyCatcher (SEQ ID NO: 46) was derived from a published plasmid (AddGene, 133449) and was synthesised BioCat. The sequence of SpyTag (SEQ ID NO: 45) was derived from Addgene (133450). E. Coli DH5a cells were transfected with the ligation product and spread on a plate with antibiotic. Single colonies were sequenced and a positive clone was grown in LB medium. The plasmid was isolated and then transfected into E. Coli BL21 cells. The cells were grown to an optical density of 0.6 at 37 °C, then IPTG was added to a concentration of 0.5 mM and the cell suspension was incubated over night at 16 °C. After cell lysis, the protein was bound to nickel beads, eluted with imizadole and the buffer was exchanged with spin columns. Were necessary, the first SUMO tag and the his tag were cleaved with a TEV protease and the protease was removed by nickel beads.
[0338] Conjugation of Calmodulin-Peptide onto DNA. The calmodulin-binding peptide (SEQ ID NO:33) was synthesised by solid-phase peptide-synthesis and capped with an N-terminal 2- azidoacetic acic. The peptide was then incubated for 60 minutes with a DBCO-DNA (SEQ ID NO: 13) and purified by HPLC using a C4 column by eluting with ammonium acetate (50 m , pH 7) / acetonitrile gradient from 1% to 99%. The pure fraction was lyophilised and pure product (Calmodulin-peptide DNA conjugate) was obtained.
[0339] Activity test of Calmodulin-TdT. Calmodulin-TdT (SEQ ID NO: 29) fusion protein comprising calmodulin (SEQ ID NO: 43), a linker (SEQ ID NO: 44) and TdTevo (SEQ ID NO: 5) was incubated with the Calmodulin-peptide DNA conjugate obtained at standard conditions. Polydenalyted products were pulled-down with poly-T beads as described above. A gel (11% Urea-PAGE) was run to determine the activity (Figure 6A) and the selectivity (Figure 6B) of the calmodulin-TdT.
[0340] Table XIII.
[0341] SEQ ID NO: 28
[0342] ATGGGCCATCATCATCATCATCATCTCGAGATGGCTGATCAGCTGACTGAAGAGC
[0343] AGATCGCAGAATTCAAAGAAGCTTTCTCCCTATTTGACAAGGACGGGGATGGGAC
[0344] AATAACAACCAAGGAGCTGGGGACGGTGATGCGGTCTCTGGGGCAGAACCCCACA
[0345] GAAGCAGAGCTGCAGGACATGATCAATGAAGTAGATGCCGACGGTAATGGCACAA
[0346] TCGACTTCCCTGAATTCCTGACAATGATGGCAAGAAAAATGAAAGACACAGACAG
[0347] TGAAGAAGAAATTAGAGAAGCGTTCCGTGTGTTTGATAAGGATGGCAATGGCTAC
[0348] ATCAGTGCAGCAGAGCTTCGCCACGTGATGACAAACCTTGGAGAGAAGTTAACAG
[0349] ATGAAGAGGTTGATGAAATGATCAGGGAAGCAGACATCGATGGGGATGGTCAGGT
[0350] AAACTACGAAGAGTTTGTACAGATGATGACTGCAAAAACTAGTAATTCGAGCTCG
[0351] AACAACAACAACAATAACAATAACAACAACCTCGGGATCGAGAGCAGGATTTCAC
[0352] ACATGTCTATGGGCGGCCGCGATATCGTCGACGGCTCCGAATTCTCACCGAGTCCC
[0353] GTTCCGGGTAGCCAGAACGTCCCTGCTCCGGCCGTGAAGAAGATCTCGCAGTACG
[0354] CCTGCCAACGGCGGACCACTCTTAATAATTATAACCAACTGTTTACAGATGCGCTG
[0355] GAAATCTTAGCTGAAAACGCTGAGTTTCGTGAGAACGAGGGACGTTGCCTTGCTTT
[0356] CATGCGTGCGGCATCAGTTCTGAAGAGTTTACCTTTCCCTATAACGAGTATGAAAG
[0357] ACCTGGAGGGCTTACCCTGCTTAGGGGACAAAGTTAAGCGTATTATAGAAGAAAT
[0358] ACTTGAGGACGGTGAGAGTTCAGAAGCGAAGGCGGTACTGAATGACGAGCGTTAC
[0359] AAATCGTTCAAGCTGTTTACATCGGTCTTCGGTGTGGGATTGAAGACTGCCGAAAA
[0360] GTGGTACAGAATGGGCTTTAGAACCTTGAGCAAGATACAGAGCGACAAGTCGCTG
[0361] CGTCTTACGCAAATGCAAAAGGCAGGTTTCTTGTACTACGAGGACCTGGTCTCCTG
[0362] TGTTAATAGACCTGAGGCCGAGGCGGTCTCCATGCTTGTAAAAGAAGCAGTTGTT
[0363] ACGTTTTTGCCGGGTGCTTTGGTGACCCTTACCGGCGGCTTCCGTAGAGGGAAGAT
[0364] GACTGGGCATGATGTGGATTTCCTGATTACCTCCCCGGAGGCCGGTGAGGACGAA
[0365] GAGCAGCAGTTACTTCATAAAGTCACAGATTTCTGGAAACAGCAAGGGCTTTTACT
[0366] TTATTGCGACATCTTGGAGTCTACATTTGAAAAGTTTAAACAACCATCGCGCAAAG
[0367] TTGATGCTCTTGACCACTTTCAGAAGTGCTTCCTGATATTGAAGTTGGATCATGGT
[0368] CGGGTGCACTCAGAGAAGTCTGGCCAGCAGGAGGGCAAAGGCTGGAAGGCAATA CGGGTTGATCTTGTGATGTGTCCCTATGACCGTAGAGCCTTTGCATTGCTTGGGTG GACGGGTTCCCGCCAATTTGAAAGAGATTTAAGACGTTATGCCACCCACGAAAGA AAAATGATGCTGGATAACCACGCACTGTATGACCGCACAAAGCGGGTGTTTCTTG AGGCGGAGTCAGAAGAAGAGATCTTTGCGCACTTAGGTTTAGACTACATCGAGCC GTGGGAACGTAACGCCTAA
[0369] SEQ ID NO: 29
[0370] MGHHHHHHLEMADQLTEEQIAEFKEAFSLFDKDGDGTITTKELGTVMRSLGQNPTEA ELQDMINEVDADGNGTIDFPEFLTMMARKMKDTDSEEEIREAFRVFDKDGNGYISAA ELRHVMTNLGEKLTDEEVDEMIREADIDGDGQVNYEEFVQMMTAKTSNSSSNNNNN NNNNNLGIESRISHMSMGGRDIVDGSEFSPSPVPGSQNVPAPAVKKISQYACQRRTTL NNYNQLFTDALEILAENAEFRENEGRCLAFMRAASVLKSLPFPITSMKDLEGLPCLGD
[0371] KVKRIIEEILEDGESSEAKAVLNDERYKSFKLFTSVFGVGLKTAEKWYRMGFRTLSKIQ SDKSLRLTQMQKAGFLYYEDLVSCVNRPEAEAVSMLVKEAWTFLPGALVTLTGGFR RGKMTGHDVDFLITSPEAGEDEEQQLLHKVTDFWKQQGLLLYCDILESTFEKFKQPSR KVDALDHFQKCFLILKLDHGRVHSEKSGQQEGKGWKAIRVDLVMCPYDRRAFALLG WTGSRQFERDLRRYATHERKMMLDNHALYDRTKRVFLEAESEEEIFAHLGLDYIEPW
[0372] ERNA
[0373] SEQ ID NO: 30
[0374] ATGGGCCATCATCATCATCATCATGAAGTTAAACCGGAAACCCACATCAACCTGAA AGTTTCTGACGGTTCTTCTGAAATCTTCTTCAAAATCAAAAAAACCACCCCGCTGC GTCGTCTGATGGAAGCTTTCGCTAAACGTCAGGGTAAAGAAATGGACTCTCTGCG TTTCCTGTACGACGGTATCCGTATCCAGGCTGACCAGACCCCGGAAGACCTGGAC ATGGAAGACAACGACATCATCGAAGCTCACCGTGAACAGCTGGCTGAGAATCTTT
[0375] ATTTTCAGGGCCATATGGTAACCACCTTATCAGGTTTATCAGGTGAGCAAGGTCCG TCCGGTGATATGACAACTGAAGAAGATAGTGCTACCCATATTAAATTCTCAAAACG TGATGAGGACGGCCGTGAGTTAGCTGGTGCAACTATGGAGTTGCGTGATTCATCT GGTAAAACTATTAGTACATGGATTTCAGATGGACATGTGAAGGATTTCTACCTGTA TCCAGGAAAATATACATTTGTCGAAACCGCAGCACCAGACGGTTATGAGGTAGCA
[0376] ACTCCAATTGAATTTACAGTTAATGAGGACGGTCAGGTTACTGTAGATGGTGAAG CAACTGAAGGTGACGCTCATACTGGAATCCTCGAGTCAGTGAGCATGAGCGATTC CGAAGTGAACCAGGAAGCCAAACCCGAAGTCAAACCGGAGGTGAAACCAGAAAC CCACATTAATTTGAAGGTGTCGGACGGCTCATCAGAAATTTTTTTCAAAATTAAAA
[0377] AAACCACCCCGTTACGTAGGCTGATGGAAGCGTTCGCGAAACGCCAAGGGAAGGA
[0378] AATGGATAGTCTCCGGTTTTTATATGATGGCATTCGCATTCAAGCGGATCAAACGC
[0379] CAGAAGATTTAGACATGGAAGATAATGATATAATTGAGGCGCATCGCGAACAGAC
[0380] TAGTAATTCGAGCTCGAACAACAACAACAATAACAATAACAACAACCTCGGGATC
[0381] GAGGGAAGGATTTCACACATGTCTATGGGCGGCCGCGATATCGTCGACGGCTCCG
[0382] AATTCTCACCGAGTCCCGTTCCGGGTAGCCAGAACGTCCCTGCTCCGGCCGTGAAG
[0383] AAGATCTCGCAGTACGCCTGCCAACGGCGGACCACTCTTAATAATTATAACCAACT
[0384] GTTTACAGATGCGCTGGAAATCTTAGCTGAAAACGCTGAGTTTCGTGAGAACGAG
[0385] GGACGTTGCCTTGCTTTCATGCGTGCGGCATCAGTTCTGAAGAGTTTACCTTTCCC
[0386] TATAACGAGTATGAAAGACCTGGAGGGCTTACCCTGCTTAGGGGACAAAGTTAAG
[0387] CGTATTATAGAAGAAATACTTGAGGACGGTGAGAGTTCAGAAGCGAAGGCGGTAC
[0388] TGAATGACGAGCGTTACAAATCGTTCAAGCTGTTTACATCGGTCTTCGGTGTGGGA
[0389] TTGAAGACTGCCGAAAAGTGGTACAGAATGGGCTTTAGAACCTTGAGCAAGATAC
[0390] AGAGCGACAAGTCGCTGCGTCTTACGCAAATGCAAAAGGCAGGTTTCTTGTACTA
[0391] CGAGGACCTGGTCTCCTGTGTTAATAGACCTGAGGCCGAGGCGGTCTCCATGCTTG
[0392] TAAAAGAAGCAGTTGTTACGTTTTTGCCGGGTGCTTTGGTGACCCTTACCGGCGGC
[0393] TTCCGTAGAGGGAAGATGACTGGGCATGATGTGGATTTCCTGATTACCTCCCCGG
[0394] AGGCCGGTGAGGACGAAGAGCAGCAGTTACTTCATAAAGTCACAGATTTCTGGAA
[0395] ACAGCAAGGGCTTTTACTTTATTGCGACATCTTGGAGTCTACATTTGAAAAGTTTA
[0396] AACAACCATCGCGCAAAGTTGATGCTCTTGACCACTTTCAGAAGTGCTTCCTGATA
[0397] TTGAAGTTGGATCATGGTCGGGTGCACTCAGAGAAGTCTGGCCAGCAGGAGGGCA
[0398] AAGGCTGGAAGGCAATACGGGTTGATCTTGTGATGTGTCCCTATGACCGTAGAGC
[0399] CTTTGCATTGCTTGGGTGGACGGGTTCCCGCCAATTTGAAAGAGATTTAAGACGTT
[0400] ATGCCACCCACGAAAGAAAAATGATGCTGGATAACCACGCACTGTATGACCGCAC
[0401] AAAGCGGGTGTTTCTTGAGGCGGAGTCAGAAGAAGAGATCTTTGCGCACTTAGGT
[0402] TTAGACTACATCGAGCCGTGGGAACGTAACGCCTA
[0403] SEQ ID NO: 31
[0404] MGHHHHHHEVKPETHINLKVSDGSSEIFFKIKKTTPLRRLMEAFAKRQGKEMDSLRFL
[0405] YDGIRIQADQTPEDLDMEDNDIIEAHREQLAENLYFQGHMVTTLSGLSGEQGPSGDM
[0406] TTEEDSATHIKFSKRDEDGRELAGATMELRDSSGKTISTWISDGHVKDFYLYPGKYTF VETAAPDGYEVATPIEFTVNEDGQVTVDGEATEGDAHTGILESVSMSDSEVNQEAKP EVKPEVKPETHINLKVSDGSSEIFFKIKKTTPLRRLMEAFAKRQGKEMDSLRFLYDGIRI QADQTPEDLDMEDNDIIEAHREQTSNSSSNNNNNNNNNNLGIEGRISHMSMGGRDIV DGSEFSPSPVPGSQNVPAPAVKKISQYACQRRTTLNNYNQLFTDALEILAENAEFRENE GRCLAFMRAASVLKSLPFPITSMKDLEGLPCLGDKVKRIIEEILEDGESSEAKAVLNDE RYKSFKLFTSVFGVGLKTAEKWYRMGFRTLSKIQSDKSLRLTQMQKAGFLYYEDLVS CVNRPEAEAVSMLVKEAVVTFLPGALVTLTGGFRRGKMTGHDVDFLITSPEAGEDEE QQLLHKVTDFWKQQGLLLYCDILESTFEKFKQPSRKVDALDHFQKCFLILKLDHGRVH SEI<SGQQEGI<GWI<AIRVDLVMCPYDRRAFALLGWTGSRQFERDLRRYATHERI<MM LDNHALYDRTKRVFLEAESEEEIFAHLGLDYIEPWERNA
[0407] SEQ ID NO: 32
[0408] GHMVTTLSGLSGEQGPSGDMTTEEDSATHIKFSKRDEDGRELAGATMELRDSSGKTIS TWISDGHVKDFYLYPGKYTFVETAAPDGYEVATPIEFTVNEDGQVTVDGEATEGDAH TGILESVSMSDSEVNQEAKPEVKPEVKPETHINLKVSDGSSEIFFKIKKTTPLRRLMEAF AKRQGKEMDSLRFLYDGIRIQADQTPEDLDMEDNDIIEAHREQTSNSSSNNNNNNNN NNLGIEGRISHMSMGGRDIVDGSEFSPSPVPGSQNVPAPAVKKISQYACQRRTTLNNY NQLFTDALEILAENAEFRENEGRCLAFMRAASVLKSLPFPITSMKDLEGLPCLGDKVK RIIEEILEDGESSEAKAVLNDERYKSFKLFTSVFGVGLKTAEKWYRMGFRTLSKIQSDK SLRLTQMQKAGFLYYEDLVSCVNRPEAEAVSMLVKEAWTFLPGALVTLTGGFRRGK MTGHDVDFLITSPEAGEDEEQQLLHKVTDFWKQQGLLLYCDILESTFEKFKQPSRKVD ALDHFQKCFLILKLDHGRVHSEKSGQQEGKGWKAIRVDLVMCPYDRRAFALLGWTG SRQFERDLRRYATHERKMMLDNHALYDRTKRVFLEAESEEEIFAHLGLDYIEPWERN A
[0409] SEQ ID NO: 33
[0410] AAARWI<I<NFIAVSAANRFAI<IS
[0411] SEQ ID NO: 41
[0412] MGHHHHHHSSGLVPRGSRGVPHIVMVDAYKRYKGSGGTSGSGHMAHHWGYGKHN GPEHWHKDFPIAKGERQSPVDIDTHTAKYDPSLKPLSVSYDQATSLRILNNGHTFNVE FDDSQDKAVLKGGPLDGTYRLIQFHFHWGSHDGQGSEHTVDKKKYAAELHLVHWNT KYGDFGKAVQQPDGLAVLGIFLKVGSANPGLQKVVDVLDSIKTKGKSADFTNFDPRG LLPESLDYWTYPGSLTTPPLLECVTWIVLKEPISVSSEQVSKFRKLNFNGEGEPEEPMV
[0413] DNWRPTQPLKNRQIKASFK
[0414] SEQ ID NO: 42
[0415] ATGGGCCATCATCATCATCATCACAGCAGCGGCCTGGTGCCGCGCGGCAGCCGTG
[0416] GCGTGCCTCATATCGTGATGGTGGACGCCTACAAGCGTTACAAGGGTAGTGGTGG
[0417] CACTAGTGGCAGTGGCCATATGGCGCATCATTGGGGTTACGGTAAACACAACGGT
[0418] CCGGAGCATTGGCACAAAGATTTTCCAATTGCGAAGGGCGAACGTCAAAGCCCGG
[0419] TTGACATTGATACGCACACGGCAAAGTACGACCCGAGCCTGAAACCGCTGAGCGT
[0420] TTCCTATGACCAGGCTACGAGCCTGCGTATCCTGAACAATGGCCACACCTTCAACG
[0421] TGGAGTTTGATGATTCCCAAGATAAGGCGGTTCTGAAAGGTGGTCCGTTGGATGG
[0422] CACCTACCGCCTGATCCAATTTCACTTTCACTGGGGTAGCCACGACGGTCAGGGCA
[0423] GCGAGCATACCGTGGACAAAAAGAAGTATGCAGCCGAACTGCACCTGGTGCATTG
[0424] GAACACGAAGTACGGCGACTTCGGTAAAGCGGTCCAGCAACCGGACGGTCTGGCT
[0425] GTTCTGGGTATTTTCCTGAAGGTCGGCAGCGCGAACCCGGGTCTGCAGAAAGTGG
[0426] TTGACGTGTTGGACTCTATCAAGACCAAAGGCAAGAGCGCGGACTTCACCAATTT
[0427] CGATCCGCGTGGTCTGCTGCCGGAGAGCCTGGATTACTGGACTTATCCGGGCAGC
[0428] CTGACCACCCCGCCATTGCTGGAGTGCGTGACCTGGATCGTCTTGAAAGAACCGA
[0429] TCAGCGTTAGCTCTGAACAGGTCAGCAAGTTCCGCAAGCTGAATTTCAATGGTGA
[0430] GGGCGAGCCGGAAGAACCGATGGTCGATAATTGGCGTCCTACCCAACCGCTGAAA
[0431] AACCGCCAGATTAAAGCATCCTTTAAGTAA
[0432] SEQ ID NO: 43
[0433] MADQLTEEQIAEFKEAFSLFDKDGDGTITTKELGTVMRSLGQNPTEAELQDMINEVD
[0434] ADGNGTIDFPEFLTMMARI<MI<DTDSEEEIREAFRVFDI<DGNGYISAAELRHVMTNLG
[0435] EKLTDEEVDEMIREADIDGDGQVNYEEFVQMMTAK
[0436] SEQ ID NO:44
[0437] TSNS S SNNNNNNNNNNLGIESRISHMSMGGRDIVDGSEF
[0438] SEQ ID NO: 45
[0439] MGHHHHHHSSGLVPRGSRGVPHIVMVDAYKRYKGSGGTSGSGH
[0440] SEQ ID NO: 46 GHMVTTLSGLSGEQGPSGDMTTEEDSATHIKFSKRDEDGRELAGATMELRDSSG KTISTWISDGHVKDFYLYPGKYTFVETAAPDGYEVATPIEFTVNEDGQVTVDGEAT EGDAHTGILESVSMSDSEVNQEAKPEVKPEVKPETHINLKVSDGSSEIFFKIKKTTP LRRLMEAFAKRQGKEMDSLRFLYDGIRIQADQTPEDLDMEDNDIIEAHREQTSNS S SNNNNNNNNNNLGIEGRISHMSMGGRDIVDGSEF
[0441] Conjugating BnG-DNA to SNAP-TdT. DNA-SNAP-TdT was prepared by incubating SNAP-SUMO-TdTEvo (1 pL, 157 ng pL-1 in storage buffer, 2 pmol, 1 eq., SEQ ID NO: 12) together with the BnG modified DNA (2 pL, 1 pM in H2O, 2 pmol, 1 eq., SEQ ID NO:34) for 15 min at 25 °C (See Figure 7A). The coupling was freshly done before each assay and for each experiment individually (no bulk). For untemplated assays, the SNAP-SUMO-TdTEvo was incubated under the same conditions with H2O. Dissociation constant of the DNAs to SEQ ID NO: 34 were calculated with the nearest neighbor method and given in Table XIV. [BnG] designates a benzoguanine moiety, [FAM] designates a fluorescein (FAM) moiety.
[0442] SEQ ID NO: 34
[0443] [BnG] GGA GOT TGT [3d A]
[0444] SEQ ID NO: 35
[0445] [BnG] ATG ACG TGG [3d A]
[0446] SEQ ID NO: 36
[0447] [FAM] ACA GAC AAG CTC OCT GAT
[0448] SEQ ID NO: 37 [FAM] ACA TAC GAG CTC CCT GAT
[0449] SEQ ID NO: 38
[0450] [FAM] ACA GAC AAG CAC CCT GAT
[0451] SEQ ID NO: 39
[0452] [FAM] ACA AAC AAG CAC GCT GAT
[0453] SEQ ID NO: 40
[0454] [FAM] ACA AAC AAG CCC GCT GAT
[0455] SpyTag-CAII and SpyCatcher-TdT extensions. 8 pL of SpyCatcher-SUMO-TdT (6 pM, SEQ ID NO:32) and 1.8 pL of SpyTag-CAII (30 pM, SEQ ID NO:41) were mixed and incubated on ice for 20 minutes. Then, the TdT extension assay was set up was using standard conditions at 4 time larger scale (800 pL) without sssDNA to improve the gel readout. For extension, conjugates 1-3 (Table VII) were used the substrate. After inactivation of TdT, the product was enriched using dT25 beads as described above and the elutes were run on a 11% UREA-PAGE gel. Figure 6C shows the Spytag-C AIL SpyCatcher-TdT conjugate as well as the extension of the DNA small molecule conjugates.
[0456] Loading multiple TdT fusion proteins on a POL A POI (lOuL, 30uM) was incubated with BG-maleimid (lOuL, 1 mM) for 2 hours at room temperature in PBS buffer and purified by spin column. Subsequently, the protein (10 uL, 3uM) was incubated with Snap-TdT (SEQ ID NO: 12) (10 uL, 9uM) for 2 hours at room temperature to give, on average, X TdTs per POI.
[0457] B) Results
[0458] Proximity-induced extension of a small-molecule DNA conjugates. Fig. IB shows the extension of the individual small-molecule DNA conjugates with the CAILTdT fusion protein. The larger the affinity, the more extended the conjugated DNA strand gets (marked with grey bars). The darker parts of the gel especially on the top originates from either unspecific binding of sheared salmon-sperm DNA or background-extensions thereof. The conjugate bearing only an amine does not show a band for enriched and extended. qPCR analysis of proximity-extended small molecule-DNA conjugates. Fig. 2 shows the same experiments as Fig IB, but measured quantitatively by qPCR instead. Furthermore, the amount of small-molecule DNA conjugate is measured instead of the amount and the distribution of the poly- A tail. All three binders are enriched at least 30 times above the background conjugate.
[0459] Proximity-induced extension DNA bound to SNAP-conjugated benzoguanine-DNA conjugates. Fig. 7 shows a model system in which the affinity of a small molecule-DNA conjugate towards the POI is simulated with a DNA-DNA pair of different affinities. The conjugate of Snap-TdT (SEQ ID NO: 12) and BG-DNA (SEQ ID NO: 34) is mixed with the DNAs of varying affinity indicated above. The affinity towards SEQ ID NO: 34 is decreasing from SEQ ID NO: 36 to SEQ ID NO: 40). The larger the affinity, the more dATP gets incorportated onto the DNA strand and the longer the DNA tail gets. If no DNA is conjugated to SNAP-TdT, the extension of the substrate DNA (SEQ ID NO:36) is minor.
[0460] Affinity-enrichment of DNA encoded libraries. When comparing the counts of compounds of the single-stranded DNA encoded library selection against His-CAII (SEQ ID NO: 10) with the selection against empty beads, 143 out of 192 compounds containing an acetazolamide (73%) were retrieved (Fig 3A, solid arrow) and 979 out of 2606 compounds containing a phenylsulfonamide (38 %, dotted arrow). In the case of the chemically identical doublestranded DNA encoded library, 157 out of 192 acetazolamides were found (82 %, Fig 3B, solid arrow)) and 1158 out of 2606 phenylsulfonamides (44 %, dotted arrow).
[0461] Proximity-induced extension of DNA-encoded libraries. When comparing the counts of compounds of the single-stranded DNA-encoded library selected against a CAII-SUMO-TdT fusion protein (SEQ ID NO: 11) against the counts of a selection against a SUMO-TdT fusion protein (SEQ ID NO: 7), 145 out of 192 compounds containing an acetazolamide (74 %) were retrieved (Fig 4 A, solid arrow), and 2319 / 2606 compounds containing a phenylsulfonamide (89 %, Fig 4A, dotted arrow) - significantly more than found with the affinity enrichment. Moreover, we found 21 compounds containing an A-methoxyphenyl sulfonamide motif (11%, dashed arrow). When incubated with half the amount of protein, 151 acetazolamide-containing compounds were retrieved (79 %, Fig 4B) but with a stronger log-Fold values. Furthermore, 2243 phenylsulfonamides were found (86 %). The A-methoxyphenyl sulfonamide motif was found 20 times (10%). In the case of a double-stranded DNA-encoded library, 126 acetazolamides where found (66 %, Fig 5 A) and 1843 phenylsulfonamides (71 %). With half the amount of protein, 115 acetazolamides where found (60%, Fig 5B) but with stronger log- Fold enrichments, and 1419 phenyl sulfonamides (54 %).
[0462] Activity test of Calmodulin-TdT and proximity-induced extension of a binder-DNA conjugate. The calmodulin-TdT (SEQ ID NO: 29) is active (Figure 6A) and can extend a given DNA on the 3’ end with the dATP. A DNA modified with a calmodulin-peptide binder (SEQ ID NO: 33) extends in higher amounts than a DNA conjugated with only a DBCO small molecule (Figure 6B), showing the induced-proximity effect of the system.
[0463] Proximity-Induced of a conjugated CAII-SpyTag-SpyCatcher-SUMO-TdT. Figure 6C shows the formation of the conjugate of SpyTag-CAII (37 kDa) and the SpyCatcher- SUMO- TdT (92 kDa). If both are present, a new band forms at around 130 kDa as the conjugation between SpyCatcher and SpyTag. Figure 6D shows the proximity-induced extension of Conjugates 1-3 (Table VII) with SpyTag-CAII- SpyCatcher-TdT. Non-binding conjugates (SEQ ID NO: 16) does not show proximity-extension.
[0464] C) Summary
[0465] The proximity-induced poly- A tailing with a CAII-TdT fusion protein with subsequent enrichment for the extended tail performs equally well in retrieving binders with a KDin the two-digit nanomolar range, as shown in Fig. 3-5. 74-79 % of these binders were retrieved with a single-stranded DNA-encoded library, compared to 73 % retrieved with the traditional affinity enrichment approach. Surprisingly, the present invention is significantly better at retrieving micromolar binders than state-of-the art affinity enrichement methods [7], Nearly 90 % of phenylsulfonamids were found that were contained in the library compared to only 38 % with state-of-the art aflfiniy enrichment methods [7], Furthermore, with the methods of the present invention even 20-21 out of 192 (10-11%) N-methoxyphenyl sulfonamides were identified which are known to be very weak inhibitors of CAII with Ki that around one magnitude weaker than the non-methoxylated analogue [7], This motif was completely invisible in the state-of-the art aflfiniy enrichment and was only captured by the present invention. This makes the method of the present invention a perfect screening tool for proteins that are hard to target.
[0466] The present method still identifies more two thirds of the strong binders (66 %, 82 % in affinity enrichment) in double-stranded DNA-encoded library. Of the weak binders, 71 % were identified. Albeit this is smaller than obtained with the method using single-stranded libraries, it is still significantly more than with the state-of-the art affinity enrichement methods [7](44 %).
[0467] Furthermore, the present method works with a peptide-protein interaction on subnanomolar scale, as seen with the proximity-induced extension of a Calmodulin-Binding-Peptide-DNA conjugate and Calmodulin-TdT. The proximity-induced extension also works if an artificial conjugate using the SpyTag- SpyCatcher technology is created between the protein of interest and TdT.
[0468] List of references
[0469] [1] Barthel, S., Palluk, S., Hillson, N. J., Keasling, J. D. & Arlow, D. H. “Enhancing Terminal Deoxynucleotidyl Transferase Activity on Substrates with 3' Terminal Structures for Enzymatic De Novo DNA Synthesis”. Genes 11, 102 (2020).
[0470] [2] More, K. N. et al. “Acetazolamide-based [18F]-PET tracer: In vivo validation of carbonic anhydrase IX as a sole target for imaging of CA-IX expressing hypoxic solid tumors”. Bioorganic & Medicinal Chemistry Letters 28, 915-921 (2018).
[0471] [3] Katzl, K. & Ruis, H. «Uber tricyclische 1,2,4-Benzothiadiazine». Monatshefte fur Chemie 96, 1603-1610 (1965).
[0472] [4] Akocak, 2016, “PEGylated Bis-Sulfonamide Carbonic Anhydrase Inhibitors Can Efficiently Control the Growth of Several Carbonic Anhydrase IX-Expressing Carcinomas”. Journal of Medicinal Chemistry 59, 5077-5088 (2016).
[0473] [5] (a) Stress, C., Sauter, B., Schneider, L., Sharpe, T. & Gillingham, D. “A DNA-encoded chemical library incorporating elements of natural macrocycles”. Angew. Chem. Int. Ed. 58, 9570-9574 (2019). (b) Bassi, G. et al. “A Single- Stranded DNA-Encoded Chemical Library Based on a Stereoisomeric Scaffold Enables Ligand Discovery by Modular Assembly of Building Blocks”. Advanced Science 7, 2001970 (2020). (c) Sauter, B. Doctoral Thesis, “Applications of Next Generation Sequencing in the Field of Chemical Biology”. (University of Basel, 2020) doi: 10.5451 / UNIBAS-EP88004.
[0474] [6] (a) Wichert, M. et al. «Dual-display of small molecules enables the discovery of ligand pairs and facilitates affinity maturation.” Nature Chemistry 7, 241-249 (2015). (b) Kazmi erski, W. M. et al. “DNA-Encoded Library Technology-Based Discovery, Lead Optimization, and Prodrug Strategy toward Structurally Unique Indoleamine 2, 3 -Dioxygenase- 1 (IDO1) Inhibitors.” Journal of Medicinal Chemistry 63, 3552-3562 (2020). (c) Stress, C., Sauter, B., Schneider, L., Sharpe, T. & Gillingham, D. “A DNA-encoded chemical library incorporating elements of natural macrocycles.” Angew. Chem. Int. Ed. 58, 9570-9574 (2019).
[0475] [7] (a) Brigand, F., Pierattelli, R., Scozzafava, A. & Supuran, C. Carbonic anhydrase inhibitors. Part 37. Novel classes of isozyme I and II inhibitors and their mechanism of action. Kinetic and spectroscopic investigations on native and cob alt- substituted enzymes. European Journal of Medicinal Chemistry 31, 1001-1010 (1996). (b) Fiore, A. D., Maresca, A., Alterio, V., Supuran, C. T. & Simone, G. D. Carbonic anhydrase inhibitors: X-ray crystallographic studies for the binding of N- substituted benzenesulfonamides to human isoform II. Chemical Communications 47, 11636 (2011).
[0476] [8] Keeble A. H. et al. “Approaching infinite affinity through engineering of peptide-protein interaction” PNAS, 116, 26523-26533 (2019).
Claims
Claims1. A fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof, with the proviso that the fusion protein is not a fusion protein comprising the terminal deoxynucleotidyl transferase (TDT) and Cas9.
2. The fusion protein of claim 1, wherein the fusion protein comprises a linker between the DNA polymerase with terminal transferase activity or a fragment therof and the protein of interest (POI) or a fragment therof.
3. The fusion protein of claim 1 or 2, wherein the DNA polymerase with terminal transferase activity or a fragment therof is a terminal deoxynucleotidyl transferase (TDT) or a fragment thereof.
4. The fusion protein of claim 3, wherein the terminal deoxynucleotidyl transferase (TDT) or a fragment thereof comprises the sequence as shown in SEQ ID NO: 5.
5. The fusion protein of any one of claims 1 to 4, wherein the POI or a fragment therof is selected from the group consisting of an enzyme or a fragment thereof and a protein or a fragment thereof targeting or involved in cellular proliferation.
6. The fusion protein of any one of claims 1 to 4, wherein the fusion protein comprises a sequence selected from the group as shown in SEQ ID NOs: 7-12.
7. The fusion protein of any one of claims 1 to 6, wherein the fusion protein is bound to a conjugate compound comprising a binder of the POI or a fragment thereof and a nucleic acid moiety.
8. The fusion protein of claim 7, wherein the nucleic acid moiety is ssDNA or dsDNA.
9. The fusion protein of any one of claims 1 to 8, wherein the fusion protein comprises a DNA polymerase with terminal transferase activity or a fragment thereof and calmodulin or a fragment thereof.
10. The fusion protein of any one of claims 1 to 9, wherein the fusion protein comprises a DNA polymerase with terminal transferase activity or a fragment thereof, and two or more copies of the protein of interest (POI) or of a fragment thereof.
11. The fusion protein of any one of claims 1 to 9, wherein the fusion protein comprises two or more copies of the DNA polymerase with terminal transferase activity or of a fragment thereof, and a protein of interest (POI) or a fragment therof.
12. The fusion protein of any one of claims 1 to 9, wherein the fusion protein comprises two or more copies of the DNA polymerase with terminal transferase activity or of a fragment thereof, and two or more copies of the protein of interest (POI) or of a fragment thereof.
13. The fusion protein of any one of claims 1 to 12, wherein the fusion protein comprises one or more linkers between the DNA polymerase with terminal transferase activity or a fragment therof and the protein of interest (POI) or a fragment therof, wherein the linker is a chemical entity, and wherein the one or more linkers are fused each to an amino acid of the protein of interest (POI) or of a fragment therof different from the N-terminal and / or C-terminal amino acid of the protein of interest (POI) or of a fragment therof, on one part of the linker, and wherein the same one or more linkers are fused each to a DNA polymerase with terminal transferase activity or to a fragment therof on another part of the linker.
14. A fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and calmodulin or a fragment thereof, wherein the fusion protein is bound to a further fusion protein comprising a calmodulin-binding peptide and a POI or a fragment thereof.
15. A method for screening a molecule for binding to a protein of interest (POI) or a fragment therof, the method comprising the steps of:a) providing i) an incubation medium comprising a deoxynucleotide triphosphate (dNTP), ii) a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof and iii) a conjugate compound comprising a molecule and a nucleic acid moiety; and incubating the fusion protein and the conjugate compound in the incubation medium; b) separating conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended from conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is not extended; c) amplifying and analyzing the extended nucleic acid moiety of the conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended obtained in step b); d) correlating the extended nucleic acid moiety analyzed in step c) with the corresponding molecule of the conjugate compound used in step a), wherein the molecules correlated in step d) are selected as binder of the POI.
16. The method for screening a molecule of claim 15, wherein in step a) the DNA polymerase with terminal transferase activity or a fragment thereof is inactivated prior to performing step b).
17. The method for screening a molecule of claim 15 or 16, wherein step b) comprises adding a solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound to the incubation medium and separating the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound hybridized to the extended nucleic acid moiety of the conjugate compound from the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound not hybridized to the extended nucleic acid moiety of the conjugate compound.
18. The method for screening a molecule of claim 17, wherein the conjugate compounds with the extended nucleic acid moiety are eluted from the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound.
19. The method for screening a molecule of claim 18, wherein a washing step is performed after the conjugate compounds with the extended nucleic acid moiety are eluted from the solid phase comprising deoxynucleotides complementary to the extended nucleic acid moiety of the conjugate compound.
20. The method for screening a molecule of any one of claims 15 to 19, wherein the extended nucleic acid moiety is amplified and sequenced in step c).
21. The method for screening a molecule of any one of claims 15 to 20, wherein the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof is the fusion protein of anyone of claims 1-8.
22. The method for screening a molecule of any one of claims 15 to 20, wherein the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof is the fusion protein of anyone of claims 1-14.
23. A method for screening a DNA-encoded library of molecules for binding to a protein of interest (POI) or a fragment therof, the method comprising the steps of: a) providing i) an incubation medium comprising a deoxynucleotide triphosphate (dNTP), ii) a fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof and iii) a DNA-encoded library of chemical molecules comprising multiple instances of one sole molecule of the library, each instance being covalently linked to a nucleic acid moiety to form conjugate compoundscomprising a molecule and a nucleic acid moiety; and incubating the fusion protein and the conjugate compounds; b) separating conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended from conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is not extended; c) amplifying and analyzing the extended nucleic acid moiety of the conjugate compounds in which the nucleic acid moiety of the conjugate compound as provided in step a) is extended obtained in step b); d) correlating the extended nucleic acid moiety analyzed in step c) with the corresponding multiple instances of one sole chemical molecule of the library used in step a), wherein the multiple instances of one sole chemical molecule of the library so correlated in step d) are selected as binder of the POI.
24. The method for screening a DNA-encoded library of molecules of claim 23, wherein the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof is the fusion protein of anyone of claims 1-8.
25. The method for screening a DNA-encoded library of molecules of claim 23, wherein the fusion protein comprising a DNA polymerase with terminal transferase activity or a fragment thereof and a protein of interest (POI) or a fragment therof is the fusion protein of anyone of claims 1-14.