RNAization

By attaching a 5'-nicotinamide nucleobase dinucleotide cap nucleic acid sequence to proteins using ADP-ribosyltransferase, the method addresses the challenge of understanding and manipulating protein post-translational modifications, enhancing protein functionality and stability.

JP7894095B2Active Publication Date: 2026-07-23MAX PLANCK GESELLSCHAFT ZUR FOERDERUNG DER WISSENSCHAFTEN EV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
MAX PLANCK GESELLSCHAFT ZUR FOERDERUNG DER WISSENSCHAFTEN EV
Filing Date
2022-04-21
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Current technologies lack the ability to effectively identify and understand post-translational modifications (PTMs) of proteins, which are crucial for cell biology and disease treatment, as they play a significant role in various biological processes and disruptions can lead to diseases.

Method used

A method is developed to attach a 5'-nicotinamide nucleobase dinucleotide (NND) cap nucleic acid sequence to a fusion protein or complex using ADP-ribosyltransferase (ART), leveraging a specific arginine recognition motif to covalently bind the NND cap nucleic acid sequence to the protein under physiological conditions.

Benefits of technology

This method enables the novel PTM 'RNAization' of proteins, allowing for the attachment of NND cap nucleic acid sequences, enhancing protein functionality and stability, and providing a new approach for studying and manipulating protein modifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007894095000014
    Figure 0007894095000014
  • Figure 0007894095000015
    Figure 0007894095000015
  • Figure 0007894095000016
    Figure 0007894095000016
Patent Text Reader

Abstract

The present invention relates to a method for attaching a 5'-nicotinamide nucleobase dinucleotide (NND) cap nucleic acid sequence to a fusion protein or complex, the method comprising the steps of: (a) contacting a heterologous fusion protein or complex comprising (i) a poly(peptide) of interest fused to a tag, with a 5'-NAD cap nucleic acid sequence and an ADP-ribosyltransferase (ART), wherein the protein is complexed with the tag under physiological conditions, and the 5'-NND cap nucleic acid sequence is covalently attached to the tag, the tag comprising a recognition motif for the ART, preferably (i) the amino acid motif DV R (ii) SEQ ID NO:1 or a sequence at least 80% identical thereto, assuming that the underlined Arg in PVRD (SEQ ID NO:7) is conserved and preferably SEQ ID NO:7 is conserved; R (iii) SEQ ID NO:2 or a sequence at least 80% identical thereto, assuming that the underlined Arg in PVRD (SEQ ID NO:7) is conserved and preferably SEQ ID NO:7 is conserved; R (iv) SEQ ID NO:3 or a sequence at least 80% identical thereto, assuming that the underlined Arg in PVRD (SEQ ID NO:7) is conserved and preferably SEQ ID NO:7 is conserved; (iv) the amino acid motif LADGVEGYLRASEASRD R (v) SEQ ID NO:4 or a sequence at least 80% identical thereto, assuming that the underlined Arg in VE (SEQ ID NO:8) is conserved and preferably SEQ ID NO:8 is conserved; (v) the amino acid motif LADGVEGYLRASEASRD R SEQ ID NO:5 or a sequence at least 80% identical thereto, assuming that the underlined Arg in VE (SEQ ID NO:8) is conserved and preferably SEQ ID NO:8 is conserved; or (vi) the amino acid motif LADGVEGYLRASEASRD RAssuming that the underlined Arg in VE (SEQ ID NO:8) is conserved and preferably SEQ ID NO:8 is conserved, it comprises or consists of SEQ ID NO:6 or a sequence at least 80% identical thereto.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for attaching a 5'-nicotinamide nucleobase dinucleotide (NND) cap nucleic acid sequence to a fusion protein or complex, the method comprising the steps of (a) contacting a heterofusion protein containing (i) a target poly(peptide) to be fused to a tag, or (ii) a complex, with the 5'-NND cap nucleic acid sequence and ADP-ribosyltransferase (ART) under conditions in which the 5'-NND cap nucleic acid sequence is covalently bound to the tag, wherein the protein is complexed with the tag under physiological conditions, and the tag contains an ART recognition motif, preferably (i) an amino acid motif DV R (ii) A sequence that is at least 80% identical to sequence number 1, assuming that the underlined Arg in PVRD (sequence number 7) is preserved and preferably sequence number 7 is preserved; (ii) Amino acid motif DV R Assuming that the underlined Arg in PVRD (SEQ ID NO: 7) is preserved and preferably SEQ ID NO: 7 is preserved, a sequence that is at least 80% identical to SEQ ID NO: 2; (iii) Amino acid motif DV R Assuming that the underlined Arg in PVRD (SEQ ID NO: 7) is preserved, and preferably SEQ ID NO: 7 is preserved, then a sequence that is at least 80% identical to SEQ ID NO: 3; (iv) Amino acid motif LADGVEGYLRASEASRD R Assuming that the underlined Arg in VE (SEQ ID NO: 8) is preserved, and preferably SEQ ID NO: 8 is preserved, then the sequence is at least 80% identical to SEQ ID NO: 4; (v) Amino acid motif LADGVEGYLRASEASRD R Assuming that the underlined Arg in VE (SEQ ID NO: 8) is preserved and preferably SEQ ID NO: 8 is preserved, then a sequence that is at least 80% identical to SEQ ID NO: 5; or (vi) amino acid motif LADGVEGYLRASEASRD RAssuming that the underlined Arg in VE (sequence number: 8) is preserved, and preferably sequence number: 8 is preserved, then the sequence contains or consists of sequence number: 6 or a sequence that is at least 80% identical to it.

[0002] In this specification, several documents, including patent applications and manufacturers' manuals, are referenced. While the disclosures in these documents are not considered relevant to the patentability of the present invention, they are incorporated herein by reference in their entirety. More specifically, all referenced documents are incorporated by reference to the same extent as each individual document is specifically and individually indicated as being incorporated by reference. [Background technology]

[0003] Post-translational modifications (PTMs) are biochemical alterations that occur to one or more amino acids on a protein after it has been translated by ribosomes. PTMs increase the functional diversity of the proteome through the covalent attachment of functional groups or proteins, proteolytic cleavage of regulatory subunits, or degradation of the entire protein. These modifications include phosphorylation, glycosylation, ubiquitination, nitrosylation, methylation, acetylation, lipidation, and proteolysis, and affect almost every aspect of normal cell biology and pathogenesis. There are over 400 different types of PTMs that affect many aspects of protein function. These modifications are crucial molecular regulatory mechanisms for controlling diverse cellular processes. These processes have significant impacts on protein structure and function. Disruptions in PTMs can lead to abnormalities in biological processes essential for life and, consequently, various diseases.

[0004] Therefore, there is a need to identify and understand further PTMs. This is extremely important in cell biology and in the study of disease treatment and prevention. This need is addressed by the present invention.

[0005] ADP-ribosyltransferase (ART) catalyzes the transfer of one or more ADP-ribose (ADPr) units from nicotinamide adenine dinucleotide (NAD) to a target protein. 1 In bacteria and archaea, they act as toxins and are involved in host defense or drug resistance mechanisms. 8 In eukaryotes, they play a role in distinct processes ranging from DNA damage repair to macrophage activation and stress response. 9 The virus uses ART as a weapon to reprogram the host's gene expression system. 7 Mechanistically, the nucleophilic groups of the target protein (mostly Arg, Glu, Asp, Ser, Cys) attack the glycosidic carbon atoms in the nicotinamide riboside moiety of NAD, forming covalent bonds as N-, O-, or S-glycosides (Figure 1a). 1 .

[0006] First, it is experimentally shown that bacteriophage T4 ART accepts not only NAD but also NND-RNA as a substrate, thereby covalently binding the entire RNA chain to the acceptor protein in the "RNAylation" reaction. As demonstrated using ART ModB, ART can effectively RNA-lyse its host protein target, ribosomal protein S1, at an arginine residue, and strongly prefers NAD-RNA to NAD. A single arginine mutation at position 139 completely disrupts ADP-ribosylation and RNAylation. ART is also further shown to work with NGD (5'-nicotinamide guanine dinucleotide)-RNA, NCD (5'-nicotinamide cytosine dinucleotide)-RNA, and (5'-nicotinamide uracil dinucleotide)NUD-RNA.

[0007] These findings reveal a novel PTM, "RNAization," which is a protein PTM mediated by the attachment of NND-nucleotide sequences via ART.

[0008] Accordingly, the present invention relates in a first aspect to a method for attaching a 5'-nicotinamide nucleobase dinucleotide (NND, also indicated herein as NXD) cap nucleic acid sequence to a fusion protein or complex, the method comprising (a) a heterofusion protein or (ii) a complex comprising (i) a poly(peptide) of interest to be fused to a tag, and contacting the 5'-NND cap nucleic acid sequence and ADP-ribosyltransferase (ART) under conditions in which the 5'-NND cap nucleic acid sequence is covalently attached to the tag, wherein the protein is complexed with the tag under physiological conditions, and the tag comprises a recognition motif of ART, preferably (i) an amino acid motif DV R (ii) A sequence that is at least 80% identical to sequence number 1, assuming that the underlined Arg in PVRD (sequence number 7) is preserved and preferably sequence number 7 is preserved; (ii) Amino acid motif DV R Assuming that the underlined Arg in PVRD (SEQ ID NO: 7) is preserved and preferably SEQ ID NO: 7 is preserved, a sequence that is at least 80% identical to SEQ ID NO: 2; (iii) Amino acid motif DV R Assuming that the underlined Arg in PVRD (SEQ ID NO: 7) is preserved, and preferably SEQ ID NO: 7 is preserved, then a sequence that is at least 80% identical to SEQ ID NO: 3; (iv) Amino acid motif LADGVEGYLRASEASRD R Assuming that the underlined Arg in VE (SEQ ID NO: 8) is preserved, and preferably SEQ ID NO: 8 is preserved, then the sequence is at least 80% identical to SEQ ID NO: 4; (v) Amino acid motif LADGVEGYLRASEASRD R Assuming that the underlined Arg in VE (SEQ ID NO: 8) is preserved and preferably SEQ ID NO: 8 is preserved, then a sequence that is at least 80% identical to SEQ ID NO: 5; or (vi) amino acid motif LADGVEGYLRASEASRD RAssuming that the underlined Arg in VE (sequence number: 8) is preserved, and preferably sequence number: 8 is preserved, then the sequence includes or consists of sequence number: 6 or a sequence that is at least 80% identical to it.

[0009] Nicotinamide adenine dinucleotide (NAD) is a central coenzyme for metabolism. As found in all living cells, NAD is called a dinucleotide because it consists of two nucleotides linked by their phosphate groups. Each nucleotide contains the adenine nucleobase and another nicotinamide. NAD exists in two forms: oxidized and reduced forms, respectively, NAD+ and NADH (where H is hydrogen).

[0010] The term “nucleic acid sequence” (also referred to herein as “nucleic acid molecule”) includes, in accordance with the present invention, DNA, e.g., double-stranded or single-stranded DNA and RNA. Nucleic acid sequences are preferably single-stranded, e.g., single-stranded DNA and RNA. In this regard, “DNA” (deoxyribonucleic acid) refers to any strand or sequence of nucleotide bases, which are chemically constructed blocks adenine (A), guanine (G), cytosine (C), and thymine (T) linked together on a deoxyribose sugar backbone. DNA may have one strand of nucleotide bases or two complementary strands that can form a double helix structure. “RNA” (ribonucleic acid) refers to any strand or sequence of nucleotide bases, which are chemically constructed blocks adenine (A), guanine (G), cytosine (C), and uracil (U) linked together on a ribose sugar backbone. RNA typically has one strand of nucleotide bases, e.g., mRNA. This also includes single-stranded and double-stranded hybrid molecules, namely DNA-DNA, DNA-RNA, and RNA-RNA.

[0011] Nucleic acid molecules can also be modified by any means known in the art. Non-limiting examples of such modifications include methylation, substitution with one or more naturally occurring nucleotide analogs, and internucleotide modifications, such as those by non-charged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.) and those by charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.). Nucleic acid molecules, also hereinafter referred to as polynucleotides, may contain one or more further covalently bonded moieties, such as proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), intercalators (e.g., acridine, psoralens, etc.), chelating agents (e.g., metals, radiometals, iron, metal oxides, etc.), alkylating agents, fluorophores (e.g., Alexa or Cy dyes), fluorescent quenchers, or biotin. Polynucleotides can be derivatized by the formation of methyl or ethyl phosphotriester or alkylphosphoamidate linkages. The present invention further includes nucleic acid mimetic molecules known in the art, such as synthetic or semi-synthetic derivatives of DNA or RNA, and mixed polymers.

[0012] As will be further detailed below in this specification, such nucleic acid mimetic molecules or nucleic acid derivatives according to the present invention include phosphorothioate nucleic acids, phosphoramidate nucleic acids, 2'-O-methoxyethyl ribonucleic acid, morpholino nucleic acid, hexitol nucleic acid (HNA), peptide nucleic acid (PNA), and located nucleic acid (LNA) (see Braasch and Corey, Chem Biol 2001, 8: 1).

[0013] This also includes nucleic acids containing modified bases, such as thiouracil, thioguanine, and fluorouracil. Nucleic acid molecules typically contain genetic information, such as information used by cellular mechanisms for producing proteins and / or polypeptides. Nucleic acid molecules considered herein may further include promoters, enhancers, response factors, signal sequences, polyadenylated sequences, introns, 5'- and 3'-noncoding regions, etc.

[0014] The 5'-nicotinucleotide dinucleotide (NND) capped nucleic acid sequence refers to a nucleic acid sequence in which NND is linked to the 5'-end of the nucleic acid sequence via a diphosphate linkage. A nucleobase is a nitrogen-containing biological compound that forms a nucleoside. A nucleoside is a component of a nucleotide in which all of these monomers constitute the basic building blocks of nucleic acids.

[0015] The term "protein" (also referred to as "polypeptide") when used interchangeably with the term "polypeptide" herein, describes a linear molecular chain of amino acids such as a single-chain protein containing at least 50 amino acids or fragments thereof. The term "peptide" when used herein, describes a group of molecules consisting of up to 49 amino acids. The term "peptide" when used herein, describes a group of molecules corresponding to an increasing preference of at least 15 amino acids, at least 20 amino acids, at least 25 amino acids and at least 40 amino acids. The groups of peptides and polypeptides are referred to together using the term "(poly)peptide". (Poly)peptides can further form oligomers consisting of at least two identical or different molecules. The corresponding higher-order structures of such multimers are accordingly referred to as homo- or heterodimers, homo- or heterotrimers, etc. Furthermore, peptidomimetics of such proteins / (poly)peptides in which the amino acid(s) and / or peptide bond(s) are replaced by functional analogs are also encompassed by the present invention. Such functional analogs include all known amino acids other than the 20 amino acids encoded by genes such as selenocysteine. The terms "(poly)peptide" and "protein" also refer to naturally modified (poly)peptides and proteins in which the modification is carried out by, for example, glycosylation, acetylation, phosphorylation and similar modifications well known in the art.

[0016] The polypeptide of interest can be any polypeptide that desirably is modified by the novel PTM RNA modification disclosed herein. Examples of polypeptides of interest are described below in this specification.

[0017] In both the fusion proteins according to the invention and the complexes according to the invention, the polypeptide of interest is attached to a tag. The tag can be attached to the N-terminus and C-terminus of both ends of the polypeptide of interest.

[0018] In the complex, the tag is non-covalently linked (under physiological conditions) to the polypeptide of interest, preferably via the binding of biotin to avidin, streptavidin or neutravidin. Biotin can be linked to the polypeptide of interest and avidin, streptavidin or neutravidin, either to the tag or vice versa. Avidin is a protein derived from both birds and amphibians that exhibits a significant affinity for biotin and is a cofactor that plays a role in multiple eukaryotic biological processes. Avidin, streptavidin and neutravidin have the ability to bind up to four biotin molecules. The avidin-biotin complex is the strongest known non-covalent interaction between a protein and a ligand (Kd = 10 -15 M). The formation of the bond between biotin and avidin is very rapid and once formed, is not affected by excessive pH, temperature, organic solvents and other denaturing agents. Therefore, the indication "under physiological conditions" means that the binding must occur under such conditions, but does not exclude the possibility that the binding is also possible under non-physiological conditions. The term "physiological conditions" refers to the conditions of the external or internal environment that can occur naturally for such an organism or cell line, in contrast to artificial laboratory conditions.

[0019] Protein ADP-ribosylation is an important post-translational modification that plays a versatile role in multiple biological processes. ADP-ribosylation is catalyzed by a group of enzymes known as ADP-ribosyltransferases (ARTs). Prior art has shown that using nicotinamide adenine dinucleotide (NAD) as a donor, ARTs covalently link from NAD to their substrate (poly)peptides, forming mono-ADP-ribosylation or poly-ADP-ribosylation (PAR).

[0020] The present invention has unexpectedly revealed a novel function of ART. ART can transfer not only NAD but also 5'-NND capped nucleic acid sequences to their substrate (poly)peptides. The following examples of this specification particularly demonstrate that ART ModB can covalently ligate 5'-NND capped nucleic acid sequences to substrate (poly)peptides. The ADP-ribose ligation was particularly found to be located between the RNA and the arginine side chain within the substrate (poly)peptide.

[0021] It has been further found that the substrate (poly)peptide of ART ModB contains a motif specifically recognized by ART, and that this motif contains a specific arginine, where the side chain of the arginine acts as a site for linkage of a 5'-NND cap nucleic acid sequence to the substrate (poly)peptide. This motif can be attached as a tag to any desired (poly)peptide, as described herein, and as a result, any desired (poly)peptide can be modified by novel PTM RNAization.

[0022] For this reason, the tag of the present invention comprises at least one arginine that acts as a recognition motif for ART and preferably as a site for linking the 5'-NND cap nucleic acid sequence to a substrate (poly)peptide.

[0023] As discussed, RNAization is described in the following examples herein using ART ModB. The substrate (poly)peptide of ModB is protein rS1. Protein rS1 contains six domains (DI~DVI), and the recognition motif of ART was found in domains DII and DVI. The wild-type motifs in domains DII and DVI correspond to SEQ ID NOs: 1 and 4, respectively. Two wild-type motifs were further investigated, and SEQ ID NOs: 2 and 5 were identified as “essential” motifs. The wild-type motifs (motives) were also processed into so-called “vokuhila” motifs, which have a short side chain and a long side chain adjacent to the site of conjugation of a 5'-NND cap nucleic acid sequence.

[0024] All of these motifs contain two β-sheets and a central loop. The loops in domains DII and DVI are shown in sequence numbers 7 and 8, respectively. The β-sheets are thought to position the loops, but the amino acids in the loops are recognized by ART, and here the conserved and functionally essential arginine (R) acts as an attachment site for the 5'-NND cap nucleic acid sequence via its side chain.

[0025] Therefore, the tag is preferably (i) amino acid motif DV R (ii) A sequence that is at least 80% identical to sequence number 1, assuming that the underlined Arg in PVRD (sequence number 7) is preserved and preferably sequence number 7 is preserved; (ii) Amino acid motif DV R Assuming that the underlined Arg in PVRD (SEQ ID NO: 7) is preserved and preferably SEQ ID NO: 7 is preserved, a sequence that is at least 80% identical to SEQ ID NO: 2; (iii) Amino acid motif DV R Assuming that the underlined Arg in PVRD (SEQ ID NO: 7) is preserved, and preferably SEQ ID NO: 7 is preserved, then a sequence that is at least 80% identical to SEQ ID NO: 3; (iv) Amino acid motif LADGVEGYLRASEASRD RAssuming that the underlined Arg in VE (SEQ ID NO: 8) is preserved, and preferably SEQ ID NO: 8 is preserved, then the sequence is at least 80% identical to SEQ ID NO: 4; (v) Amino acid motif LADGVEGYLRASEASRD R Assuming that the underlined Arg in VE (SEQ ID NO: 8) is preserved and preferably SEQ ID NO: 8 is preserved, then a sequence that is at least 80% identical to SEQ ID NO: 5; or (vi) amino acid motif LADGVEGYLRASEASRD R Assuming that the underlined Arg in VE (sequence number: 8) is preserved, and preferably sequence number: 8 is preserved, then the sequence includes or consists of sequence number: 6 or a sequence that is at least 80% identical to it.

[0026] According to the present invention, the term “percent (%) sequence identity” refers to the number of identical nucleotide / amino acid matches ("hits") of two or more aligned nucleic acids or amino acid sequences compared to the number of nucleotides or amino acid residues that make up a template nucleic acid or amino acid sequence. In other terms, when alignment is used for two or more sequences or subsequences, the percentage of identical amino acid residues or nucleotides (e.g., 70%, 75%, 80%, 85%, 90%, or 95% identity) may be determined when (sub)sequences are compared and aligned for the greatest match across a comparison window or over a specified region, such as measured using a sequence comparison algorithm known in the art, or when they are manually aligned and visually examined. This definition also applies to complements of any sequences being aligned.

[0027] Nucleotide and amino acid sequence analysis according to the present invention is preferably performed using the NCBI BLAST algorithm (Stephen F. Altschul, Thomas L. Madden, Alejandro A. Schaeffer, Jinghui Zhang, Zheng Zhang, Webb Miller, and David J. Lipman (1997), Nucleic Acids Res. 25:3389-3402). BLAST can be used for nucleotide sequences (nucleotide BLAST) and amino acid sequences (protein BLAST). Those skilled in the art know of further suitable programs for aligning nucleic acid sequences.

[0028] As defined herein, sequence identity of at least 80%, preferably at least 85%, more preferably at least 90%, and most preferably at least 95%, is conceived by the present invention. However, with increasing preference, sequence identity of at least 97.5%, at least 98.5%, at least 99%, at least 99.5%, at least 99.8%, and 100% is also conceived by the present invention.

[0029] According to a preferred embodiment of the first aspect of the present invention, the nucleobase of NND is a purine base or a pyrimidine base, preferably selected from adenine, guanine, cytosine, thymine, and uracil.

[0030] Purine bases are preferred over pyrimidine bases.

[0031] The five preferred nucleobases adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U) are referred to as major or normative. They function as the basic units of the genetic code; bases A, G, C, and T are found in DNA, while A, G, C, and U are found in RNA. Thymine and uracil are distinguished by the presence or absence of a methyl group on the fifth carbon (C5) of their heterocyclic six-membered rings, respectively. Adenine and guanine have a fused ring skeleton structure derived from purines and are therefore members of the purine base class. The ring structures of cytosine, uracil, and thymine are derived from pyrimidines, and therefore they are members of the pyrimidine base class.

[0032] Of the five preferred nucleobases adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), adenine (A) is preferred because it is a natural substrate of ART, or cytosine (C), guanine (G), thymine (T), and uracil (U) are preferred because they are unnatural substrates of ART. Unnatural substrates are not removed by human ADP-ribose hydrolase ARH1, thereby advantageously demonstrating increased RNA-protein stability (see Example 7). Of the cytosine (C), guanine (G), thymine (T), and uracil (U), three nucleobases—cytosine (C), guanine (G), and uracil (U)—are preferred because they form RNA.

[0033] For example, further nucleobases include xanthine, hypoxanthine, 7-methylguanine, 2,6-diaminopurine and 6,8-diaminopurine (purine bases), pseudouridine, N1-methylpseudridine, or 5,6-dihydrouracil, 5-methyluracil and 5-hydroxymethylcytosine (pyrimidine bases).

[0034] According to a preferred embodiment of the first aspect of the present invention, ART comprises or consists of a sequence that is at least 80% identical to sequence number 9 or sequence number 10.

[0035] Sequence ID: 9 is the amino acid sequence of ART ModB derived from Escherichia virus T4 deposited under Acc. No. CAA67254.1. Sequence ID: 10 is the amino acid sequence of cloned ModB used in the following examples herein, further comprising a His6 tag that works for the purification of ModB after it has been recombinantly produced from an expression vector as described below herein for heterofusion proteins, for example.

[0036] According to a preferred embodiment of the first aspect of the present invention, the method comprises, prior to step (a), a step of fusing a tag to a poly(peptide) of interest as defined with respect to the first aspect, thereby obtaining a heterofusion protein comprising the poly(peptide) of interest to be fusing to the tag.

[0037] In an alternative preferred embodiment of the first aspect of the present invention, the method includes, prior to step (a), a step of (a') complexing the tag with the poly(peptide) of interest, as defined with respect to the first aspect.

[0038] For example, nucleic acid sequences encoding the target poly(peptide), tag, and optionally a peptide linker can be introduced into a frame within an expression vector in an expressible form. The expression vector can then be introduced into host cells, which can be cultured under conditions that produce the heterofusion protein. The heterofusion protein can then be isolated from the cells.

[0039] As a substitute for the target poly(peptide), tags and optionally peptide linkers can be synthesized and linked via peptide synthesis to form heterofusion proteins.

[0040] According to a preferred embodiment of the first aspect of the present invention, the method comprises, after step (a), a step of (b) purifying or isolating the fusion protein or the complex to which the NND-5' cap nucleic acid sequence is attached.

[0041] Means and methods for the isolation or purification of proteins or peptides are known in the art. These means and methods include, but are not limited to, method techniques and processes such as ion exchange chromatography, gel filtration chromatography (size exclusion chromatography), affinity chromatography, high-pressure liquid chromatography (HPLC), reverse-phase HPLC, disk gel electrophoresis, or immunoprecipitation, see, for example, Sambrook, 2001, Molecular Cloning: A laboratory manual, 3rd ed, Cold Spring Harbor Laboratory Press, New York.

[0042] In a second aspect, the present invention relates to a fusion protein comprising a poly(peptide) of interest to be fused to a tag as defined in relation to the first aspect of the present invention, or to a complex comprising a poly(peptide) of interest to be complexed with a tag as defined in relation to the first aspect of the present invention.

[0043] The definitions and preferred embodiments of the first aspect of the present invention are applicable to the second aspect of the present invention with necessary modifications, insofar as they are applicable to the second aspect of the present invention.

[0044] Therefore, the fusion protein in the second phase is also a heterogeneous fusion protein, meaning that the amino acid sequence of the poly(peptide) does not occur naturally, and it should be noted that the tag may be part of the protein rS1 in nature.

[0045] Similarly, the complex of the second aspect of the present invention is also preferably formed by the binding of biotin to avidin, streptavidin, or neutraavidin.

[0046] Furthermore, with respect to the second aspect, the tag also includes the recognition motif of ART, preferably (i) a sequence that is at least 80% identical to sequence number 1 or thereto, assuming that the underlined Arg in the amino acid motif DVRPVRD (sequence number 7) is preserved and preferably sequence number 7 is preserved; (ii) a sequence that is at least 80% identical to sequence number 2 or thereto, assuming that the underlined Arg in the amino acid motif DVRPVRD (sequence number 7) is preserved and preferably sequence number 7 is preserved; (iii) a sequence that is at least 80% identical to sequence number 3 or thereto, assuming that the underlined Arg in the amino acid motif DVRPVRD (sequence number 7) is preserved and preferably sequence number 7 is preserved; (iv) amino (v) A sequence that is at least 80% identical to sequence number 4, assuming that the underlined Arg in the acid motif LADGVEGYLRASEASRDRVE (SEQ ID NO: 8) is preserved and preferably sequence number 8 is preserved; (v) A sequence that is at least 80% identical to sequence number 5, assuming that the underlined Arg in the amino acid motif LADGVEGYLRASEASRDRVE (SEQ ID NO: 8) is preserved and preferably sequence number 8 is preserved; or (vi) A sequence that is at least 80% identical to sequence number 6, assuming that the underlined Arg in the amino acid motif LADGVEGYLRASEASRDRVE (SEQ ID NO: 8) is preserved and preferably sequence number 8 is preserved.

[0047] The present invention also relates, in a second aspect, to vectors encoding nucleic acid molecules, preferably fusion proteins of the second aspect.

[0048] In the present invention, the term “vector” preferably means a plasmid, cosmid, virus, bacteriophage, or another vector conventionally used in genetic engineering to carry, for example, the nucleic acid molecule of the present invention. The nucleic acid molecule of the present invention can be inserted into, for example, several commercially available vectors. Non-limiting examples include prokaryotic plasmid vectors, such as the pUC-series, pBluescript (Stratagene), pET-series expression vectors (Novagen), or pCRTOPO (Invitrogen), as well as vectors suitable for expression within mammalian cells, such as pREP (Invitrogen), pcDNA3 (Invitrogen), pCEP4 (Invitrogen), pMC1neo (Stratagene), pXT1 (Stratagene), pSG5 (Stratagene), EBO-pSV2neo, pBPV-1, pdBPVMMTneo, pRSVgpt, pRSVneo, pSV2-dhfr, pIZD35, pLXIN, pSIR (Clontech), pIRES-EGFP (Clontech), pEAK-10 (Edge Biosystems), pTriEx-Hygro (Novagen), and pCINeo (Promega). Examples of suitable plasmid vectors for Pichia pastris include plasmids pAO815, pPIC9K, and pPIC3.5K (all from Invitrogen).

[0049] Nucleic acid molecules to be inserted into the vector can be synthesized, for example, by standard methods. Ligation of coding sequences to transcription regulators and / or other amino acid coding sequences can also be performed using established methods. Transcription regulators (parts of the expression cassette) that ensure expression in prokaryotic or eukaryotic cells are well known to those skilled in the art. These factors include regulatory sequences that ensure the initiation of transcription (e.g., transcription start codons, promoters, e.g., natively related or heterologous promoters and / or insulators; see above), internal ribosome entry sites (IRES) (Owens, Proc. Natl. Acad. Sci. USA 98 (2001), 1471-1476), and optionally poly-A signals that ensure the termination of transcription and stabilization of the transcript. Further regulators may include transcription and translation enhancers. Preferably, the polynucleotide encoding the polypeptide / protein or fusion protein of the present invention is operably ligated to such expression regulators that enable expression in prokaryotic or eukaryotic cells. The vector may further include nucleic acid sequences encoding secretory signals as further regulatory factors. Such sequences are well known to those skilled in the art. Furthermore, depending on the expression system used, a leader sequence capable of directing the expressed polypeptide to a cellular compartment may be added to the polynucleotide coding sequence of the present invention. Such leader sequences are well known in the art.

[0050] Furthermore, it is preferable that the vector contains a selectable marker. Examples of selectable markers include genes encoding resistance to neomycin, ampicillin, hygromycin, chloramphenicol, and kanamycin. Specifically designed vectors enable the transport of DNA between different hosts, such as bacterial-fungal cells or bacterial-animal cells (e.g., Gateway systems available from Invitrogen). Expression vectors according to the present invention can induce replication and expression of the polynucleotides and encoded fusion proteins of the present invention. Apart from introduction via vectors such as phage vectors or viral vectors (e.g., adenoviruses, retroviruses), the nucleic acid molecules described herein may be designed for direct introduction or introduction into cells via liposomes. Furthermore, baculovirus systems or systems based on vaccinia virus or Semryki forest virus may be used as eukaryotic expression systems for the nucleic acid molecules of the present invention.

[0051] According to a more preferred embodiment of the second aspect of the present invention, the nucleic acid sequence is covalently attached to a tag at its 5' end, preferably to the conserved Arg side chain of the tag, via a nicotinamide nucleobase dinucleotide (NND), as defined with respect to the first aspect of the present invention.

[0052] As described herein above, according to the method of the present invention, a 5'-nicotinamide nucleobase dinucleotide (NAD) cap nucleic acid sequence can be attached to a fusion protein or complex, as defined with respect to the first aspect, and preferably to a conserved Arg of a tag contained in the fusion protein or complex.

[0053] Therefore, the fusion proteins or complexes of the above preferred embodiments can be obtained, are available, or will be obtained by the methods of the first aspect of the present invention.

[0054] In a third aspect, the present invention relates to a composition comprising a fusion protein or complex obtainable by the method of the first aspect, or a fusion protein and / or complex of the second aspect, preferably a pharmaceutical or diagnostic composition.

[0055] The definitions and preferred embodiments of the first and second aspects of the present invention are applicable to the third aspect of the present invention with necessary modifications, insofar as they are applicable to the third aspect of the present invention.

[0056] As used herein, the term “composition” means a composition comprising at least one fusion protein and / or complex as defined above, or a combination thereof, which is also collectively referred to below as a compound.

[0057] According to the present invention, the term “pharmaceutical composition” refers to a composition for administration to a patient, preferably a human patient. The pharmaceutical compositions of the present invention comprise the compounds described above. Optionally, the composition may comprise further molecules that can alter the properties of the compounds of the present invention, thereby stabilizing, modifying, and / or activating their functions, for example. The composition may be in solid, liquid, or gaseous form, and may particularly be in the form of powder(s), tablets(s), solutions(s), or aerosols(s). The pharmaceutical compositions of the present invention may comprise optional and further pharmaceutically acceptable carriers. Examples of suitable pharmaceutical carriers are well known in the art and include phosphate-buffered saline solutions, water, emulsions, e.g., oil / water emulsions, various types of wetting agents, sterile solutions, organic solvents, e.g., DMSO. Compositions comprising such carriers may be formulated by well-known conventional methods. These pharmaceutical compositions may be administered to a subject in an appropriate dose. Dosage and administration methods are determined by the attending physician and clinical factors. As is well known in the field of medicine, the dose for any given patient depends on many factors, including the patient's size, weight, surface area, age, the specific compound administered, sex, time and route of administration, general health condition, and other drugs administered concurrently. The therapeutically effective dose for a given situation can be easily determined by routine experimentation and is within the scope of the usual clinician's or physician's skill and judgment. Generally, the therapeutic regimen as a regular administration of a pharmaceutical composition should be in the range of 1 μg to 5 g per day. However, a more preferable dose may be in the range of 0.01 mg to 100 mg per day, even more preferably 0.01 mg to 50 mg, and most preferably 0.01 mg to 10 mg. Furthermore, if the compound is, for example, siRNA, the total effective dose of the administered pharmaceutical composition is typically less than about 75 mg per kg of body weight, for example, less than about 70, 60, 50, 40, 30, 20, 10, 5, 2, 1, 0.5, 0.1, 0.05, 0.01, 0.005, 0.001, or 0.0005 mg per kg of body weight. More preferably, the amount is less than 2000 nmol of iRNA agent per kg of body weight (for example, about 4.4 x 10⁻⁶). 16(copy), for example, iRNA agents less than 1500, 750, 300, 150, 75, 15, 7.5, 1.5, 0.75, 0.15, 0.075, 0.015, 0.0075, 0.0015, 0.00075, or 0.00015 nmol per kg of body weight. The length of treatment required to observe the post-treatment interval during which changes and responses occur varies depending on the desired effect. The specific amount can be determined by conventional tests well known to those skilled in the art.

[0058] In pharmaceutical compositions and the medical uses described below herein, the active compound may be the (poly)peptide and / or nucleic acid sequence of interest on the tag. A non-limiting example of a category of pharmaceutically active (poly)peptides of interest is an antibody. Therapeutic antibodies against several types of cancer and autoimmune diseases are commercially available. A non-limiting example of a category of pharmaceutically active nucleic acid sequences is siRNA, which can be designed to silence the expression of virtually any desired gene. It is also possible to combine the desirable properties of antibodies and siRNA. For example, an antibody may bind to a tissue-specific antigen, thereby conferring the activity of siRNA to a particular tissue or at least focusing it.

[0059] The cosmetic compositions according to the present invention are for use in non-therapeutic applications. Cosmetic compositions may also be defined by their intended use as compositions intended to be rubbed, poured, sprinkled or sprayed, or otherwise applied to the human body to cleanse, beautify, enhance attractiveness or alter appearance. The specific formulations of the cosmetic compositions according to the present invention are not limited. Formulations that can be conceived include rinse solutions, emulsions, creams, emulsions, gels such as hydrogels, ointments, suspensions, powders, solid sticks, foams, sprays and shampoos. For this purpose, the cosmetic compositions according to the present invention may further comprise a diluent and / or carrier that is acceptable for cosmetic use. Selecting a suitable carrier and diluent depending on the desired formulation is within the scope of the art. Suitable diluents and carriers that are acceptable for cosmetic use are well known in the art, including those referenced in Bushell et al. (WO 2006 / 053613). Preferred formulations for the cosmetic compositions are rinse solutions and creams. The preferred amount of the cosmetic composition according to the present invention to be applied in a single application is 0.1 to 10 g, more preferably 0.1 to 1 g, and most preferably 0.5 g. The amount applied also needs to be adapted to the size of the area being treated.

[0060] In preferred embodiments of the three aspects of the present invention described above, the nucleic acid in the nucleic acid sequence is RNA, DNA, PNA, morpholino, or LNA, or a combination thereof, and is preferably RNA.

[0061] In the following examples of this specification, it is shown that ART can attach a 5'-nicotinamide nucleobase dinucleotide (NND)-cap RNA sequence to a tag having a tag recognition site. For this reason, the nucleic acid in the nucleic acid sequence is most preferably RNA.

[0062] Since ART can not only attach 5'-nicotinamide nucleobase dinucleotide (NND)-cap RNA sequences to proteins, but can also covalently link single or multiple ADP-ribose moieties derived from NAD to their substrate (poly)peptides, it is thought that ART can attach not only 5'-nicotinamide nucleobase dinucleotide (NND)-cap RNA sequences, but generally 5'-nicotinamide nucleobase dinucleotide (NND)-cap nucleic acid sequences to tags.

[0063] Following RNA, DNA, PNA, morpholino, or LNA are nucleic acids that can be used to form useful nucleic acid sequences.

[0064] PNAs are oligonucleotide analogs in which the sugar-phosphate backbone is replaced by a pseudopeptide backbone. They bind DNA and RNA with high specificity and selectivity, producing PNA-RNA and PNA-DNA hybrids that are more stable than the corresponding nucleic acid complexes.

[0065] LNAs are RNA derivatives in which the ribose ring is constrained by a methylene bond between the 2'-oxygen and 4'-carbon atoms. The ribose portion of an LNA nucleotide is modified by excessive crosslinking linking the 2'-oxygen and 4'-carbon atoms. This crosslinking "fixes" the ribose to the 3'-end (North) conformation, often found in A-form double helix. This fixed ribose conformation enhances base stacking and pre-skeletal organization. This significantly increases the hybridization properties (melting point) of the oligonucleotide.

[0066] Morpholinos are synthetic, uncharged, P-chiral analogs of nucleic acids. Morpholino oligonucleotides are typically constructed by linking together 25 subunits, each having one of four nucleic acid bases.

[0067] The aforementioned types of nucleotides can be combined into a single sequence whenever desired. Such sequences can be chemically synthesized or commercially available.

[0068] The length of the nucleic acid sequence used herein is not particularly limited, but the nucleic acid sequence preferably contains about 100 or fewer nucleotides, preferably about 50 or fewer nucleotides, and most preferably about 25 or fewer nucleotides.

[0069] In this specification, the term "approximately" means ±20%, ±10%, and ±5%, with increasing preference.

[0070] In a more preferred embodiment of the three aspects of the present invention described above, the nucleic acid sequence is siRNA, an antisense molecule shRNA.

[0071] The nucleic acid sequence is preferably an antisense molecule (e.g., antisense oligonucleotide, e.g., LNA-GapmeR, Antagomir, or antimiR) capable of inhibiting the expression of a target nucleic acid molecule, typically an mRNA expressed in a host or organism (e.g., a human), siRNA, or shRNA. Such nucleic acid sequences may include DNA sequences (e.g., LNA-GapmeR) or RNA sequences (e.g., siRNA). As will be further detailed below herein, the nucleotide compounds that inhibit the expression of the target nucleic acid molecule may be single-stranded (e.g., LNA-GapmeR) or double-stranded (e.g., siRNA).

[0072] Antisense techniques for downregulating target nucleic acid molecules are well-established and widely used in the art to treat various diseases. The basic idea of ​​antisense techniques is the use of oligonucleotides to silence selected target RNAs through strong specificity of complementarity-based pairing (Re, Ochsner J., 2000 Oct; 2(4): 233-236). Details of the antisense construct compound classes of siRNA, shRNA, and antisense oligonucleotides are provided below herein. As will be further detailed below herein, antisense oligonucleotides are single-stranded antisense constructs, while siRNA and shRNA are double-stranded antisense constructs in which one strand contains an antisense oligonucleotide sequence (i.e., the so-called antisense strand). All of these compound classes can be used to achieve downregulation or inhibition of target RNA.

[0073] According to the present invention, the term "siRNA" refers to short interfering RNA, also known as silencing RNA. siRNA is a class of 12-30, preferably 18-30, more preferably 20-25, and most preferably 21-23 or 21 nucleotide-long double-stranded RNA molecules that play various roles in biology. Most notably, siRNA is involved in RNA interference (RNAi) pathways, in which siRNA interferes with the expression of specific genes. In addition to their roles in RNAi pathways, siRNA also acts, for example, as an antiviral mechanism or in RNAi-related pathways in the formation of chromatin structure in the genome. siRNA has a well-defined structure: a short double-stranded RNA (dsRNA) having at least one RNA strand with a favorable overhang. Each strand typically has a 5' phosphate group and a 3' hydroxyl (-OH) group. This structure is the result of processing by dicers, an enzyme that converts either long dsRNA or small hairpin RNA to siRNA. siRNA can also be introduced into cells exogenously (artificially) to produce specific knockdown of a gene of interest. Therefore, any gene whose sequence is publicly known can, in principle, be targeted based on sequence complementarity with appropriately tailored siRNA. Double-stranded RNA molecules or their metabolites can mediate target-specific nucleic acid modifications, particularly RNA interference and / or DNA methylation. Preferably, at least one RNA strand has 5'- and / or 3'-overhangs. Preferably, one or both of the double strands have 3'-overhangs of 1-5 nucleotides, more preferably 1-3 nucleotides, and most preferably 2 nucleotides. In general, any RNA molecule suitable to function as siRNA is conceived in this invention. To date, the most effective silencing has been obtained using siRNA double strands paired in a manner consisting of 21-nt sense and 21-nt antisense strands having 2-nt 3'-overhangs.The 2-nt 3' overhang sequence contributes little to the specificity of target recognition, which is limited to non-paired nucleotides adjacent to the first base pair (Elbashir et al. Nature. 2001 May 24; 411(6836):494-8). 2'-deoxynucleotides in the 3' overhang are effective as ribonucleotides but are often less expensive to synthesize and possibly more nuclease-resistant. The siRNA according to the present invention comprises an antisense strand containing or consisting of a sequence complementary to at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, or at least 21 nucleotides of the target nucleic acid sequence, with increasing preference.

[0074] A preferred example of siRNA is endoribonuclease-prepared siRNA (esiRNA). esiRNA is a mixture of siRNA oligos, produced by cleaving long double-stranded RNA (dsRNA), and an endoribonuclease such as E. coli RNase III or Dicer. esiRNA is an alternative concept to the use of chemically synthesized siRNA for RNA interference (RNAi). For esiRNA preparation, the cDNA of an lncRNA template can be amplified by PCR and tagged with two bacteriophage promoter sequences. RNA polymerase is then used to produce a long double-stranded RNA complementary to the target gene cDNA. This complementary RNA can then be digested by E. coli-derived RNase III to produce short duplicate fragments of siRNA having a length of 18-25 base pairs. This complex mixture of short double-stranded RNAs is similar to the mixture produced by Dicer cleavage in vivo, and is therefore referred to as endoribonuclease-prepared siRNA or short esiRNA. Therefore, esiRNA is a heterogeneous mixture of siRNAs, all of which target the same mRNA sequence. esiRNA results in highly specific and effective gene silencing.

[0075] The “shRNA” according to the present invention is a short hairpin RNA, which is an RNA sequence that forms a (dense) hairpin turn and can also be used to silence gene expression via RNA interference. The shRNA preferably uses the U6 promoter for its expression. The shRNA hairpin structure is cleaved into siRNA by cellular mechanisms, which is then bound to an RNA-induced silencing complex (RISC). This complex binds to and cleaves mRNA, which then conforms to the shRNA it binds to. The shRNA according to the present invention includes, with increasing preference, an antisense strand containing or consisting of a sequence complementary to at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, or at least 21 nucleotides of the target nucleic acid sequence.

[0076] The term "antisense oligonucleotide" according to the present invention refers to a single-stranded nucleotide sequence complementary to a target nucleic acid sequence by Watson-Crick base pair hybridization, thereby blocking the target nucleic acid sequence. Antisense oligonucleotides may or may not be modified. Generally, they are relatively short (preferably 13-25 nucleotides). Furthermore, they are specific to the target nucleic acid sequence, i.e., they hybridize to specific sequences in the entire pool of targets present in the target cell / organism. Antisense oligonucleotides according to the present invention contain or consist of sequences complementary to at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, or at least 25 nucleotides of the target nucleic acid sequence, with increasing preference.

[0077] The antisense oligonucleotide is preferably LNA-GapmeR, Antagomir, or antimiR.

[0078] LNA-GapmeR, or simply GapmeR, is a potential antisense oligonucleotide used for highly effective inhibition of mRNA and lncRNA function. GapmeRs function by RNase H-dependent degradation of complementary RNA targets. They are a superior alternative to siRNA for mRNA and lncRNA knockdown. They are advantageously taken up into cells without the use of transfection reagents. GapmeRs contain a central strand of DNA monomers adjacent to a block of LNA. GapmeRs are preferably 14-16 nucleotides long and optionally well-phosphorothiolated. The DNA gap is also suitable for activating RNase H-mediated degradation of targeted RNA and for directly targeting transcripts in the nucleus. LNA-GapmeRs according to the present invention contain a sequence complementary to at least 13 nucleotides, at least 14 nucleotides, or at least 15 nucleotides of the target nucleic acid sequence, with increasing preference.

[0079] As described, AntimiR is an oligonucleotide inhibitor initially designed to be complementary to miRNA. AntimiR against miRNA has been widely used as a tool for gaining an understanding of specific miRNA function and as a potential therapeutic agent. As used herein, AntimiR is designed to be complementary to a target nucleic acid sequence. AntimiR is preferably 14 to 23 nucleotides long. More preferably, with increasing preference, AntimiR according to the present invention comprises or consists of a sequence complementary to at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, or at least 23 nucleotides of the target nucleic acid sequence.

[0080] AntimiR is preferably AntagomiR. AntagomiR is preferably a 21-23 nucleotide synthetic 2-O-methylRNA oligonucleotide, preferably well complementary to a selected target nucleic acid sequence. Although AntagomiR was initially designed for miRNA, they can also be designed for mRNA. Therefore, AntagomiR according to the present invention preferably contains a sequence complementary to 21-23 nucleotides of the target nucleic acid sequence. AntagomiR is preferably synthesized by adding a 2'-OMe modified base (the 2'-hydroxyl group of ribose is replaced with a methoxy group), phosphorothioates on the first two and last four bases (phosphodiester bonds are replaced with phosphorothioates), and a cholesterol motif at the 3' end via a hydroxyprolinol modified bond. The addition of 2'-OMe and phosphorothioate modifications improves biostability, and cholesterol conjugation enhances the distribution and cellular penetration of AntagomiR.

[0081] Useful antisense molecules (antisense oligonucleotides, e.g., LNA-GapmeR, Antagomir, antimiR, etc.), siRNAs, and shRNAs according to the present invention are preferably synthesized chemically using conventional nucleic acid synthesizers. Sources of nucleic acid sequence synthesis reagents include Proligo (Hamburg, Germany), Dharmacon Research (Lafayette, CO, USA), Pierce Chemical (part of Perbio Science, Rockford, IL, USA), Glen Research (Sterling, VA, USA), ChemGenes (Ashland, MA, USA), and Cruachem (Glasgow, UK).

[0082] Due to their potent yet reversible ability to silence or inhibit target nucleic acid sequences in vivo, antisense molecules (antisense oligonucleotides, e.g., LNA-GapmeR, Antagomir, antimiR, etc.), siRNA, and shRNA are particularly well-suited for use in the pharmaceutical compositions of the present invention.

[0083] Antisense molecules (antisense oligonucleotides, e.g., LNA-GapmeR, Antagomir, antimiR, etc.), siRNA, and shRNA may contain modified nucleotides such as locked nucleic acid (LNA).

[0084] According to another preferred embodiment of the three aspects of the present invention, the nucleic acid sequence comprises a fluorescent label, preferably Alexa Fluor or Cy dye, at its 3' end, or biotin at its 3' end.

[0085] Since fluorescent labeling is particularly advantageous in the diagnostic compositions of the present invention, the fusion or complex of the present invention to which the nucleic acid sequence is attached can be deployed in vivo by a subject (preferably a human subject).

[0086] The Cy dye is preferably Cy2, Cy3, or Cy5. The Alex Fluor is preferably Alexa Fluor 488, 532, 546, 555, 568, 594, 647, 660, 680, 700, and 750.

[0087] Biotin at the 3' end of the nucleic acid sequence needs to be retained apart from any biotin that may be present in the complex according to the present invention. Biotin at the 3' end of the nucleic acid sequence can, via avidin, streptavidin, or neutraavidin, then act as a site for further attachment to the nucleic acid sequence, rather than for tagging and linking of the (poly)peptide of interest.

[0088] According to a more preferred embodiment of the three aspects of the present invention described above, the nucleic acid sequence includes a portion that can be used in click chemistry, such as an alkyne or azide.

[0089] "Click chemistry" is an established term in the field; for example, Kolb et al. (2001) Click chemistry: diverse chemical function from a few good reactions. Angew. Chem. Int. Ed. 40 (11):2004; Sletten et al. (2009) Bioorthogonal Chemistry: Fishing for Selectivity in a Sea of ​​Functionality. Angew. Chem. Int. Ed. 48:6998; Jewett et al. (2010) Cu-free click cycloaddition reactions in chemical biology. Chem. Soc. Rev. 39(4):1272; Best et al. (2009) Click Chemistry and Bioorthogonal Reactions: Unprecedented Selectivity in the Labeling of Biological Molecules. Biochemistry. 48:6571; and Lallana et al. (2011) Reliable and Efficient Procedures for the See Conjugation of Biomolecules through Huisgen Azide-Alkyne Cycloadditions. Angew. Chem. Int. Ed. 50:8794. While several reactions meet this criterion, Huisgen 1,3-dipolar cycloaddition of azides and terminal alkynes emerged as a pioneer.

[0090] According to preferred embodiments of the three aspects of the present invention, the poly(peptide) of interest is an antibody, antibody mimetic, cytokine, interleukin, transmembrane protein, membrane-anchored protein, enzyme, or DNA and / or RNA-binding protein.

[0091] As used in the present invention, the term "antibody" includes, for example, polyclonal or monoclonal antibodies. Furthermore, derivatives or fragments thereof that still retain binding specificity to a target are also included in the term "antibody." Antibody fragments or derivatives include, in particular, Fab or Fab' fragments, Fd, F(ab')2, Fv or scFv fragments, single-stranded domains VH or V-like domains, such as VhH or V-NAR-domains, and multimeric forms, such as minibodies, diabodies, tribodies or triplebodies, tetrabodies or chemically conjugated Fab'-multimers (see, for example, Harlow and Lane, "Antibodies, A Laboratory Manual," Cold Spring Harbor Laboratory Press, 1988; Harlow and Lane, "Using Antibodies: A Laboratory Manual," Cold Spring Harbor Laboratory Press, 1999; Altshuler EP, Serebryanaya DV, Katrukha AG. 2010, Biochemistry (Mosc)., vol. 75(13), 1584; Holliger P, Hudson PJ. 2005, Nat Biotechnol., vol. 23(9), 1126). In particular, the multimeric form includes a bispecific antibody capable of simultaneously binding to two different types of antigens. The first antigen may be found on the (poly)peptide of interest of the present invention. The second antigen may be, for example, a tumor marker specifically expressed on cancer cells or certain types of cancer cells. Non-limiting examples of bispecific antibody forms include Biclonics (bispecific, full-length human IgG antibody), DART (biaffinity retargeting antibody), and BiTE (consisting of two single-strand variable fragments (scFvs) of different antibodies) molecules (Kontermann and Brinkmann (2015), Drug Discovery Today, 20(7):838-847).

[0092] The term "antibody" also includes forms such as chimeric antibodies (human constant domain, non-human variable domain), single-chain antibodies, and humanized antibodies (human antibodies excluding non-human CDRs).

[0093] Various techniques for antibody production are well known in the art, as described, for example, in Harlow and Lane (1988) and (1999) and Altshuler et al., 2010, loc. cit. Therefore, polyclonal antibodies can be obtained from animal blood after immunization with antigens mixed with additives and adjuvants, while monoclonal antibodies can be produced by any technique that provides antibodies produced by serial cell line culture. Examples of such techniques include hybridoma techniques, trioma techniques, and human B-cell hybridoma techniques (see, e.g., Kozbor D, 1983, Immunology Today, vol.4, 7; Li J, et al. 2006, PNAS, vol. 103(10), 3557), as well as EBV hybridoma techniques for producing human monoclonal antibodies (Cole et al., 1985, Alan R. Liss, Inc, 77-96), as described in Harlow E and Lane D, Cold Spring Harbor Laboratory Press, 1988; Harlow E and Lane D, Using Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory Press, 1999, and first described in Koehler and Milstein, 1975), and EBV hybridoma techniques for producing human monoclonal antibodies. Furthermore, recombinant antibodies can be obtained from monoclonal antibodies or newly prepared using various display methods such as phage, ribosome, mRNA, or cell display. Suitable systems for the expression of recombinant (humanized) antibodies may be selected from, for example, bacteria, yeast, insects, mammalian cell lines, or transgenic animals or plants (see, e.g., U.S. Patent 6,080,560; Holliger P, Hudson PJ. 2005, Nat Biotechnol., vol. 23(9), 11265). Furthermore, the techniques described for the production of single-chain antibodies (see, in particular, U.S. Patent 4,946,778) may be adapted to produce single-chain antibodies specific to an epitope.Surface plasmon resonance, as used in the BIAcore system, can be used to improve the efficiency of phage antibodies.

[0094] As used herein, the term “antibody mimetic” refers to a compound that can bind specifically to an antigen in the same way as an antibody, but is structurally unrelated to an antibody. Antibody mimetics are typically artificial peptides or proteins with a molar mass of approximately 3 to 20 kDa. For example, antibody mimetics may be selected from the group consisting of affibody, adnectin, anticalin, DARPin, avimer, nanofitin, affilin, Kunitz domain peptides, Fynomers®, triple-specificity binding molecules, and prododies. These polypeptides are well known in the art and are described in further detail below herein.

[0095] As used herein, the term "aphibody" refers to a family of antibody mimetic molecules derived from the Z domain of Staphylococcus aureus protein A. Structurally, aphibody molecules are based on a triple-helical bundle domain that can also be incorporated into fusion proteins. In itself, aphibody molecules have a molecular weight of approximately 6 kDa and are stable at high temperatures and under acidic or alkaline conditions. Target specificity is obtained by randomization of 13 amino acids located in the two α-helices involved in the binding activity of the parent protein domain (Feldwisch J, Tolmachev V.; (2012) Methods Mol Biol. 899:103-26).

[0096] The term “adnectin” (also referred to as “monobody”), as used herein, refers to a molecule based on the 10th extracellular domain (10Fn3) of human fibronectin III, employing a 94-residue Ig-like β-sandwich folding with 2-3 exposed loops but lacking a central disulfide crosslink (Gebauer and Skerra (2009) Curr Opinion in Chemical Biology 13:245-255). Adnectin with desired target specificity can be genetically reconstructed by introducing modifications into specific loops of the protein.

[0097] As used herein, the term "anticalin" refers to a genetically engineered protein derived from lipocalin (Beste G, Schmidt FS, Stibora T, Skerra A. (1999) Proc Natl Acad Sci US A. 96(5):1898-903; Gebauer and Skerra (2009) Curr Opinion in Chemical Biology 13:245-255). Anticalin has an eight-stranded β-barrel that forms a highly conserved core unit within lipocalin and naturally forms a ligand binding site with four structurally distinct loops at its open end. Although anticalin is not homologous to the IgG superfamily, it exhibits characteristics previously considered typical of antibody binding sites: (i) high structural plasticity as a result of sequence variation and (ii) high conformational flexibility that allows for induced adaptation to targets with different shapes.

[0098] As used herein, the term “DARPin” refers to the indicated ankyrin repeat domain (166 residues) that typically provides a rigid interface arising from three repeated β-turns. DARPin usually has three repeats corresponding to an artificial consensus sequence, where six positions per repeat are randomized. Consequently, DARPin lacks structural flexibility (Gebauer and Skerra, 2009).

[0099] As used herein, the term "avimer" refers to a class of antibody mimetic compounds consisting of two or more peptide sequences, each 30-35 amino acids long, derived from the A domain of various membrane receptors and linked by a linker peptide. Binding of the target molecule occurs via the A domain, and the domain with the desired binding specificity can be selected, for example, by phage display technology. However, the binding specificity of different A domains contained in the avimer does not necessarily have to be identical (Weidle UH, et al., (2013), Cancer Genomics Proteomics; 10(4):155-68).

[0100] Nanophytin (also known as affitin) is an antibody-mimicking protein derived from Sac7d, a DNA-binding protein of Sulfolobus acidocaldarius. Nanophytin typically has a molecular weight of approximately 7 kDa and is designed to bind specifically to target molecules by randomizing the amino acids on its binding surface (Mouratou B, Beehar G, Paillard-Laurance L, Colinet S, Pecorari F., (2012) Methods Mol Biol.; 805:315-31).

[0101] As used herein, the term "affilin" refers to an antibody mimetic developed by randomly mutagenerating amino acids on the surface of either γ-B crystallin or ubiquitin as the backbone. Selection of affilins with desired target specificity is performed, for example, by phage display or ribosome display techniques. Depending on the backbone, affilins have a molecular weight of approximately 10 or 20 kDa. As used herein, the term affilin also refers to dimerized or multimerized forms of affilins (Weidle UH, et al., (2013), Cancer Genomics Proteomics; 10(4):155-68).

[0102] "Kunitz domain peptides" are derived from the Kunitz domain of Kunitz-type protease inhibitors such as bovine pancreatic trypsin inhibitors (BPTIs), amyloid precursor protein (APP), or tissue factor pathway inhibitors (TFPIs). The Kunitz domain has a molecular weight of approximately 6 kDA, and the domain with the required target specificity can be selected using display technologies such as phage display (Weidle et al., (2013), Cancer Genomics Proteomics; 10(4):155-68).

[0103] As used herein, the term "Fynomer®" refers to a non-immunoglobulin-derived polypeptide of the human Fyn SH3 domain. Fyn SH3-derived polypeptides are well known in the art and are described, for example, in Grabulovski et al. (2007) JBC, 282, p. 3196-3204, WO 2008 / 022759, Bertschinger et al (2007) Protein Eng Des Sel 20(2):57-68, Gebauer and Skerra (2009) Curr Opinion in Chemical Biology 13:245-255, or Schlatter et al. (2012), MAbs 4:4, 1-12).

[0104] The term “triple-specific binding molecule,” as used herein, refers to a polypeptide molecule having three binding domains and therefore preferably capable of specifically binding to three different epitopes. A preferred triple-specific binding molecule is TriTac. TriTac is a solid tumor T-cell engager consisting of three binding domains designed to have an extended serum half-life and to be about one-third the size of a monoclonal antibody.

[0105] As used herein, the term “probody” refers to a protease-activatable antibody prodrug. The probody consists of a faithful IgG heavy chain and a modified light chain. A masking peptide is fused to the light chain via a peptide linker that is cleavable by tumor-specific proteases. The masking peptide prevents the probody from binding to healthy tissue, thereby minimizing toxic side effects.

[0106] The cytokines are preferably selected from the group consisting of IL-2, IL-12, TNF-α, IFNα, IFNβ, IFNγ, IL-10, IL-15, IL-24, GM-CSF, IL-3, IL-4, IL-5, IL-6, IL-7, IL-9, IL-11, IL-13, LIF, CD80, B70, TNFβ, LT-β, ​​CD-40 ligand, Fas-ligand, TGF-β, IL-1α, and IL-1β.

[0107] The chemokine is preferably selected from the group consisting of IL-8, GROα, GROβ, GROγ, ENA-78, LDGF-PBP, GCP-2, PF4, Mig, IP-10, SDF-1α / β, BUNZO / STRC33, I-TAC, BLC / BCA-1, MIP-1α, MIP-1β, MDC, TECK, TARC, RANTES, HCC-1, HCC-4, DC-CK1, MIP-3α, MIP-3β, MCP-1-5, Eotaxin, Eotaxin-2, I-309, MPIF-1, 6Ckine, CTACK, MEC, Lymphotactin, and Fractalkine.

[0108] Enzymes are proteins that act as biological catalysts (biocatalysts). Catalysts accelerate chemical reactions. The molecules on which an enzyme can act are called substrates, and enzymes convert the substrate into different molecules known as products. Almost all metabolic processes within a cell require enzyme catalysis to occur at a rate fast enough to sustain life. The enzyme is preferably a sequence-specific DNA or RNA nuclease, and the DNA or RNA nuclease is most preferably a Cas nuclease (e.g., Cas 9, Cpf1, or Cms1). Note that Cas nucleases can specifically cleave a desired location within the cell's genome, the exact location being determined by a guide RNA. The guide RNA can be attached to the poly(peptide) of interest as described herein.

[0109] Proteins that bind to both DNA and RNA consolidate the ability of a single gene product to perform multiple functions. Such DNA- and RNA-binding proteins (DRBPs) regulate many cellular processes, including transcription, translation, gene silencing, microRNA biosynthesis, and telomere maintenance. An example is zinc finger-binding protein. DNA / RNA-binding proteins are preferably those responsible for RNA transport and / or localization.

[0110] In a more preferred embodiment of the three aspects of the present invention described above, the target poly(peptide) further comprises a purification tag, preferably a His-tag.

[0111] Therefore, the fusion protein or complex may contain a purification tag to facilitate its purification. Non-limiting examples of such tags include the ALFA-tag, V5-tag, Myc-tag, HA-tag, Flag-tag, Spot-tag, T7-tag, or NE-tag. The His-tag is used in the examples and is therefore preferred.

[0112] According to a more preferred embodiment of the three aspects described above, the poly(peptide) and the tag are fused via a peptide linker, preferably a G-linker or a GS-linker.

[0113] In a fusion protein, the tag is covalently linked to the target (poly)peptide by one or more peptide bonds. In the case of a single peptide bond, the target (poly)peptide and the tag are directly fused to each other.

[0114] According to the preferred embodiment described above, the (poly)peptide and tag of interest are fused to each other via a linker (containing two or more peptide bonds), such as a GS-linker or a G-linker. A G-linker is used in the appended examples.

[0115] In a fourth aspect, the present invention relates to a kit for attaching a 5'-nicotinamide adenine dinucleotide (NND) cap nucleic acid sequence to a (poly)peptide of interest, the kit comprising (a) a tag as defined in the first aspect, (b) an ADP-ribosyltransferase (ART) capable of covalently attaching the 5'-NND cap nucleic acid sequence to a nucleic acid molecule encoding the tag or ART, and (c) optionally instructions on how to covalently attach the tag having the ART to a (poly)peptide of interest.

[0116] The definitions and preferred embodiments of the first, second, and third aspects of the present invention are applicable to the fourth aspect of the present invention with necessary modifications, insofar as they are applicable to the fourth aspect of the present invention.

[0117] The kit preferably includes multiple compartments for the kit's components, with each compartment filled with one component. These compartments may be, for example, tubes or vials, bags or other packages.

[0118] ART is preferably provided in glycerol, for example, about 50% glycerol. Tag is preferably provided in about 50 mM Tris-HCl (pH 7.5), about 300 mM NaCl and about 50% glycerol.

[0119] Instructions on how to co-attach the tag having ART to the (poly)peptide of interest preferably further include guidance on the reaction conditions used with respect to the kit, such as temperature, time, and reaction buffer components. Preferred but not limited temperatures and / or times are about 2 hours and about 15°C. Preferred but not limited reaction buffers include about 50 mM Tris-HCl (pH 7.5), about 10 mM Mg(OAc)2, about 22 mM NH4Cl, about 1 mM EDTA, about 10 mM β-mercaptoethanol, about 1% glycerol, about 1 μM ADP-ribosyltransferase (ART, e.g., ModB), about 0.1–10 μM (poly)peptide, and about 1–10 μM 5'-NND capped nucleic acid sequence. These reaction conditions may also apply to the methods of the first aspect of the present invention.

[0120] The instructions may be included as part of the kit or as a leaflet, but if stored on the internet, they may also be in the form of a web link or QR code directing to the instructions.

[0121] According to a preferred embodiment of the fourth aspect, the kit further comprises a reaction buffer or buffer stock solution, wherein the reaction buffer or final reaction buffer is preferably prepared from the buffer stock solution. Mg(OAc)2 at a concentration of 50-200 mM, preferably about 100 mM; NH4Cl at a concentration of 100-500 mM, preferably about 220 mM; Tris-acetic acid at a concentration of 250-1000 mM, preferably about 500 mM, with a pH of 7.5; EDTA at a concentration of 5-15 mM, preferably about 10 mM; β-mercaptoethanol at a concentration of 50-200 mM, preferably about 100 mM; and Glycerol at a concentration of 5-15%, preferably about 10% Includes.

[0122] In the case of the buffer stock solution in the kit, the instructions on how to commonly attach the tag with ART to the target (poly)peptide preferably further include information on the dilution of the stock solution in order to obtain the reaction buffer from the buffer stock solution. Non-limiting examples of the buffer stock solution are 5x buffer stock solution and 10x buffer stock solution.

[0123] The reaction buffer according to the above preferred embodiment was used in the attached examples for ART ModB and was found to work very well for this enzyme when used in the method of the present invention. For this reason, the reaction buffer according to the above preferred embodiment of the kit is preferably also used in the method of the first aspect of the present invention.

[0124] According to a preferred embodiment of the fourth aspect, the kit further contains one or more of MgCl2 at a concentration of at least 0.25 M, preferably 0.5 M to 2 M, most preferably about 1 M. Imidazolidenicotinamide mononucleotide (Im-NMN), nuclease-free water and a positive control, preferably an oligonucleotide containing a fluorescent label at its 3'-end and / or a control fusion protein comprising a control poly(peptide) that is fused or complexed with a tag with respect to the first aspect of the present invention.

[0125] When MgCl2 is included in the kit, the instructions preferably further include the information that the concentration of EDTA should not be higher than the concentration of Mg 2+ ions when using the kit. This is because EDTA forms a complex with Mg 2+ [[ID=z16]]and the enzyme cannot come into contact with it.

[0126] When Im-NMN is included in the kit, the instructions preferably convey that Im-NMN is used at least 1000-fold in excess with respect to the 5'-nicotinamide nucleobase dinucleotide (NND) capped nucleic acid sequence when using the kit.

[0127] Nuclease-free water is included in the kit, especially if the kit contains a buffer stock solution. This nuclease-free water is then used to prepare the final buffer solution.

[0128] When using a kit in a method for attaching a 5'-nicotinamide nucleobase dinucleotide (NND) cap nucleic acid sequence to a fusion protein or complex as described herein, a positive control may be used to check whether the kit components themselves are working, but the desired RNAized fusion protein or complex may not be obtained. Possible reasons for such failure include, for example, that the tag was not sufficiently attached to the (poly)peptide of interest.

[0129] With respect to the embodiments characterized in this specification, particularly in the claims, each embodiment described in a dependent claim is intended to be combined with each embodiment of the respective claim (independent or dependent) to which the dependent claim depends. For example, in the case of independent claim 1 describing three options A, B, and C, dependent claim 2 describing three options D, E, and F, and claim 3 dependent on claims 1 and 2 and describing three options G, H, and I, it is understood that, unless otherwise specifically stated, the specification expressly discloses embodiments corresponding to combinations A, D, G; A, D, H; A, D, I; A, E, G; A, E, H; A, E, I; A, F, G; A, F, H; A, F, I; B, D, G; B, D, H; B, D, I; B, E, G; B, E, H; B, E, I; B, F, G; B, F, H; B, F, I; C, D, G; C, D, H; C, D, I; C, E, G; C, E, H; C, E, I; C, F, G; C, F, H; C, F, I.

[0130] Similarly, and even when an independent and / or dependent claim does not specify alternatives, if a dependent claim refers back to multiple prior claims, it is understood that any combination of subject matter covered by them is to be considered expressly disclosed. For example, in the case of independent claim 1, dependent claim 2 referring back to claim 1, and dependent claim 3 referring back to both claims 2 and 1, the combination of subject matter of claim 3 and 1 is to be considered as clearly and obviously disclosed as the combination of subject matter of claim 3, 2 and 1. If there is a further dependent claim 4 referring to any one of claims 1-3, the combinations of subject matter of claim 4 and 1, claim 4, 2 and 1, claim 4, 3 and 1, and claim 4, 3, 2 and 1 are to be considered clearly and obviously disclosed.

[0131] The above considerations apply to all attached claims with necessary modifications. [Brief explanation of the drawing]

[0132] The drawing is shown below: [Figure 1-1] Figure 1: Mechanism of ADP-ribosylation and proposed "RNAization". a. Here, the mechanism of ADP-ribosylation is experimentally shown for arginine. First, the N-glycosidic bond between ribose and nicotinamide is destabilized by the glutamine residue of ART. This leads to the formation of the ADP-ribose oxocarbenium ion. Nicotinamide acts as a leaving group. This electrophilic ion is attacked by the nucleophilic arginine residue of the receptor protein after glutamate-mediated proton extraction. This leads to the formation of the N-glycosidic bond.30 [Figure 1-2] b. Similar to ADP-ribosylation in the presence of NAD, we propose that ART may catalyze an "RNAization" reaction using NAD-RNA, thereby covalently attaching RNA to the receptor protein. [Figure 2-1]Figure 2: Post-translational protein modification of rS1 in vitro by ART ModB. a. Time course of ADP-ribosylation of rS1 by ModB (SDS-PAGE gel at completion shown in Figure 5b). b. Time course of RNAization of rS1 by ModB (SDS-PAGE gel at completion shown in Figure 5c). [Figure 2-2] c. In vitro kinetics of rS1 RNAization by ModB in the presence of excess NAD. d. In vitro kinetics of rS1 RNAization by ModB using 5'-NAD-100nt-RNA (Qβ-RNA) as a substrate, analyzed by SDS-PAGE (top panel). Shifted RNAized rS1 is highlighted with pink asterisks. 5'-P-100-nt-RNA is used as a negative control (bottom panel). e. Nuclease P1 digestion of RNAized protein rS1. Covalently attached 100nt-length RNA causes a shift in RNAized protein rS1 (approximately 100kDa) on SDS-PAGE. Treatment of RNAized protein rS1 with nuclease P1, which cleaves phosphodiester bonds, results in the degradation of mononucleotide-attached RNA. Nuclease P1 can convert RNA-modified rS1 to ADP-ribosylated rS1 (approximately 70 kDa), which can be visualized as a downward-shifted protein band on an SDS-PAGE gel. [Figure 3-1] Figure 3: Identification of the RNAization site of rS1. a-d. Specific removal of ADP-ribosylation and RNAization by ARH1. Enzyme kinetics of ARH1 in the presence of ADP-ribosylated or RNAized protein rS1, as analyzed by SDS-PAGE. [Figure 3-2]e. Pipeline for identifying modified amino acid residues by mass spectrometry. f. MALDI-TOF-MS of in vitro modified protein rS1. Isolation scan (MS1) and pseudo-MS2 (LIFT) spectra of peptide-ADPR conjugate. The given peptide AFLPGSLVDVRPVR (SEQ ID NO: 11) produces a peak at 2067 when conjugated to ADPR. MALDI-TOF-MS using in vitro modified protein rS1 produced the two spectra shown. LIFT parent ion isolation produced the given MS1 ​​with little interference. Note: The shifted b12 ion at m / z = 1483Th corresponding to the peptide with R5P modification shows the fragile nature of ADP-ribosylation. The resulting pseudo-MS2 produces sufficient sequencing ions to confirm the peptide sequence and ADPr modification on the arginine residue in the yellow box. [Figure 4-1] Figure 4: In vivo characterization of ADP-ribosylation and RNAization. a. Visualization of quantification of protein rS1 RNAization using nuclease P1 digestion and Western blot analysis. b. In vivo quantification of rS1 RNAization. [Figure 4-2] c) Quantification of ADP-ribosylation and d) RNAization. Modification of rS1 domains 1-6. Two biologically independent repeats (n=2). [Figure 4-3] e, ADP-ribosylation and RNAization of proteins containing the S1 motif by ModB, graphical representation. [Figure 4-4] f. SDS-PAGE analysis of RNAization and ADP-ribosylation of proteins rS1, RNase E, inactive NudC variants (*V157A, E174A, E177A, E178A) and BSA by ModB. n=2 biologically independent replicates. [Figure 4-5] g. Quantification of rS1 levels in the presence (+T4) or absence (-T4), n=4. h. ARH1-mediated removal of ADP-ribosylation and RNA modification during T4 infection. i. Time course of bacteriophage T4-mediated lysis of E. coli expressing plasmid carrier copies of ARH1-WT or its inactive mutant ARH1 D55,56A. [Figure 5-1] Figure 5: ADP-ribosylation and RNAization by T4 ART. a. Functional characterization of ART Alt and ModA. Self and targeted modification by Alt using NAD or NAD-RNA, analyzed by SDS-PAGE and autoradiography. b. Time course of ADP-ribosylation of rS1 by ModB, analyzed by SDS-PAGE. [Figure 5-2] c) Time course of rS1 RNAization by ModB, as analyzed by SDS-PAGE. d) Negative control for rS1 RNAization by ModB. RNAization assays were performed in the presence of 32P-RNA, in the absence of rS1 (-rS1), or in the absence of ModB (-ModB). [Figure 6-1] Figure 6: Characterization of RNAization of protein rS1 by ModB. a. Inhibition of in vitro RNAization of protein rS1 by ModB via the ART inhibitor 3-methoxybenzamide (3-MB). The reaction was performed using 32P-NAD-RNA 8-mers (32P-NAD-8-mer) and 32P-RNA 8-mers (negative control). b. In vitro digestion of RNAization and ADP-ribosylated protein rS1 by RNase T1. The reaction performed in the absence of RNase T1 (-) served as a negative control. Protein rS1 ADP-ribosylation in the presence of 32P-NAD was applied as a reference (S1-ADPr). [Figure 6-2] c. In vitro treatment of ADP-ribosylated and RNA-transferred protein rS1 with NudC (indicated by arrows) and alkaline phosphatase (AP), as well as trypsin digestion of ADP-ribosylated and RNA-transferred protein rS1. All samples were analyzed by 12% SDS-PAGE. Left panel: Coomassie-stained gel, middle panel: autoradiography scan, right panel: Coomassie-stained gel and autoradiography scan overlay. [Figure 7-1]Figure 7: Characterization of ModB specificity for NAD-RNA as a substrate. a, Competitive experiments using 32P-NAD-RNA and excess unlabeled NAD revealed ModB's preference for the former (compare Figures 2c and d). ADPr-rS1 acts as a reference. b, Analysis of RNAization dependence to the presence of the 5'-NAD-cap of RNA. 10% SDS-PAGE analysis of in vitro RNAization of protein rS1 by ModB in the presence of either 5'-NAD-(NAD-32P-Qβ), 5'-monophosphate-(5'-P32-Qβ), or 5'-triphosphate-Qβ-RNA (5'-P32PP-Qβ). [Figure 7-2] c. Characterization of ADPr-RNA as a substrate for ModB. NAD-8 mer was applied as a positive control. All reactions were analyzed by 12% SDS-PAGE. Left panel: Coomassie-stained gel, middle panel: autoradiography scan, right panel: Coomassie-stained gel and autoradiography scan overlay. [Figure 8] Figure 8: Specific removal of RNAization using chemical and enzymatic treatment. a, Different ADP-ribose-protein ligatures were shown to be either stable or unstable in the presence of HgCl2 and neutral hydroxylamine, indicating a relatively direct and rapid approach to identifying ADP-ribosylation sites. Treatment with hydroxylamine hydrolyzes the ligatures between glutamic acid and aspartic acid and ADP-ribose. HgCl2 specifically cleaves the thiol-glycosidic bond. ADP-ribosylated and RNAized protein rS1 were treated with hydroxylamine or HgCl2. Removal of ADPr or RNA resulted in a decrease in the radioactive signal of protein rS1. All samples were analyzed by 12% SDS-PAGE. No decrease in radioactive signal was measured in comparison with the control (untreated). b, In vitro kinetics of RNAized proteins in the presence of ARH1 or ARH3, analyzed by 12% SDS-PAGE. [Figure 9-1]Figure 9: LC-MS2 spectrum of rS1 peptide with peptide-ribose-5-phosphate. LC-MS2 spectrum of rS1 peptide with peptide-ribose-5-phosphate (R5P) modification at R139 and R426 from in vivo experiments. Sufficient peptide sequence coverage of the manually validated spectra reveals that arginine is the only amino acid modified in vivo. ADPr evaded LC-MS detection, but the inventors reliably and clearly identified ribose-5-phosphate (R5P), m / z = 212.0086Th, as a shorter fragment of ADPr. R5P-linked arginine residues are enclosed in yellow boxes. [Figure 9-2] Figure 9: LC-MS2 spectrum of rS1 peptide with peptide-ribose-5-phosphate. LC-MS2 spectrum of rS1 peptide with peptide-ribose-5-phosphate (R5P) modification at R139 and R426 from in vivo experiments. Sufficient peptide sequence coverage of the manually validated spectra reveals that arginine is the only amino acid modified in vivo. ADPr evaded LC-MS detection, but the inventors reliably and clearly identified ribose-5-phosphate (R5P), m / z = 212.0086Th, as a shorter fragment of ADPr. R5P-linked arginine residues are enclosed in yellow boxes. [Figure 10-1] Figure 10: Characterization of ADP-ribosylation and RNAization of R139 of rS1 in vivo. The peptide AFLPGSLVDVRPVRTHLEGK isolated from in vivo samples has R5P modification at R139, as indicated by the yellow box. The peptide sequence and modification site were reliably determined even though the peptide was longer due to erroneous cleavage at R142. The peptide AFLPGSLVDVAPVRTHLEGK identified from rS1 R139A or R139K mutants does not have R5P modification at position 139. [Figure 10-2]Figure 10: Characterization of ADP-ribosylation and RNAization of R139 of rS1 in vivo. The peptide AFLPGSLVDVRPVRTHLEGK isolated from in vivo samples has R5P modification at R139, as indicated by the yellow box. The peptide sequence and modification site were reliably determined even though the peptide was longer due to erroneous cleavage at R142. The peptide AFLPGSLVDVAPVRTHLEGK identified from rS1 R139A or R139K mutants does not have R5P modification at position 139. [Figure 11] Figure 11: In vivo characterization of ADP-ribosylation and RNAization by Western blot analysis. a. Analysis of substrate specificity of pan-ADPr antibody. In vitro prepared ADP-ribosylated or RNAized protein rS1 was applied to evaluate antibody characterization. b. Quantification of RNAization using a combination of nuclease P1 digestion and detection of protein-linked ADP-ribose by Western blotting. Visualization of protein packing by TCE staining. Removal of ADP-ribose signal by ARH1 treatment. Corresponding bar graphs are shown in Figure 4. [Figure 12-1] Figure 12: In vitro ADP-ribosylation and RNAization of rS1 domains D1-D6 and the S1 motif of PNPase by ModB. a, Schematic representation of the crystal structures (PDB) of the rS1 motif, domains 1 (2MFI), 2 (2MFL), 4 (2KHI), 5 (5XQ5), and 6 (2KHJ) of rS1, and the NMR structure of domain 3.1 b, Alignment of D2 and D6 of rS1 and the S1 domain of PNPase using T-coffee expresso.2 R139 of D2 is highlighted with an arrow. [Figure 12-2] c, ADP-ribosylation and [Figure 12-3]d. RNAization experiments were performed in triple replication and analyzed by 16% tricine SDS-PAGE (L=ladder). ModB and S1 domains are indicated by black arrows. RNAized rS1 domains, characterized by significant shifts compared to unmodified proteins, are highlighted with red arrows. n=2 biologically independent replicates. Reactions were performed using 32P-NAD or 32P-NAD-RNA 8-mers as substrates for ModB. [Figure 13-1] Figure 13: Characterization of the effect of R139 of rS1 domain 2 on ADP-ribosylation and RNAization. a, Analysis of ADP-ribosylation of rS1 domain 2 and its mutants R139A and R139K by 16% tricine-SDS-PAGE. b, Quantification of the relative intensity of ADP-ribosylation of rS1 domain 2 and its mutants R139A and R139K. Two biologically independent replicates. [Figure 13-2] c. Analysis of RNAization of rS1 domain 2 and its variants R139A and R139K by 16% tricine-SDS-PAGE. Inactive versions of NudC V157A, E174A, E177A, and E178A (NudC*) were used. d. Quantification of the relative intensity of RNAization of rS1 domain 2 and its variants R139A and R139K. Two biologically independent replicates were used. [Figure 14-1] Figure 14: In vivo characterization of ADP-ribosylation and RNAization. a, Time course of ADP-ribosylation in T4-infected (+T4) or uninfected (-T4) E. coli with chromosomal fusion of Flag-tag to rS1, analyzed by Western blotting. In vitro prepared rS1-ADPr served as a positive control, n=4. b, Western blotting to characterize the abundance of FLAG-rS1 during bacteriophage T4 infection in the presence of ARH1 WT and inactive ARH1 D55,56A. Equivalent expression of ARH1 WT or ARH1 D55,56A was verified by Western blotting using His-tag specific antibodies. Overexpression of ARH1 WT resulted in a significant decrease in pan-ADPr signaling. [Figure 14-2]c. Quantification of FLAG-rS1 levels in T4 phage-infected Escherichia coli overexpressing ARH1 WT or inactive ARH1 D55,56A. FLAG-rS1 levels were determined by Western blotting, as shown in b. [Figure 15] Figure 15: RNAization of rS1 by ModB using 3'Cy5-labeled NAD-cap RNA in vitro. Enzyme kinetics of ModB were performed in the presence of a 3'-Cy5-labeled NAD-cap-10 mer. rS1 was used as a target for ModB, which fluoresces (Cy5) during RNAization. Fluorescence signals were visualized using a Typhoon scanner. Samples were analyzed by 12% SDS-PAGE. [Figure 16] Figure 16: Preparation of 5'PX-10mer-Cy5 RNA. A) 5'-PX-RNA was incubated at 50°C for 5 hours in the presence of 1000-fold excess Im-NMN and 50 mM MgCl2. 5'-NXD-RNA was prepared by coupling Im-NMN to the 5'-monophosphate group as described in (23). B) Analysis of NXD-capping by APB gel electrophoresis. 5'-PX-RNA served as a negative control (n=1). C) Comparison of calculated yields of NXD-RNA capping reactions. [Figure 17] Figure 17: RNAization reactions of rS1 and rS1 DII by ModB in the presence of 5'-NXD-RNA. A) Shows the proposed RNAization reaction mechanism for the rS1 protein by ModB. RNAized rS1 was produced by incubating the rS1 protein with each 5'-NXD-RNA in the presence of ModB. B) Shows the RNAization reaction of rS1 DII by ModB. In the presence of each 5'-NXD-RNA, the entire RNA strand was covalently ligated to rS1 DII to produce RNAized rS1 DII. C) Relative RNAization efficiency of rS1 and rS1 DII using different NXD-RNAs as substrates for ModB. [Figure 18]Figure 18: In vitro RNAization of rS1 and rS1 DII in the presence of differently capped RNAs by ModB. A, B) RNAization reactions of rS1 in the presence of 5'-PX-10-mer-Cy5 or 5'-NXD-10-mer-Cy5 RNA were analyzed by 12% SDS-PAGE. RNAized rS1 was detected as a shifted band in the presence of 5'-NXD-RNA and ModB (n=2). C, D) RNAization reactions of rS1 DII using 5'-PX-10-mer-Cy5 or 5'-NXD-10-mer-Cy5 RNA were analyzed by 15% trichine electrophoresis. Shifted RNAized rS1 DII was observed in the presence of 5'-NXD-RNA and ModB (n=2). [Figure 19] Figure 19: In vitro ARH1 digestion kinetics of RNA-modified rS1 using differently capped RNAs. RNA-modified ADPr-RNA-rS1 (A), GDPr-RNA-rS1 (B), CDPr-RNA-rS1 (C), or UDPr-RNA-rS1 (D) proteins were subjected to ARH1 digestion for 0, 2, 5, 10, 30, 60, 120, and 180 minutes. Reactions were analyzed by 12% SDS-PAGE. E) Calculated mean relative RNA-modification levels during ARH1 treatment, N=2. F) Schematic representation of the mechanism by which ARH1 removes XDPr-RNA from RNA-modified (XDPr-RNA)-rS1. [Figure 20] Figure 20: In vitro RNAization of rS1 and rS1 DII by ModB in the presence of differently capped DNA. A, B) RNAization reactions of rS1 in the presence of 5'-PX-10mer(DNA)-Cy5 or 5'-NXD-10mer(DNA)-Cy5 RNA were analyzed by 12% SDS-PAGE. RNAized rS1 was detected as a shifted band in the presence of 5'-NXD-DNA and ModB (n=2). (N=3). [Figure 21] Figure 21: A) Relative RNAization efficiency of rS1 using different NXD-DNAs as ModB substrates (N=3). [Figure 22]Figure 22: In vitro ARH1 digestion kinetics of RNA-modified rS1 using differently capped DNA. RNA-modified dADPr-DNA-rS1, dGDPr-DNA-rS1, dCDPr-DNA-rS1, or dUDPr-DNA-rS1 proteins were subjected to ARH1 digestion for 0, 2, 5, 10, 30, 60, 120, and 180 minutes. Reactions were analyzed by 12% SDS-PAGE, and the mean relative RNAization levels during ARH1 treatment were calculated. [Modes for carrying out the invention]

[0133] The present invention will be explained by examples. [Examples]

[0134] Example 1 - T4 ART catalyzes RNAization in vitro. To test the hypothesis that ART can accept NAD-RNA as a substrate, three T4 ARTs were purified and their synthesis site-specific properties were investigated. 32 The proteins were incubated with 8-mers of 5'-NAD-RNA labeled with 3P and tested for either self-modification or modification of the target protein. Modifications were identified by either ART or the target protein, respectively. 32 This is shown by the acquisition of P labeling. Both Alt and ModA showed only low levels of self and target RNAization (Figure 5a), but ModB rapidly RNAized its known target ribosomal protein S1 (rS1) without detectable self-RNAization, as indicated by the radioactive band with expected mobility in the SDS-PAGE gel. In contrast, 32 ADP-ribosylation in the presence of P-NAD resulted in modifications of both proteins (ModB and rS1) with similar radioactive band intensities (Figures 2a, b and 5b, c). The radioactive bands were observed when either ModB or rS1 was deficient or when the 5'- of the same sequence was present. 32 P-monophosphate-RNA(5'- 32 This was not observed when P-RNA was used as a substrate for ModB (Figure 5d).

[0135] Example 2 - ModB prefers NAD-RNA over NAD. ModB-catalyzed RNAization of rS1 was strongly inhibited by the ART inhibitor 3-methoxybenzamide (3-MB) (Figure 6a). The radioactive rS1 band did not disappear when the reaction product was treated with Rnase T1. This treatment was performed when the RNA was non-covalently bound to rS1 or covalently linked via a position other than the 5' terminal. 32 The P label was removed (Figure 6b). NudC is a bacterial enzyme that hydrolyzes pyrophosphate bonds in various non-canonical cap structures. 21 This resulted in a 53% decrease in the radioactive signal (Figure 6c), indicating the generation of ribose-5'-phosphate-modified rS1. However, the radioactive band disappeared entirely upon treatment with trypsin (which digests rS1) (Figure 6c). In summary, these data strongly support covalent linking of RNA to rS1 via diphosphoriboside ligation, as shown in Figure 1b.

[0136] 32 Competitive experiments using P-NAD-RNA and excess unlabeled NAD revealed ModB's preference over the former, which is crucial for the modification reaction in vivo, and where NAD is considerably more abundant than NAD-RNA (Figures 2c and 7a). ModB also exhibits comparable activity with longer, biologically relevant RNAs (e.g., Qβ-RNA fragments of approximately 100 nt). 22 The proteins were received as shown in Figures 2d and 7b). RNAization using this NAD-cap-Qβ-RNA caused the protein rS1 (approximately 70 kDa) to move on an SDS-PAGE gel in the same way as a 100 kDa protein (Figure 2e). This shift was reversed by treatment of the RNAized protein with nuclease P1, which hydrolyzes the 3'-5' phosphodiester bond but does not attack the pyrophosphate bond of 5'-ADP-ribose. The radioactive product then moved in the same way as unmodified rS1 or ADPr-rS1 (Figure 2e), reaffirming the proposed properties of covalent linkage.

[0137] ModB eliminates the possibility that hydrolysis could remove just the nicotinamide moiety from NAD-RNA, creating a highly reactive ribosyl moiety that could spontaneously react with a nucleophile nearby (via its coated aldehyde group). 23 , true ADP-ribose-modified RNA (site-specific) 32 A 3P-labeled sample was prepared and tested as a substrate. No radioactive band was observed (Figure 7c), and no support for spontaneous ADP-ribosylation was provided.

[0138] Example 3 - ModB modifies specific arginine in rS1. To identify the amino acid residues in the protein rS1, where RNA chains are covalently linked during RNAization, we utilized a tool developed for analyzing protein ADP-ribosylation. The radioactive signal of the RNAized protein rS1 (prepared in Figure 2b) remained unchanged during treatment with HgCl2 (cleaving S-glycoside residues generated by Cys), NH2OH (hydrolyzing O-glycosides) (Figure 8a), and recombinant enzyme ARH3 (specifically hydrolyzing O-ADPr glycosides at serine residues) (Figure 8b), and was effectively removed by treatment with human ARH1 (Figures 3a-d). 24 These findings indicate that the main product(s) of the ModB-catalyzed RNAization reaction are linked as N-glycosides via arginine residues (similar to those shown in Figures 3a and 3b).

[0139] To identify the amino acid residues targeted by ModB, in vitro modified rS1 was subjected to trypsin digestion, chromatographic purification, and mass spectrometry. This LC / MS / MS analysis revealed three specific modification sites in rS1: R19, R139, and R426 (Figure 9).

[0140] To establish the biological significance of T4 ART-mediated RNAization in vivo, the (untagged) protein rS1 was endogenously isolated from both uninfected and T4-infected E. coli. E. coli possesses significant amounts of endogenous NAD-RNA. 4、6 Ribosomes were isolated, rS1 was extracted using poly-U-Sepharose, and subjected to LC / MS / MS analysis (Figure 3e). This experiment confirmed in vitro data and identified the same three sites, R19, R139, and R426, that were rich in phosphoribose modification only in T4-infected samples (Figure 3f). Site-induced mutagenesis further identified modified residues: R139K and R139A mutants of protein rS1 were expressed in T4-infected E. coli, purified, and analyzed to demonstrate that these mutations inhibit modification (Figure 10).

[0141] Example 4 - Detection of RNAization in vivo ADP-ribosylation and RNAization were detected using the same method via a mass spectrometer pipeline, i.e., as ribose-5'-phosphate or ADPr fragments. To distinguish between the two modifications, an immunoblotting assay using an antibody-like ADP-ribose conjugate ("pan-ADPr") was considered. The specificity of pan-ADPr was investigated by Western blotting with in vitro prepared ADP-ribosylated or RNAized proteins, respectively (Figure 11a). As expected, both rS1-ADPr and ModB-ADPr were recognized by pan-ADPr, producing bands of high intensity, but no signal was observed for rS1-RNA, suggesting that pan-ADPr does not tolerate 3'-extension of the ribose moiety. However, when rS1-RNA was digested with nuclease P1 before pan-ADPr treatment, thereby degrading the RNA and leaving rS1-ADPr, a strong signal comparable to that of true rS1-ADPr was observed in the blot (Figure 4a).

[0142] This immunoblotting assay was applied to investigate in vivo ADP-ribosylation and RNAization. Plasmid-carrier copies of rS1 were applied to uninfected or T4-infected E. coli. Subsequently, rS1 was affinity-purified, and its ADP-ribosylation was analyzed by pan-ADPr blotting (data Figure 11b). Consistent with our mass spectrometer data, this experiment revealed widespread ADP-ribosylation of rS1 only in T4-infected samples. After nuclease P1-treatment, the pan-ADPr signal intensity of the rS1 band significantly increased (Figure 4b), indicating that approximately 30% of modified rS1 was RNAized in vivo (measured as the difference between P1-treated and nuclease-untreated samples). Furthermore, the signal for ADP-ribose disappeared upon ARH1 treatment, reconfirming the RNA-protein linkage properties (Figure 11b).

[0143] Example 5 - Recognition motif for ModB How ModB identifies its targets remains a challenge. The target protein rS1 contains an oligonucleotide-binding (OB) domain. 22 One structural variant of OB folding is the S1 domain present in rS1 in six sequences with different sequences (Figure 12a, b). The S1 domain was hypothesized to be important for substrate recognition by ModB. To characterize the specificity of MobB to different rS1 domains, each S1 domain position of rS1 was individually cloned, expressed, purified, and applied in RNAization assays (Figure 4c, d and Figure 12c, d). For the rS1 domain, high RNAization signals of D2 and D6 were determined. For comparison, the rS1 D1, D3, D4, and D5 domains were modified to a considerably lower degree. Alignment of D2 and D6 of rS1, as well as the S1 domain of PNPase and another protein with an S1 domain in E. coli, revealed that these S1 motifs share an arginine residue as part of a loop linking β-barrel chains 3 and 4. 25(Figure 12b). This loop is packed at the top of the β-barrel, thereby making it accessible to ModB as well. For rS1 D2, this particular residue is R139, which was shown to be modified by mass spectrometry (Figure 3f). Mutagenesis analysis confirmed that the ADP-ribosylation level of D2 was dramatically reduced when R139 was substituted with alanine or lysine (Figure 13). Based on these findings, other E. coli proteins possessing an S1 domain with arginine in the loop between chains 3 and 4 were screened, and Rnase E was identified. In our in vitro assay, Rnase E with the S1 motif in its active site was effectively modified by ModB, while control proteins without the S1 domain (BSA, NudC inactive 4x mutant) were not modified, supporting the identification of a subgroup of S1 domains with arginine embedded as an RNA target motif (Figure 4e, f).

[0144] Example 6 - Modification and T4 Replication Cycle rS1 is a crucial RNA-binding protein required for the translation of virtually all cellular mRNA in E. coli. To investigate the biological consequences of rS1 modification by ModB, rS1 levels were analyzed during T4 infection using E. coli strains containing rS1 and FLAG-tag chromosome fusions (Figure 4g and Figure 13a). Immediately after infection, rS1 levels rapidly decreased, and they moderately increased over 20 minutes in the absence of T4. Therefore, it was hypothesized that ADP-ribosylation and / or RNAization may affect rS1 stability. To test this hypothesis, human ARH1, which is thought to remove ADP-ribose and ligating RNA, was overexpressed in E. coli during T4 infection. As a control, large, inactive ARH1 D55, 56A mutants were overexpressed. Indeed, with active ARH1, the ADP-ribosylation signal was dramatically reduced (Figure 14b), and the mutants showed a similar pattern to the parental strain (Figure 14a). Using these constructed E. coli strains, the effects of ADP-ribosylation and RNAization on rS1 levels were analyzed during phage infection. Indeed, strains expressing active ARH1 showed an increase in rS1 levels over time, similar to uninfected samples, while mutant strains showed decreased levels, similar to T4-infected samples lacking ARH1 (Figure 14b, c). Therefore, the removal of ADPr and RNA strands during phage infection occurs concurrently with the stabilization of rS1 levels.

[0145] To investigate whether these modifications are important for the phage lysogenic behavior, E. coli strains expressing either ARH1 or an inactive variant were infected with T4 cells, and optical density was monitored over time (Figure 4i). In the inactive variant strain, bacterial lysis began 50 minutes after infection, and a delay in lysis (120 minutes) was observed when active ARH1 was overexpressed (Figure 4i). In summary, these data suggest that ADP-ribosylation and / or RNAization impair protein stability and modulate the pathway and efficiency of T4 infection.

[0146] Example 7 - RNAization of proteins using NND (=NXD)-cap RNA and DNA This example demonstrates that ModB accepts 5'-NGD-, NCD-, or NUD-capped RNA as a substrate for RNAization reactions, in addition to 5'-NAD-RNA. The exchange of RNA-caps from NAD to NGD, NCD, or NUD does not alter ModB's catalytic activity. This finding indicates that ModB's catalytic pocket does not sense the adenosine moiety of NAD. In contrast, the nicotinamide moiety may be decisive for substrate recognition by ModB. Furthermore, the application of naturally occurring NGD-, NCD-, or NUD-RNA caps allows for a flexible and adaptable design of 5'-NXD-RNA as a substrate for RNAization reactions. Thus, ModB's target proteins can be RNAized with any preferred RNA sequence. Finally, this example demonstrates that GDPr-, UDPr-, and CDPr-linked RNAs are not removed by human ADP-ribose hydrolase ARH1, thereby exhibiting increased RNA-protein stability. These properties lay the foundation for creating novel in vitro RNA-protein conjugates that may be applicable to eukaryotic systems in vivo in the future.

[0147] 7.1 The imidazolide reaction resulted in effective 5'-NXD-capping of all monophosphorylated RNAs. Compared to NAD-RNA, which is explained in all biological kingdoms, NUD-RNA, NCD-RNA, and NGD-RNA are not yet explained in biological systems. Therefore, to verify whether NXD-capped RNA can be applied as a substrate for RNAization, they were synthesized chemically. Here, 5'-NXD-capping of 5'-monophosphorylated RNA was achieved using an imidazolide reaction by coupling Im-NMN to the 5'-monophosphate group of RNA (Figure 16A). To investigate the capping reaction efficiency, the reaction products were characterized by APB gel electrophoresis (Figure 16B). The calculated 5'-NXD-RNA yields indicated that the capping efficiency of the reaction ranged from 42.8% corresponding to NUD capping to 66.2% observed for NAD capping. For NGD and NCD, capping of 52.4% and 45.0% were calculated, respectively (Figure 16C). The capping efficiencies were consistent with previous reports.

[0148] 7.2 ModB accepts 5'-NXD-cap RNA as a substrate for the RNAization reaction. The successful preparation of NXD-cap RNA made it possible to test the substrate range of ModB. It was hypothesized that all tested 5'-NXD-cap RNAs could be accepted by ModB for RNAization.

[0149] Figures 17A and 17B show the proposed mechanisms of rS1 and rS1 DII RNAization reactions. In the presence of 5'-NXD-10mer-Cy5 RNA, ModB can covalently ligate its entire RNA chain to the target protein via the RNAization reaction.

[0150] To investigate whether NXD-cap RNA could be applied as a novel substrate for ModB, in vitro RNAization reactions were performed (Figures 17C and 18). RNAization reactions were performed using 5'-PX-RNA (negative control) and 5'-NXD-RNA in the presence and absence of ModB. NAD-RNA served as a positive control and reference for RNAization. The data indicate that RNAization of the rS1 protein by ModB was achieved regardless of the second nucleotide located within the cap structure. NGD-RNA, NCD-RNA, or NUD-RNA were identified as novel substrates accepted for RNAization by ModB. Furthermore, RNAization of rS1 using NXD-RNA altered protein size and caused changes in the current behavior of the modified protein compared to the unmodified protein (Figures 18A and 18B).

[0151] The calculated RNAization yields of rS1 in the presence of 5'-NGD-RNA or 5'-NUD-RNA were similar to those of RNAization using 5'-NAD-RNA. Surprisingly, the RNAization reaction using 5'-NCD-RNA yielded four times higher yields than that using 5'-NAD-RNA (Figure 3C). Assuming that rS1 does not covalently bind to RNA in a non-enzymatic manner, no RNAization was detected in the absence of ModB. Furthermore, 5'-PX-RNA was not accepted as a substrate by ModB, and the RNAization reaction occurred only in the presence of 5'-NXD-cap RNA (Figures 18A and 18B).

[0152] It can be shown that rS1 can be RNA-activated by ModB in the presence of 5'-NXD-RNA. The question was also raised as to whether other target proteins can be RNA-activated by ModB using NXD-RNA as a substrate. Therefore, the RNA-activated synthesis of another target protein, rS1 DII, by ModB was characterized in the presence of 5'-NXD-RNA. In contrast to the already studied rS1 (68 kDa), rS1 DII is a small protein with a molecular weight of 9.7 kDa.

[0153] Similar to the rS1 protein, rS1 DII can be shown to be RNAized by ModB in the presence of 5'-NXD-RNA. Furthermore, different size shifts of the RNAized protein can be observed (Figures 18C and 18D). The calculated RNAization efficiencies reveal the same trends as those described for rS1. Again, the highest RNAization of rS1 DII was observed in the presence of NCD-RNA (Figure 17C).

[0154] Therefore, the data indicate that both rS1 and rS1 DII were successfully RNAized in the presence of 5'-NXD-RNA and ModB. Furthermore, the RNAization efficiency did not differ among target proteins, meaning that various target proteins can be RNAized with the same efficiency regardless of their molecular weight.

[0155] 7.3 ARH1 specifically hydrolyzes the N-glycosidic bond of the ADP-ribosyl-arginine residue. In eukaryotes, ARH1 plays a major role in the removal of ADP-ribosylation. Therefore, the stability of in vitro prepared RNA-modified protein conjugates applied to eukaryotes depends on the enzymatic activity of ARH1. It was hypothesized that the exchange of ADP-ribose-RNA covalently bound to GDPr-RNA, CDPr-RNA, or UDPr-RNA would alter substrate recognition by ARH1.

[0156] To test whether covalently linked XDPr-RNA is removed by ARH1, rS1 proteonucleotides using 5'-NXD-RNA were digested in vitro with ARH1 (Figure 19). rS1 proteonucleotides using 5'-NAD-RNA (ADPr-RNA-rS1) served as a positive control for ARH1 digestion. ARH1 may be shown to effectively remove ADPr-RNA from ADPr-RNA-rS1. Relative RNAization levels decreased to 20% 30 minutes after ARH1 treatment (Figure 19A, E). In contrast, ARH1 was unable to remove GDPr-RNA, CDPr-RNA, or UDPr-RNA from RNAized GDPr-RNA-rS1, CDPr-RNA-rS1, or UDPr-RNA-rS1 proteins (Figure 19B-E). Compared to the hydrolysis of RNAization in the presence of ADPr-RNA, the reaction proceeds 40-fold and 20-fold slower in the presence of CDPr-RNA- and UDPr-RNA, respectively. Therefore, ARH1 specifically hydrolyzes the N-glycosidic bond of the ADP-ribosyl-arginine residue.

[0157] 7.4 RNAization of proteins using NND (=NXD)-capped DNA In vitro RNAization of rS1 and rS1 DII in the presence of DNA capped differently by ModB is shown in Figure 20. The RNAization reaction of rS1 in the presence of 5'-PX-10mer(DNA)-Cy5 or 5'-NXD-10mer(DNA)-Cy5 RNA was analyzed by 12% SDS-PAGE. RNAized rS1 was detected as a shifted band in the presence of 5'-NXD-DNA and ModB (n=2).

[0158] Furthermore, the relative RNAization efficiency of rS1 was tested using different NXD-DNAs as substrates for ModB (etsed) (Figure 21).

[0159] Finally, the in vitro ARH1 digestion kinetics of RNA-modified rS1 using differently capped DNA were analyzed (Figure 22). For this analysis, RNA-modified dADPr-DNA-rS1, dGDPr-DNA-rS1, dCDPr-DNA-rS1, or dUDPr-DNA-rS1 proteins were subjected to ARH1 digestion for 0, 2, 5, 10, 30, 60, 120, and 180 minutes, and the reactions were analyzed.

[0160] Example 8 - Discussion 8.1 Discussion of Examples 1-7 To date, all interactions between RNA and proteins have been described as being based on non-covalent interactions. 26 In contrast, this specification shows that ADP-ribosyltransferase can attach NND-cap RNA to target proteins in a cohesive manner. This finding indicates a different biological function of NND-caps to RNA in bacteria, namely RNA activation for enzymatic transfer to receptor proteins. RNAization of target proteins, a novel post-translational protein modification that plays a role in bacterial E. coli infection by bacteriophage T4, has been discovered. Our data show that T4 ART ModB modifies proteins having an S1 RNA-binding domain. Specific arginine residues modified were identified, thereby increasing the molecular weight and negative charge of the target protein and reliably causing significant changes in the properties and function of the modified protein. Post-translational modifications of key players in bacterial translation and transcription highlight the importance of known ADP-ribosylation and the novel RNAization reaction for bacteriophage pathogenicity. Introducing human ADP-ribosylhydrolase ARH1, which removes these modifications, into E. coli resulted in a significant delay in bacterial lysis during phage infection.

[0161] The reason ART attaches RNA to translation-related proteins may be that these RNAs help preferentially associate mRNA encoding phage proteins with ribosomes (e.g., through base pairing), thereby ensuring their biosynthesis. Similarly, the observation that RNAase E, a major player in RNA turnover in E. coli, is RNA-modified by ModB at its catalytic center may suggest that T4 phages, after transcriptional reprogramming by Alt and ModA, block RNA degradation in the host, ensuring a long half-life for phage mRNA. We are working on methods to identify RNAs that attach to target proteins, which will enable the elucidation of their biochemical mechanisms.

[0162] ART is not known to occur exclusively in bacteriophages; ADP-ribosylated proteins have been detected in hosts during infections by various viruses, such as influenza, coronavirus, and HIV. In addition to viruses that use ART as a weapon, the mammalian antiviral defense system applies host ART to inactivate viral proteins. Furthermore, mammalian ART and poly-(ADP-ribose) polymerase (PARP) are known to be crucial regulators of cellular pathways and interact with RNA. 27 Therefore, ARTs in different organisms can catalyze RNAization reactions, and RNAization can be expected to be a phenomenon of wide-ranging biological relevance.

[0163] Ultimately, RNAization can be considered both a post-translational protein modification and a post-transcriptional RNA modification. Our findings challenge established views on how RNA and proteins can interact with each other. The discovery of these novel RNA-protein conjugates marks a time when the structural and functional boundaries between different classes of biopolymers are gradually blurring. 28、29 .

[0164] 8.2 Further consideration of Example 7 In contrast to recently identified NAD-RNAs, NGD-, NCD-, or NUD-RNAs had not yet been discovered in biological systems. Therefore, 5'-NXD-capped RNAs were prepared by chemical synthesis using imidazolide reactions. In addition to earlier studies, this specification demonstrates that synthetic 3'-Cy5 labeled RNA can be used as a template for imidazolide reactions to prepare fluorescent NXD-capped RNA / DNA. The calculated capping efficiency for 5'-NXD-capped RNA was consistent with previous reports. Furthermore, the prepared 5'-NXD-RNAs were used to investigate the substrate specificity of ModB. In vitro RNAization reactions of rS1 and rS1 DII by ModB were performed in the presence of 5'-NXD-RNAs.

[0165] It was discovered that 5'-NXD-capped RNA / DNA is accepted as a substrate by ModB. Therefore, the RNAization reaction occurs regardless of the first base of the RNA. This means that A can be replaced with G, C, or U in the cap structure, and that capped RNA can also be used as a substrate for the RNAization reaction by ModB.

[0166] To date, the protein crystal structures of ModB and its substrate NAD are not available. For this reason, the substrate specificity of ModB remains unclear. The exchange of the RNA-cap from NAD to NGD, NCD, or NUD does not alter ModB's catalytic activity. This finding suggests that ModB's catalytic pocket does not detect the adenosine portion of NAD. In contrast, the nicotinamide portion may be decisive for substrate recognition by ModB. Therefore, it can be concluded that the essential requirement for RNA-modified substrate design is only the NMN portion of the NAD-RNA-cap.

[0167] Furthermore, the data herein demonstrate that 5'-NGD-RNA and 5'-NUD-RNA produced similar RNAization yields to the 5'-NAD-RNA used as a reference. Interestingly, the increased RNAization efficiency of ModB was identified in the presence of 5'-NCD-RNA.

[0168] Recently, it has been shown that naturally occurring RNAization affects the molecular properties of target proteins, such as molecular weight (Hoefer et al. (2021), bioRxiv, 2021.2006.2004.446905). Example 7 shows that covalent attachment of NGD-RNA, NCD-RNA, or NUD-RNA to the target proteins rS1 and rS1 DII increases protein size. In conclusion, the discovery of NXD-RNA as a novel substrate for ModB allows for the flexible design of RNA-oligos applied to RNAization reactions. RNAization substrates can be prepared by solid-phase synthesis or in vitro transcription. In particular, in vitro transcription reactions allow for the preparation of biologically relevant transcripts longer than 80 nucleotides. Here, G-start sites typically produce the high transcription yields required to prepare RNAization substrates such as NGD-RNA. Furthermore, our data show that higher RNAization yields can be achieved by using 5'-NCD-RNA as a substrate.

[0169] Furthermore, the stability of the XDPr-RNA-protein in the presence of human ARH1 was tested in Example 7. ARH1 is the only eukaryotic enzyme known to date to remove RNAization from target proteins in vivo. The catalytic activity of ARH1 in the presence of differently capped RNAs has not been previously tested. Example 7 shows that ARH1 is unable to effectively remove RNAization in the presence of GDPr-RNA, UDPr-RNA, and CDPr-RNA. In vitro dynamics data show that ARH1 strongly prefers arginine-linked ADPr-RNA as a substrate over GDPr-RNA, UDPr-RNA, and CDPr-RNA.

[0170] In conclusion, applying NXD-RNA as a substrate for protein RNAization improves the understanding of the substrate specificity of ModB and ARH1. ModB accepts all four different NXD-RNA derivatives as substrates, while ARH1 is highly specific to the hydrolysis of the N-glycosidic bond of ADP-ribosyl-arginine. As a result, the GDPr-RNA-rS1, UDPr-RNA-rS1, and CDPr-RNA-rS1 proteins increased stability in vitro in the presence of human ARH1. These properties set the basis for creating in vitro RNA-protein conjugates that may be applied to eukaryotic systems in vivo in the future.

[0171] Example 9 - Extended Data [Table 1] [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 3-1] [Table 3-2] [Table 4-1] [Table 4-2]

[0172] The following are examples of aspects of the present invention. Item 1 A method for attaching a 5'-nicotinamide nucleobase dinucleotide (NND) cap nucleic acid sequence to a fusion protein or complex, the method is (a) a heterogeneous fusion protein containing the target poly(peptide) to be fused to the tag, or (ii) a complex, comprising the step of contacting the complex with a 5'-NND cap nucleic acid sequence and ADP-ribosyltransferase (ART), wherein the protein is complexed with the tag under physiological conditions, and the 5'-NND cap nucleic acid sequence is covalently attached to the tag. The tag includes an ART recognition motif, and preferably the tag is (i) A sequence that is at least 80% identical to sequence number 1, assuming that the underlined Arg in the amino acid motif DVRPVRD (sequence number 7) is conserved and preferably sequence number 7 is conserved; (ii) A sequence that is at least 80% identical to sequence number 2, assuming that the underlined Arg in the amino acid motif DVRPVRD (sequence number 7) is preserved and preferably sequence number 7 is preserved; (iii) A sequence that is at least 80% identical to sequence number 3, assuming that the underlined Arg in the amino acid motif DVRPVRD (sequence number 7) is preserved and preferably sequence number 7 is preserved; (iv) A sequence that is at least 80% identical to sequence number 4, assuming that the underlined Arg in the amino acid motif LADGVEGYLRASEASRDRVE (Sequence ID: 8) is conserved and preferably sequence number 8 is conserved; (v) A sequence that is at least 80% identical to sequence number 5, assuming that the underlined Arg in the amino acid motif LADGVEGYLRASEASRDRVE (Sequence ID: 8) is preserved and preferably sequence number: 8 is preserved; or (vi) Assuming that the underlined Arg in the amino acid motif LADGVEGYLRASEASRDRVE (SEQ ID NO: 8) is preserved and preferably SEQ ID NO: 8 is preserved, then a sequence that is at least 80% identical to SEQ ID NO: 6. A method that includes or consists of. Section 2 The method according to claim 1, wherein the nucleobase of NND is a purine base or a pyrimidine base, preferably selected from adenine, guanine, cytosine, thymine, and uracil. Section 3 The method according to claim 1 or 2, wherein ART includes or consists of a sequence that is at least 80% identical to sequence number 9 or sequence number 10. Section 4 Before step (a), (a') A step of fusing the tag defined in item 1 to the target poly(peptide) to obtain a heterofusion protein containing the target poly(peptide) to be fused to the tag, or (a') The process of complexing the tag defined in item 1 with the target poly(peptide). A method including any of the methods described in items 1 to 3. Section 5 A fusion protein comprising a poly(peptide) intended to be fused to the tag defined in item 1, or a complex containing a poly(peptide) intended to be complexed with the tag defined in item 1. Section 6 The fusion protein or complex according to claim 5, wherein a nucleic acid sequence is covalently attached at its 5' end to a tag defined in claim 1 via a nicotinamide nucleobase dinucleotide (NND), preferably to the conserved Arg side chain of the tag. Section 7 A composition, preferably a pharmaceutical composition or a diagnostic composition, comprising a fusion protein or complex obtained by any of the methods described in items 1 to 4 and / or a fusion protein or complex described in item 5 or 6. Section 8 A method, fusion protein, complex, or composition according to any of the preceding items, wherein the nucleic acid in the nucleic acid sequence is RNA, DNA, PNA, morpholino, or LNA or a combination thereof, and preferably RNA. Section 9 A method, fusion protein, complex, or composition according to any of the preceding items, wherein the nucleic acid sequence is an antisense molecule siRNA or shRNA. Section 10 A method, fusion protein, complex, or composition according to any of the preceding items, wherein the nucleic acid sequence comprises a fluorescent label at its 3' end, preferably containing an Alexa Fluor or Cy dye, or containing biotin at its 3' end. Section 11 A method, fusion protein, complex, or composition according to any of the preceding items, wherein the target poly(peptide) is an antibody, antibody mimetic, cytokine, interleukin, transmembrane protein, membrane-anchored protein, enzyme, or DNA and / or RNA-binding protein. Section 12 A method, fusion protein, complex, or composition according to any one of the preceding items, wherein a poly(peptide) and a tag are fused via a peptide linker, preferably a G-linker or a GS-linker. Section 13 A kit for attaching a 5'-nicotinamide nucleobase dinucleotide (NND) cap nucleic acid sequence to a target (poly)peptide, (a) Tags as defined in Section 1, (b) ADP-ribosyltransferase (ART) that can co-attach a 5'-NND cap nucleic acid sequence to a tag or a nucleic acid molecule encoding ADP-ribosyltransferase (ART), and (c) Optionally, instructions on how to co-attach the ART-containing tag to the target (poly)peptide. A kit that includes this. Section 14 A kit further comprising a reaction buffer or buffer stock solution, wherein the final reaction buffer, preferably prepared from the reaction buffer or buffer stock solution, Mg(OAc)2 at concentrations of 50-200 mM; NH4Cl at concentrations of 100-500 mM; Trisacetic acid at a concentration of 250-1000 mM, pH 7.5 EDTA at a concentration of 5-15 mM; β-mercaptoethanol at concentrations of 50-200 mM; and Glycerol at a concentration of 5-15% The kit described in item 13, including the kit described in item 13. Item 15 MgCl2 at a concentration of at least 0.25 M, preferably 0.5 M to 2 M, Imidazolide nicotinamide mononucleotide (Im-NMN), Nuclease-free water, and A control fusion protein comprising a control poly(peptide) fused to or complexed with a positive control, preferably an oligonucleotide and / or a tag as defined in claim 1, which has a fluorescent label at its 3' end. A kit as described in item 13 or 14, further comprising one or more of the following. References Table 5-1 Table 5-2

Claims

1. A method for attaching a 5'-nicotinamide nucleonucleotide (NND) cap nucleic acid to a fusion protein or complex, the method is (a) a fusion protein containing the target poly(peptide) fused to a tag, or (ii) a complex in which the protein is complexed with the tag under physiological conditions, comprising the step of contacting the 5'-NND cap nucleic acid with ADP-ribosyltransferase (ART), wherein the 5'-NND cap nucleic acid is covalently attached to the tag. The tags include the ART recognition motif, and the tags are, (i) A sequence that is at least 95% identical to sequence number 1, provided that the underlined Arg in the amino acid motif DVRPVRD (sequence number 7) is conserved; (ii) A sequence that is at least 95% identical to sequence number 2, provided that the underlined Arg in the amino acid motif DVRPVRD (sequence number 7) is conserved; (iii) A sequence that is at least 95% identical to sequence number 3, provided that the underlined Arg in the amino acid motif DVRPVRD (sequence number 7) is conserved; (iv) A sequence that is at least 95% identical to sequence number 4, provided that the underlined Arg in the amino acid motif LADGVEGYLRASEASRDRVE (sequence number 8) is conserved; (v) A sequence that is at least 95% identical to sequence number 5, provided that the underlined Arg in the amino acid motif LADGVEGYLRASEASRDRVE (sequence number 8) is conserved; or (vi) A sequence that is at least 95% identical to sequence number 6, provided that the underlined Arg in the amino acid motif LADGVEGYLRASEASRDRVE (sequence number 8) is conserved. including or consisting of Here, the nucleic acid base of the NND is a purine base or a pyrimidine base, selected from adenine, guanine, cytosine, thymine, and uracil. method.

2. The method according to claim 1, wherein ART includes or consists of sequence number 9 or sequence number 10, or a sequence that is at least 95% identical to them.

3. Before step (a), (a') A step of fusing the tag defined in claim 1 to the target poly(peptide) to obtain a fusion protein containing the target poly(peptide) fused to the tag, or (a') A step of obtaining a complex by complexing the tag defined in claim 1 with a protein. The method according to claim 1, including the method described in claim 1.

4. A fusion protein comprising a target poly(peptide) fused to a tag as defined in claim 1, or a complex in which a protein is complexed with a tag as defined in claim 1 under physiological conditions, wherein the nucleic acid is covalently attached to the tag via a nicotinamide nucleotide dinucleotide (NND) at its 5' end, wherein the nucleic acid base of the NND is a purine base or a pyrimidine base, selected from adenine, guanine, cytosine, thymine, and uracil. A fusion protein or complex.

5. The fusion protein or complex according to claim 4, wherein a nucleic acid is covalently attached to the conserved Arg side chain of the tag via a nicotinamide nucleoside dinucleotide (NND) at its 5' end.

6. A pharmaceutical composition or diagnostic composition comprising the fusion protein or complex according to claim 4 or 5.

7. The method according to claim 1, wherein the nucleic acid is RNA, DNA, PNA, morpholino, or located nucleic acid or a combination thereof.

8. The fusion protein or complex according to claim 4 or 5, wherein the nucleic acid is RNA, DNA, PNA, morpholino, or located nucleic acid or a combination thereof.

9. The composition according to claim 6, wherein the nucleic acid is RNA, DNA, PNA, morpholino, or located nucleic acid or a combination thereof.

10. The method according to claim 1, wherein the nucleic acid is an antisense molecule and is siRNA or shRNA.

11. The fusion protein or complex according to claim 4 or 5, wherein the nucleic acid is an antisense molecule and is siRNA or shRNA.

12. The composition according to claim 6, wherein the nucleic acid is an antisense molecule and is siRNA or shRNA.

13. The method according to claim 1, wherein the nucleic acid contains a fluorescent label at its 3' end or contains biotin at its 3' end.

14. The fusion protein or complex according to claim 4 or 5, wherein the nucleic acid contains a fluorescent label at its 3' end or contains biotin at its 3' end.

15. The composition according to claim 6, wherein the nucleic acid contains a fluorescent label at its 3' end or contains biotin at its 3' end.

16. The method according to claim 1, wherein the target poly(peptide) is an antibody, cytokine, interleukin, transmembrane protein, membrane-anchored protein, enzyme, or DNA and / or RNA-binding protein.

17. The fusion protein or complex according to claim 4 or 5, wherein the target poly(peptide) is an antibody, cytokine, interleukin, transmembrane protein, membrane anchor protein, enzyme, or DNA and / or RNA-binding protein.

18. The composition according to claim 6, wherein the target poly(peptide) is an antibody, cytokine, interleukin, transmembrane protein, membrane anchor protein, enzyme, or DNA and / or RNA-binding protein.

19. The method according to claim 1, wherein the target poly(peptide) and tag are fused via a peptide-linker.

20. The fusion protein or complex according to claim 4 or 5, wherein the target poly(peptide) and tag are fused via a peptide-linker.

21. The composition according to claim 6, wherein the target poly(peptide) and tag are fused via a peptide-linker.

22. A kit for attaching a 5'-nicotinamide nucleonucleotide (NND) cap nucleic acid to a target (poly)peptide, (a) Tag as defined in claim 1, (b) ADP-ribosyltransferase (ART) capable of covalently attaching a 5'-NND cap nucleic acid to a tag, or a nucleic acid molecule encoding the ART, and (c) Instructions on how to attach the tag to the target (poly)peptide via covalent bond using ART. including, Here, the nucleic acid base of the NND is a purine base or a pyrimidine base, selected from adenine, guanine, cytosine, thymine, and uracil. kit.

23. A kit further comprising a reaction buffer or buffer stock solution, wherein the final reaction buffer prepared from the reaction buffer or buffer stock solution is Mg(OAc) at concentrations of 50-200 mM 2 ; NH4 at concentrations of 100–500 mM 4 Cl; Trisacetic acid at a concentration of 250–1000 mM, pH 7.5 EDTA at concentrations of 5–15 mM; β-mercaptoethanol at concentrations of 50–200 mM; and Glycerol at a concentration of 5-15% The kit according to claim 22, comprising:

24. MgCl at a concentration of at least 0.25 M 2 , Imidazolide nicotinamide mononucleotide (Im-NMN), Nuclease-free water, and A control fusion protein comprising a control poly(peptide) fused to or complexed with an oligonucleotide and / or a tag as defined in claim 1, wherein the positive control has a fluorescent label at its 3' end. The kit according to claim 22 or 23, further comprising one or more of the following.