Enzyme for RNA capping and reaction system thereof

By using T4 RNA ligase or its fusion enzyme for RNA capping reaction, the steps are simplified and efficiency and consistency are improved, and the problems of cumbersome steps and difficulty in purification in traditional methods are solved, achieving efficient RNA capping.

WO2025180432A1PCT designated stage Publication Date: 2025-09-04PIXEL BIOSCIENCES (SUZHOU) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/079459
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-28
Filing Date
2025-02-27
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

In the prior art, the enzymatic capping method has cumbersome steps, is difficult to purification, and is inefficient and costly to achieve efficient and high specific RNA capping.

Method used

T4 RNA ligase or its fusion enzyme is used for capping reaction, and RNA ligase is used for sequence ligation, which simplifies the steps and improves capping efficiency and reduces the difficulty of purification.

Benefits of technology

An efficient and simplified RNA capping process is achieved, which improves capping efficiency and consistency, and reduces purification difficulty and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2025079459-FTAPPB-I100001
    Figure PCTCN2025079459-FTAPPB-I100001
  • Figure PCTCN2025079459-FTAPPB-I100002
    Figure PCTCN2025079459-FTAPPB-I100002
  • Figure PCTCN2025079459-FTAPPB-I100003
    Figure PCTCN2025079459-FTAPPB-I100003
Patent Text Reader

Abstract

A method for capping a polynucleotide, comprising: S1, providing a composition, which at least comprises: an uncapped polynucleotide, a capping enzyme, and a cap structure analog, the capping enzyme comprising a polynucleotide ligase polypeptide; and S2, forming a capped polynucleotide on the basis of the composition of S1.
Need to check novelty before this filing date? Find Prior Art

Description

An enzyme for RNA capping and its reaction system Technical Field

[0001] The present application relates to the field of biomedicine, and specifically to a method for capping a polynucleotide. Background Art

[0002] After mRNA is transcribed, a cap structure needs to be further added to the 5' end to increase stability and improve translation efficiency. The most commonly used capping methods are enzymatic capping and co-transcriptional capping. Traditional enzymatic capping requires purification after the initial transcript is synthesized, and then a capping reaction is performed. The capping process is cumbersome and requires three enzymatic reactions. There are many intermediate products in the capping reaction, which is difficult to purify and has poor consistency. Co-transcriptional capping uses chemically synthesized cap analogs, which is relatively expensive. Capping is completed simultaneously during the transcription process, but the capping efficiency is low, and a large number of non-target products are obtained. Therefore, it is necessary to develop an efficient and specific RNA capping method. Summary of the Invention

[0003] The present application directly uses a ligase (e.g., T4 RNA ligase 1) or a fusion enzyme containing a ligase (e.g., a fusion enzyme of Sso7d and T4 RNA ligase 1) to perform a capping reaction. The use of RNA ligase for sequence ligation is highly efficient, and RNA fragments can be ligated in a short time. The capping efficiency is higher than that of the traditional three-step capping method, and the steps are simple, the purification difficulty is low, and the consistency is high.

[0004] In one aspect, the present application provides a method for capping a polynucleotide, comprising:

[0005] S1. Provide a composition comprising at least: an uncapped polynucleotide, a capping enzyme, and a cap structure analog, wherein the capping enzyme comprises a polynucleotide ligase polypeptide;

[0006] S2. Forming a capping polynucleotide according to the composition of S1.

[0007] In some embodiments, the polynucleotide ligase polypeptide comprises a DNA ligase polypeptide.

[0008] In some embodiments, the DNA ligase polypeptide is a prokaryotic DNA ligase, a prokaryotic DNA ligase variant, or a functional fragment thereof.

[0009] In some embodiments, the DNA ligase polypeptide is a bacterial DNA ligase, a bacterial DNA ligase variant, or a functional fragment thereof.

[0010] In some embodiments, the DNA ligase polypeptide is a viral DNA ligase, a viral DNA ligase variant, or a functional fragment thereof.

[0011] In some embodiments, the DNA ligase polypeptide is T4 DNA ligase, a variant thereof, or a functional fragment thereof.

[0012] In some embodiments, the polynucleotide ligase polypeptide comprises an RNA ligase polypeptide.

[0013] In some embodiments, the RNA ligase polypeptide is a prokaryotic RNA ligase, a prokaryotic RNA ligase variant, or a functional fragment thereof.

[0014] In some embodiments, the RNA ligase polypeptide is a bacterial RNA ligase, a bacterial RNA ligase variant, or a functional fragment thereof.

[0015] In some embodiments, the RNA ligase polypeptide is a viral RNA ligase, a viral RNA ligase variant, or a functional fragment thereof.

[0016] In some embodiments, the RNA ligase polypeptide is T4 RNA ligase, a variant thereof, or a functional fragment thereof.

[0017] In some embodiments, the RNA ligase polypeptide is T4 RNA ligase.

[0018] In some embodiments, the RNA ligase polypeptide is T4 RNA ligase 1.

[0019] In some embodiments, the capping enzyme is a fusion polypeptide comprising a polynucleotide binding polypeptide fused to a polynucleotide ligase polypeptide.

[0020] In some embodiments, the polynucleotide binding polypeptide comprises a DNA binding polypeptide.

[0021] In some embodiments, the polynucleotide binding polypeptide comprises an RNA binding polypeptide.

[0022] In some embodiments, the polynucleotide binding polypeptide is selected from one or more of a DNA double-strand binding domain, a DNA single-strand binding domain, an RNA / DNA composite chain binding domain, an RNA single-strand or RNA double-strand binding protein, wherein

[0023] The DNA double-strand binding domain includes Sso7d and / or NF-kappaB p50;

[0024] The RNA / DNA complex chain binding domain includes one or more of ScFV, RNaseH1 (D210N), and HBD of the monoclonal antibody S9.6;

[0025] RNA single-strand binding domains include: Sso7d, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4.

[0026] In some embodiments, the polynucleotide binding polypeptide includes one or more of Sso7d, NF-kappaB p50, ScFV of monoclonal antibody S9.6, RNaseH1 (D210N), HBD, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4.

[0027] In some embodiments, the polynucleotide binding polypeptide is Sso7d.

[0028] In some embodiments, the C-terminus of the polynucleotide ligase polypeptide is linked to the N-terminus of the polynucleotide binding polypeptide.

[0029] In some embodiments, the N-terminus of the polynucleotide ligase polypeptide is linked to the C-terminus of the polynucleotide binding polypeptide.

[0030] In some embodiments, the polynucleotide ligase polypeptide and the polynucleotide binding polypeptide are linked by a first linker.

[0031] In some embodiments, the N-terminus of the polynucleotide ligase polypeptide is linked to the first linker, and the C-terminus of the polynucleotide binding polypeptide is linked to the first linker.

[0032] In some embodiments, the C-terminus of the polynucleotide ligase polypeptide is linked to the first linker, and the N-terminus of the polynucleotide binding polypeptide is linked to the first linker.

[0033] In some embodiments, the first linker is a G4S linker.

[0034] In some embodiments, the polynucleotide ligase polypeptide comprises the amino acid sequence shown in SEQ ID NO: 9-12, 21-22.

[0035] In some embodiments, the polynucleotide-binding polypeptide comprises an amino acid sequence as shown in SEQ ID NO: 13-17.

[0036] In some embodiments, the capping enzyme comprises the amino acid sequence shown in SEQ ID NO: 18-20, 24-29.

[0037] In some embodiments, the uncapped polynucleotide is DNA or RNA.

[0038] In some embodiments, the uncapped polynucleotide is chemically synthesized.

[0039] In some embodiments, the uncapped polynucleotide is 1 to 150 nucleotides in length.

[0040] In some embodiments, the base of the 5' terminal nucleotide of the uncapped polynucleotide is guanine.

[0041] In some embodiments, the 5' terminal nucleotide of the uncapped polynucleotide is modified by phosphorylation.

[0042] In some embodiments, wherein the sugars of the nucleotides of the uncapped polynucleotide are independently selected for each position from ribose and deoxyribose, and may comprise modifications including 2'-O-alkyl, 2'-O-methoxyethyl, 2'-O allyl, 2'-O alkylamine, 2'-fluororibose, and 2'-deoxyribose;

[0043] and / or the bases of the nucleotides of the uncapped polynucleotide are independently selected for each position from adenine, uracil, guanine or cytosine, or an analog of adenine, uracil, guanine or cytosine,

[0044] And the nucleotide modified base can be selected from xanthine, allylaminouracil, allylaminothymidine, hypoxanthine, dioxyadenine, dioxycytosine, dioxyguanine, dioxyuracil, 6-chloropurine nucleoside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil, 5,6-dihydrouracil, 5-[(3-indolyl)propionamide-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxyuracil, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyluracil, 5-propynylaminocytosine, 5-propynylaminouracil, 5-propynyl cytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaadenine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin-16-aminoallyluracil, 3-5-Propynylaminocytosine, 3-6-Propynylaminouracil, cyano-3-aminoallylcytosine, cyano-3-aminoallyllurazolopyrimidine, cyano-5-6-propynylaminocytosine, cyano-5-6-propynylaminouracil, cyano-5-aminoallylcytosine, cyano-5-aminoallyluracil, cyano-7-aminoallyluracil, Dabcyl-5-3-aminoallyluracil, desthiobiotin-16-aminoallyluracil, desthiobiotin-6-aminoallylcytosine, isoguanine, N 1 -ethyl pseudouracil, N 1 -methoxymethyl pseudouracil, N1-methyladenine, N 1 -methyl pseudouracil, N 1 -propyl pseudouracil, N2-methylguanine, N 4 -Biotin-OBEA-cytosine, N4-methylcytosine, N 6 -methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thienylcytosine, thienylguanine, thienyluracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6-macroaminoguanine, 5-formamidouracil, 5-ethynyluracil, N 6-Isopentenyl adenine (i6A), 2-methylthio-N 6 -isopentenyl adenine (ms2i6A), 2-methylthio-N 6 -methyladenine (ms2m6A), N 6 -(cis-hydroxyisopentenyl)adenine (io6A), 2-methylthio-N 6 -(cis-hydroxyisopentenyl)adenine (ms2io6A), N 6 -glycylaminoformyladenine (g6A), N 6 -Threonylaminoformyladenine (t6A), 2-methylthio-N 6 -Threonylaminoformyladenine (ms2t6A), N 6 -methyl-N 6 -Threonylaminoformyladenine (m6t6A), N 6 -hydroxyvalylaminoformyl adenine (hn6A), 2-methylthio-N 6 -Hydroxyvalylcarbamoyladenine (ms2hn6A), N 6 ,N 6 -dimethyladenine (m62A) and N 6 -acetyl adenine (ac6A).

[0045] In some embodiments, the cap analog has the following general formula:

[0046] in:

[0047] R3 is selected from guanine, adenine, cytosine, uracil, a guanine analog, an adenine analog, a cytosine analog, and a uracil analog;

[0048] R4 is (N1p1) x N2, wherein N1 and N2 are ribonucleosides, and N1 is the same as or different from N2;

[0049] Each position of p1 is independently a phosphate group, a phosphorothioate, a phosphorodithioate, an alkylphosphonic acid, an arylphosphonic acid, or an N-phosphoramide bond;

[0050] X is an integer from 0 to 8, wherein if X ≥ 2, the ribonucleosides N1 in (N1-p)x are identical to or different from each other;

[0051] The R1 and R2 groups are independently selected from O-alkyl, halogen, acetylamino (AcNH), hydrogen or hydroxy.

[0052] In some embodiments, wherein the sugars in N1 and N2 are independently selected from ribose and deoxyribose for each position and may contain modifications including 2'-O-alkyl, 2'-O-methoxyethyl, 2'-O allyl, 2'-O alkylamine, 2'-fluororibose, and 2'-deoxyribose;

[0053] And / or the bases in N1 and N2 are independently selected from adenine, uracil, guanine or cytosine for each position, or analogs of adenine, uracil, guanine or cytosine, and the nucleotide modified base may be selected from xanthine, allylaminouracil, allylaminothymidine, hypoxanthine, dioxyadenine, dioxycytosine, dioxyguanine, dioxyuracil, 6-chloropurine nucleoside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil , 5,6-dihydrouracil, 5-[(3-indolyl)propionamido-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyl Uracil, 5-propynylaminocytosine, 5-propynylaminouracil, 5-propynylcytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaadenine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin Biotin-16-aminoallyl uracil, 3-5-propynylaminocytosine, 3-6-propynylaminouracil, cyano-3-aminoallyl cytosine, cyano-3-aminoallyl uracil, cyano-5-6-propynylaminocytosine, cyano-5-6-propynylaminouracil, cyano-5-aminoallyl cytosine, cyano-5-aminoallyl uracil, cyano-7-aminoallyl uracil, Dabcyl-5-3-aminoallyl uracil, desthiobiotin-16-aminoallyl uracil, desthiobiotin-6-aminoallyl cytosine, isoguanine, N 1 -ethyl pseudouracil, N 1 -methoxymethyl pseudouracil, N1-methyladenine, N 1 -methyl pseudouracil, N1 -propyl pseudouracil, N2-methylguanine, N 4 -Biotin-OBEA-cytosine, N4-methylcytosine, N 6 -methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thienylcytosine, thienylguanine, thienyluracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6-macroaminoguanine, 5-formamidouracil, 5-ethynyluracil, N 6 -Isopentenyl adenine (i6A), 2-methylthio-N 6 -isopentenyl adenine (ms2i6A), 2-methylthio-N 6 -methyladenine (ms2m6A), N 6 -(cis-hydroxyisopentenyl)adenine (io6A), 2-methylthio-N 6 -(cis-hydroxyisopentenyl)adenine (ms2io6A), N 6 -glycylaminoformyladenine (g6A), N 6 -Threonylaminoformyladenine (t6A), 2-methylthio-N 6 -Threonylaminoformyladenine (ms2t6A), N 6 -methyl-N 6 -Threonylaminoformyladenine (m6t6A), N 6 -hydroxyvalylaminoformyl adenine (hn6A), 2-methylthio-N 6 -Hydroxyvalylcarbamoyladenine (ms2hn6A), N 6 ,N 6 -dimethyladenine (m62A) and N 6 -acetyl adenine (ac6A).

[0054] In some embodiments, the cap structure analog is a dinucleotide cap analog or a trinucleotide cap analog.

[0055] In some embodiments, the cap structure analog has the following general formula:

[0056] wherein R5 and R6 groups are independently selected from O-alkyl (O-methyl), halogen, tag, hydrogen or hydroxyl;

[0057] X1 and X2 are bases, and X1 and X2 are the same as or different from each other.

[0058] In some embodiments, X1 and / or X2 are guanine or adenine.

[0059] In some embodiments, the cap structure analog is selected from one of the following:

[0060] In some embodiments, the reaction temperature of step S2 is 5°C to 80°C.

[0061] In some embodiments, the reaction temperature of step S2 is 25°C to 60°C.

[0062] In some embodiments, the composition of step S1 further comprises adenosine triphosphate.

[0063] In some embodiments, the composition of step S1 further comprises a DNA and / or RNA buffer.

[0064] In some embodiments, the composition of step S1 further comprises a DNase and / or RNase inhibitor.

[0065] On the other hand, the present application provides a capping enzyme for polynucleotide capping, wherein the capping enzyme is a fusion polypeptide comprising a polynucleotide binding polypeptide fused to a polynucleotide ligase polypeptide.

[0066] In some embodiments, the polynucleotide ligase polypeptide comprises a DNA ligase polypeptide.

[0067] In some embodiments, the DNA ligase polypeptide is a prokaryotic DNA ligase, a prokaryotic DNA ligase variant, or a functional fragment thereof.

[0068] In some embodiments, the DNA ligase polypeptide is a bacterial DNA ligase, a bacterial DNA ligase variant, or a functional fragment thereof.

[0069] In some embodiments, the DNA ligase polypeptide is a viral DNA ligase, a viral DNA ligase variant, or a functional fragment thereof.

[0070] In some embodiments, the DNA ligase polypeptide is T4 DNA ligase, a variant thereof, or a functional fragment thereof.

[0071] In some embodiments, the polynucleotide ligase polypeptide comprises an RNA ligase polypeptide.

[0072] In some embodiments, the RNA ligase polypeptide is a prokaryotic RNA ligase, a prokaryotic RNA ligase variant, or a functional fragment thereof.

[0073] In some embodiments, the RNA ligase polypeptide is a bacterial RNA ligase, a bacterial RNA ligase variant, or a functional fragment thereof.

[0074] In some embodiments, the RNA ligase polypeptide is a viral RNA ligase, a viral RNA ligase variant, or a functional fragment thereof.

[0075] In some embodiments, the RNA ligase polypeptide is T4 RNA ligase, a variant thereof, or a functional fragment thereof.

[0076] In some embodiments, the RNA ligase polypeptide is T4 RNA ligase.

[0077] In some embodiments, the RNA ligase polypeptide is T4 RNA ligase 1.

[0078] In some embodiments, the capping enzyme is a fusion polypeptide comprising a polynucleotide binding polypeptide fused to a polynucleotide ligase polypeptide.

[0079] In some embodiments, the polynucleotide binding polypeptide comprises a DNA binding polypeptide.

[0080] In some embodiments, the polynucleotide binding polypeptide comprises an RNA binding polypeptide.

[0081] In some embodiments, the polynucleotide binding polypeptide is selected from one or more of a DNA double-strand binding domain, a DNA single-strand binding domain, an RNA / DNA composite chain binding domain, an RNA single-strand or RNA double-strand binding protein, wherein

[0082] The DNA double-strand binding domain includes Sso7d and / or NF-kappaB p50;

[0083] The RNA / DNA complex chain binding domain includes one or more of ScFV, RNaseH1 (D210N), and HBD of the monoclonal antibody S9.6;

[0084] RNA single-strand binding domains include: Sso7d, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4.

[0085] In some embodiments, the polynucleotide binding polypeptide includes one or more of Sso7d, NF-kappaB p50, ScFV of monoclonal antibody S9.6, RNaseH1 (D210N), HBD, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4.

[0086] In some embodiments, the polynucleotide binding polypeptide is Sso7d.

[0087] In some embodiments, the C-terminus of the polynucleotide ligase polypeptide is linked to the N-terminus of the polynucleotide binding polypeptide.

[0088] In some embodiments, the N-terminus of the polynucleotide ligase polypeptide is linked to the C-terminus of the polynucleotide binding polypeptide.

[0089] In some embodiments, the polynucleotide ligase polypeptide and the polynucleotide binding polypeptide are linked by a first linker.

[0090] In some embodiments, the first linker is a G4S linker.

[0091] In some embodiments, the polynucleotide ligase polypeptide comprises the amino acid sequence shown in SEQ ID NO: 9-12, 21-22.

[0092] In some embodiments, the polynucleotide-binding polypeptide comprises an amino acid sequence as shown in SEQ ID NO: 13-17.

[0093] In some embodiments, the capping enzyme comprises the amino acid sequence shown in SEQ ID NO: 18-20, 24-29.

[0094] In another aspect, the present application provides a kit comprising the aforementioned capping enzyme.

[0095] In another aspect, the present application provides a use of a capping enzyme in polynucleotide capping.

[0096] Those skilled in the art can easily discern other aspects and advantages of the present application from the detailed description below. In the detailed description below, only exemplary embodiments of the present application are shown and described. As will be appreciated by those skilled in the art, the content of this application enables those skilled in the art to modify the disclosed specific embodiments without departing from the spirit and scope of the invention to which this application relates. Accordingly, the descriptions in the drawings and specification of this application are merely exemplary and not restrictive. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] The specific features of the invention involved in this application are shown in the appended claims. The features and advantages of the invention involved in this application can be better understood by referring to the exemplary embodiments described in detail below and the accompanying drawings. A brief description of the drawings is as follows:

[0098] FIG1 is a schematic diagram of a capping connection method provided by the present application;

[0099] FIG2 is a schematic diagram of the capping effect of wild-type T4 RNA ligase 1;

[0100] FIG3 is a schematic diagram of the capping effect of wild-type RM378 RNA ligase 1;

[0101] FIG4 is the Urea-PAGE detection results of different enzyme capping temperature gradient tests;

[0102] Figure 5 is a schematic diagram of connecting conventional oligonucleotides;

[0103] FIG6 is a Urea-PAGE test result of a temperature gradient test of conventional oligonucleotides ligated with different enzymes;

[0104] FIG7 is a Urea-PAGE assay result showing the concentration of cap analogs;

[0105] FIG8 is an analysis curve of the hat analog concentration test results;

[0106] FIG9 shows the Urea-PAGE detection results of different cap structure analogs and receptor RNA capping;

[0107] Figure 10 shows the results of HiBiT capping Urea PAGE detection;

[0108] Figure 11 shows the mass spectrometry results of the product of Oligo H1 linked to LzCap;

[0109] Figure 12 shows the results of HiBiT tailing Urea PAGE detection;

[0110] FIG13 is the Urea PAGE detection result of HiBiT mRNA ligation product;

[0111] FIG14 shows the mass spectrometry results of HiBiT mRNA recovery products;

[0112] FIG15 is the Urea PAGE detection result of HiBiT mRNA recovery product;

[0113] FIG16 shows a schematic diagram of a strategy for preparing oligonucleotides according to the preparation method of the present application. DETAILED DESCRIPTION

[0114] The following describes the implementation of the present invention through specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification.

[0115] Definition of terms

[0116] In this application, the terms "polynucleotide," "oligonucleotide," "nucleic acid molecule," and "nucleic acid" are used interchangeably and generally include DNA molecules, RNA molecules, analogs of DNA or RNA produced using nucleotide analogs (e.g., peptide nucleic acids and non-naturally occurring nucleotide analogs), and hybrids thereof. Nucleic acid molecules can be single-stranded or double-stranded. The following are non-limiting examples of polynucleotides: genes or gene fragments (e.g., probes, primers, EST or SAGE tags), exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, siRNA, shRNA, RNAi agents, and primers. Polynucleotides can be modified or substituted at one or more bases, sugars, and / or phosphates with any of the various modifications or substitutions described herein or known in the art. Polynucleotides can contain modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, the nucleotide structure can be modified before or after polymer assembly. Polynucleotides can be modified after polymerization, for example by conjugation with a labeling component.

[0117] The term "uncapped polynucleotide" may refer to a given polynucleotide. The uncapped polynucleotide may be a natural polynucleotide or an artificially designed non-natural polynucleotide. The target polynucleotide may be a DNA molecule, an RNA molecule, or a combination of DNA and RNA molecules (e.g., a target polynucleotide that is partially DNA and partially RNA). The uncapped polynucleotide may be a portion of the 5' untranslated region (5'UTR) of an mRNA.

[0118] The term "polypeptide" can include amino acid chains of any length, including full-length proteins, in which the amino acid residues are linked by covalent peptide bonds. The polypeptides of the present invention can be purified natural products or produced in part or in whole using recombinant or synthetic techniques. The term can refer to a polypeptide, a polypeptide polymer (e.g., a dimer or other multimer), a fusion polypeptide, a polypeptide variant, or a derivative thereof.

[0119] The term "capping enzyme" may be an enzyme used to add a cap structure analog to the 5' end or the 3' end of a polynucleotide.

[0120] The term "ligase" refers to an enzyme that can form a covalent bond between two polynucleotides. More specifically, a ligase can form a covalent bond (e.g., 5', 3'-phosphodiester bond) at the 5' end and 3' end of a polynucleotide. These ligases can include DNA ligase or RNA ligase. Within the scope of the present invention, unmodified oligonucleotide can be connected to a ligase of another unmodified oligonucleotide, unmodified oligonucleotide can be connected to a modified oligonucleotide (i.e., the 5' oligonucleotide being modified is connected to the 3' oligonucleotide being unmodified, and / or the 5' oligonucleotide being unmodified is connected to the 3' oligonucleotide being modified), and the oligonucleotide being modified can be connected to a ligase of the oligonucleotide being another modified.

[0121] The term "polynucleotide ligase polypeptide" or "polynucleotide ligase" may refer to an enzyme that ligates two polynucleotides together, generally referring to an enzyme that catalyzes the formation of a phosphodiester bond to connect two polynucleotides together, including DNA ligase polypeptides and RNA ligase polypeptides. The two polynucleotides ligated by the polynucleotide ligase polypeptide can be DNA, RNA, or one DNA and one RNA.

[0122] The term "DNA ligase polypeptide" or "DNA ligase" may refer to an enzyme that joins two DNA segments together, generally catalyzing the formation of a phosphodiester bond to join two DNA segments or one DNA segment and one RNA segment together. DNA ligase polypeptides can join two double-stranded DNA segments, two single-stranded DNA segments, or one single-stranded DNA segment and one double-stranded DNA segment.

[0123] The term "RNA ligase polypeptide" or "RNA ligase" may refer to an enzyme that ligates two RNA segments together, generally catalyzing the formation of a phosphodiester bond to join two RNA segments or a DNA segment to an RNA segment. RNA ligase polypeptides can ligate two double-stranded RNA segments, two single-stranded RNA segments, or one single-stranded RNA segment to a double-stranded RNA segment.

[0124] The term "polynucleotide-binding polypeptide" refers to a polypeptide capable of binding to a polynucleotide, including polypeptides that bind to single-stranded polynucleotides, polypeptides that bind to double-stranded polynucleotides, and polypeptides that bind to polynucleotides in other configurations. As described herein, a polynucleotide-binding polypeptide can be fused to a polynucleotide ligase polypeptide, for example, to the N-terminus or C-terminus of a polynucleotide ligase, without inactivating the polynucleotide-binding polypeptide or the ligase.

[0125] The term "DNA-binding polypeptide" refers to a polypeptide capable of binding to DNA, including polypeptides that bind to single-stranded DNA, polypeptides that bind to double-stranded DNA, and polypeptides that bind to DNA in other configurations. As described herein, a DNA-binding polypeptide can be fused to a DNA ligase polypeptide, for example, to the N-terminus or C-terminus of an RNA ligase polypeptide and / or a DNA ligase polypeptide, without inactivating the RNA ligase polypeptide and / or the DNA ligase polypeptide. It should be understood that RNA-binding polypeptides can also bind to polynucleotides other than RNA, such as DNA or known natural nucleotide analogs.

[0126] The term "RNA-binding polypeptide" refers to a polypeptide capable of binding to RNA, including polypeptides that bind to single-stranded RNA, polypeptides that bind to double-stranded RNA, and polypeptides that bind to other configurations of RNA. As described herein, the RNA-binding polypeptide can be fused to an RNA ligase polypeptide, for example, to the N-terminus or C-terminus of an RNA ligase polypeptide and / or a DNA ligase polypeptide, without inactivating the RNA ligase polypeptide and / or the DNA ligase polypeptide. It should be understood that the RNA-binding polypeptide can also bind to polynucleotides other than RNA, such as DNA or known natural nucleotide analogs.

[0127] The term "fusion polypeptide" refers to a polypeptide comprising two or more amino acid subsequences (e.g., two or more polypeptide domains) fused (e.g., via peptide chains through their respective amino and carboxyl residues) to form a single contiguous polypeptide. It should be understood that the two or more amino acid sequences can be fused directly or indirectly through their respective amino and carboxyl termini via a linker, spacer, or additional polypeptide.

[0128] A "fragment" of a polypeptide is a subsequence of the polypeptide that possesses the function required for enzymatic or binding activity and / or provides three-dimensional structure to the polypeptide.

[0129] The term "domain" refers to a unit of a protein or protein complex, including a polypeptide subsequence, a complete polypeptide sequence or multiple polypeptide sequences, which unit has a defined function.

[0130] The term "linker" refers to an amino acid or nucleotide sequence that indirectly fuses two or more polypeptides or two or more nucleic acid sequences encoding two or more polypeptides. In certain embodiments, the length of the linker or spacer is about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or about 100 amino acids or nucleotides. In other embodiments, the length of the linker or spacer is about 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950 or about 1000 amino acids or nucleotides. In other embodiments, the linker or spacer has a length of about 1 to about 1000 amino acids or nucleotides, about 10 to about 1000, about 50 to about 1000, about 100 to about 1000, about 200 to about 1000, about 300 to about 1000, about 400 to about 1000, about 500 to about 1000, about 600 to about 1000, about 700 to about 1000, about 800 to about 1000, or about 900 to about 1000 amino acids or nucleotides.

[0131] The terms "5'" and "3'" are conventional expressions used to describe nucleic acid sequence characteristics, relating to the position of genetic elements and / or the direction of events such as RNA polymerase transcription or ribosomal translation that proceed in a 5' to 3' direction (5' to 3'). Synonymous terms are upstream (5') and downstream (3'). Typically, DNA sequences, gene maps, and RNA sequences can be drawn from left to right in a 5' to 3' direction, or the 5' to 3' direction can be represented by an arrow symbol, with the arrow pointing in the 3' direction. So when following this conventional usage, 5' (upstream) refers to the placement of the genetic element to the left hand side, while 3' (downstream) refers to the placement of the genetic element to the right hand side. Among them, the term "5' end" or "5' end" can be the first nucleotide of a polynucleotide along the 5' to 3' direction or a continuous section of nucleotides with this nucleotide as the endpoint, and the "3' end" or "3' end" can be the last nucleotide of a polynucleotide along the 5' to 3' direction or a continuous section of nucleotides with this nucleotide as the endpoint.

[0132] As used herein, the term "C-terminus" generally refers to the carboxyl terminus of a polypeptide.

[0133] As used herein, the term "N-terminus" generally refers to the amino terminus of a polypeptide.

[0134] The terms "ligate," "ligation," and "ligation" are used interchangeably to refer to the formation of a covalent bond (e.g., a 5', 3'-phosphodiester bond) between two polynucleotides. Ligation may also refer to the formation of a covalent bond (e.g., a 5', 3'-phosphodiester bond) between the 5' end of one polynucleotide and the 3' end of another polynucleotide.

[0135] The term "capping" may refer to the addition of a cap structure analog at the 5' and / or 3' end of a polynucleotide to generate a 5' and / or 3' capped polynucleotide.

[0136] The term "cap analogue" includes natural or artificial cap structures (cap), such as m7G and any compound of the following general formula:

[0137] Cap structure analogs include, for example, dinucleotide cap analogs of formula m7G(5')p3(5')G, in which guanine nucleotide (G) is connected to a triphosphate bridge via its 5'OH. In some dinucleotide cap analogs, the 3'-OH group is replaced by hydrogen or OCH3 (reference US 7,074,596; Kore, Nucleotides, Nucleotides, and Nucleic Acids, 2006, 25:307-14; and Kore, Nucleotides, Nucleotides, and Nucleic Acids, 2006, 25:337-40). Dinucleotide cap analogs include m7G(5')p3G, 3'-OMe-m7G(5')p3G (ARCA). The term "cap structure analog" also includes trinucleotide cap analogs (defined below) and other longer molecules (e.g., cap analogs with four, five, or six or more nucleotides connected to a triphosphate bridge). 'Cap structure analogs can include 5' capping nucleotides connected to the 5' end of the mRNA through a 5' to 5' triphosphate internucleotide bond. In some embodiments, the nucleotides connected to the mRNA through a 5' to 5' triphosphate internucleotide bond are referred to as "natural" cap structure analogs. In some embodiments, the natural cap structure analog is a 7-methylguanosine (m7G) nucleotide. In some embodiments, the 5' cap is a modified 5' cap that includes one or more modified nucleotides, such as a 5' capping nucleotide, or one or more modified internucleotide modifications, such as modifications to the 5' to 5' triphosphate internucleotide bond. In some embodiments, the 5' cap includes one or more nucleotides with a sugar modification (such as 2'-O-methylation).

[0138] The terms "functional variant" and "functional fragment" as used herein, for example in relation to DNA / RNA ligases or DNA / RNA binding polypeptides, refer to polypeptide sequences that are different from the specific identification sequence, in which one or more amino acid residues are deleted, replaced or added, or sequences that include specific identification sequence fragments. Functional variants can be naturally occurring allelic variants or non-naturally occurring variants. Functional variants can be from the same species or other species and can include homologues, collateral relatives and direct relatives. Functional variants or functional fragments of a polypeptide have one or more biological activities of the native specific identification polypeptide, such as the ability to stimulate one or more biological effects of the native polypeptide. For example, functional fragments of DNA or RNA ligases are generally capable of catalyzing the formation of phosphodiester bonds.

[0139] The term "messenger RNA" (mRNA) refers to a ribonucleic acid (RNA) molecule that can mediate the transfer of genetic information in the cytoplasm to the ribosome, where it serves as a template for protein synthesis. It is synthesized from a DNA template during transcription. In eukaryotes, mRNA is transcribed in chromosomes in vivo by intracellular RNA polymerase. During or after in vivo transcription, a 5'-end cap (also known as an RNA end, an RNA 7-methylguanosine end, or an RNA m7G end) is added to the 5' end of the mRNA in vivo. The 5' end is a terminal 7-methylguanosine residue that is connected to the first transcribed nucleotide via a 5'-5'-triphosphate bond. In addition, most eukaryotic mRNA molecules have a polyadenylic acid portion ("poly (A) tail") at the 3' end of the mRNA molecule. In vivo, a polyadenylic acid portion (Poly A) is added after transcription in eukaryotic cells. Therefore, a typical mature eukaryotic mRNA has the following structure: starting with the terminal nucleotide of the mRNA at the 5' end, followed by a 5' untranslated region (5'UTR) of nucleotides, then an open reading frame (ORF, which may be a protein coding sequence) starting with a start codon (which is an AUG triplet of nucleotide bases) and ending with a stop codon (which may be a UAA, UAG, or UGA triplet of nucleotide bases), and then a 3' untranslated region (3'UTR) of nucleotides, ending with a polyadenosine moiety. Although the characteristics of a typical mature eukaryotic mRNA can be naturally produced in eukaryotic cells in vivo, the same or structurally and functionally equivalent characteristics can also be prepared in vitro using molecular biological methods. Therefore, any RNA with a structure similar to that of a typical mature eukaryotic mRNA can function as an mRNA and is also included within the scope of the term "messenger RNA."

[0140] The term "target polynucleotide" may refer to a given polynucleotide. The target polynucleotide may be a natural polynucleotide or an artificially designed non-natural polynucleotide. The target polynucleotide may be a DNA molecule, an RNA molecule, or a combination of DNA and RNA molecules (e.g., a target polynucleotide comprising a portion of DNA and a portion of RNA).

[0141] The term "polynucleotide to be linked" refers to a portion of a target polynucleotide. The polynucleotide to be linked can be a continuous or discontinuous sequence of the target polynucleotide. The polynucleotide to be linked can include the 5' end or the 3' end of the target polynucleotide. The polynucleotide to be linked can be a non-Poly A region of the target polynucleotide.

[0142] The terms "first polynucleotide" and "second polynucleotide" may be part of a target polynucleotide. In the target polynucleotide, the polynucleotide to be connected, the first polynucleotide, and the second polynucleotide are arranged in a 5' to 3' direction. In some embodiments, the 3' end of the polynucleotide to be connected may form a covalent bond with the 5' end of the first polynucleotide, and the 3' end of the first polynucleotide may form a covalent bond with the 5' end of the second polynucleotide. In some embodiments, the first polynucleotide may include a PolyA region and a non-PolyA region. In some embodiments, the first polynucleotide may only include a PolyA region and a non-PolyA region. In some embodiments, the second polynucleotide is entirely a PolyA region.

[0143] In this application, the terms "poly(A) portion," "polyA," "poly(A) tail," "poly(A)," "poly(A) portion," "PolyA region," "PolyA," "PolyA tail," and "PolyA tail" are used interchangeably to refer to a nucleic acid sequence that is attached to the 3' end of a nucleic acid (e.g., RNA) and is composed primarily of adenosine nucleotides (either adenosine or deoxyadenosine). The poly(A) portion can be composed of 10%-100%, 25%-100%, 30%-100%, 40%-100%, 50%-100%, 60%-100%, 70%-100%, 80%-100%, 90%-100%, 95%-100%, 96%-100%, 97%-100%, 98%-100%, or 99%-100% adenosine nucleotides. In some embodiments, the polyadenosine portion (PolyA region) may be composed entirely of adenosine nucleotides (which may be adenosine or deoxyadenosine). The adenosine nucleotides included in the polyadenosine portion may be standard adenosine nucleotides or modified (non-standard) adenosine nucleotides.

[0144] In the present application, the term "non-Poly A region" can be a sequence other than the Poly A region in the target polynucleotide. For example, a typical mature eukaryotic mRNA has the following structure: starting from the mRNA terminal nucleotide at the 5' end, followed by a 5' untranslated region (5'UTR) of nucleotides, then an open reading frame (which can be a protein coding sequence), and then a 3' untranslated region (3'UTR) of nucleotides, ending with a polyadenosine moiety. The non-Poly A region of the mRNA can include at least a portion of the 5'UTR, at least a portion of the 3'UTR, at least a portion of the open reading frame, or any combination thereof.

[0145] In this application, the term "first functional region" can refer to a segment of a target polynucleic acid that has a certain function. The function can be related to gene translation, expression, or regulation. For example, the first functional region in DNA can include a coding region or a non-coding region. In mRNA, the first functional region can include a 5' UTR, a 3' UTR, or an open reading frame.

[0146] The term "splint oligonucleotide" generally refers to an oligonucleotide having a first sequence of nucleotides that are regionally complementary or reversely complementary to a first oligonucleotide adjacent to an end, and a second sequence of nucleotides that are regionally complementary or reversely complementary to a second oligonucleotide adjacent to an end. The splint oligonucleotide is capable of binding to the first oligonucleotide and the second oligonucleotide to produce a complex, thereby bringing the ends of the first oligonucleotide and the second oligonucleotide into close spatial proximity. By bringing the ends of the first oligonucleotide and the second oligonucleotide into close proximity, the splint oligonucleotide increases the chance of connection between the first oligonucleotide and the second oligonucleotide. For example, the splint oligonucleotide can increase the chance of connection between the polynucleotide to be connected and the first polynucleotide. For another example, the splint oligonucleotide can increase the chance of connection between the polynucleotide to be connected and the second complex.

[0147] Detailed Description of the Invention

[0148] Capping enzyme

[0149] In one aspect, the present application provides a capping enzyme, wherein the capping enzyme may include a polynucleotide ligase polypeptide.

[0150] In the present application, polynucleotide ligase polypeptides include DNA ligase polypeptides. For example, the DNA ligase polypeptide can be a prokaryotic DNA ligase, a prokaryotic DNA ligase variant, or a functional fragment thereof. For another example, the DNA ligase polypeptide can be a bacterial DNA ligase, a bacterial DNA ligase variant, or a functional fragment thereof. For another example, the DNA ligase polypeptide can be a viral DNA ligase, a viral DNA ligase variant, or a functional fragment thereof.

[0151] In the present application, the DNA ligase polypeptide may be T4 DNA ligase, a variant thereof, or a functional fragment thereof.

[0152] In the present application, the polynucleotide ligase polypeptide may include an RNA ligase polypeptide. For example, the RNA ligase polypeptide may be a prokaryotic RNA ligase, a prokaryotic RNA ligase variant, or a functional fragment thereof. For another example, the RNA ligase polypeptide may be a bacterial RNA ligase, a bacterial RNA ligase variant, or a functional fragment thereof. For another example, the RNA ligase polypeptide may be a viral RNA ligase, a viral RNA ligase variant, or a functional fragment thereof.

[0153] In the present application, the RNA ligase polypeptide is T4 RNA ligase, a variant thereof or a functional fragment thereof.

[0154] In the present application, the RNA ligase polypeptide is T4 RNA ligase. For example, the RNA ligase polypeptide is T4 RNA ligase 1 or a functional fragment of T4 RNA ligase 1. For another example, the RNA ligase polypeptide is T4 RNA ligase 2 or a functional fragment of T4 RNA ligase 2.

[0155] In the present application, the capping enzyme can be a fusion polypeptide comprising a polynucleotide binding polypeptide fused to a polynucleotide ligase polypeptide, wherein the polynucleotide binding polypeptide comprises a DNA binding polypeptide and / or an RNA binding polypeptide.

[0156] In the present application, the polynucleotide binding polypeptide is selected from one or more of a DNA double-strand binding domain, a DNA single-strand binding domain, an RNA / DNA composite chain binding domain, an RNA single-strand or RNA double-strand binding protein.

[0157] For example, the double-stranded DNA binding domain may include Sso7d and / or NF-kappaB p50; the RNA / DNA complex chain binding domain may include one or more of ScFV, RNaseH1 (D210N), and HBD of monoclonal antibody S9.6; the single-stranded RNA binding domain may include: Sso7d, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4.

[0158] In the present application, the polynucleotide binding polypeptide includes one or more of Sso7d, NF-kappaB p50, ScFV of monoclonal antibody S9.6, RNaseH1 (D210N), HBD, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4.

[0159] For example, the polynucleotide binding polypeptide is Sso7d.

[0160] In the present application, the capping enzyme is a fusion polypeptide, wherein the C-terminus of the polynucleotide ligase polypeptide is linked to the N-terminus of the polynucleotide binding polypeptide, or the N-terminus of the polynucleotide ligase polypeptide is linked to the C-terminus of the polynucleotide binding polypeptide.

[0161] In the present application, the polynucleotide ligase polypeptide and the polynucleotide binding polypeptide are connected by a first joint. The first joint indirectly fuses two or more polypeptides or two or more nucleic acid sequences encoding two or more polypeptides or nucleotide sequences. In certain embodiments, the length of the first joint is about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or about 100 amino acids or nucleotides. In other embodiments, the length of the connector or spacer is about 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950 or about 1000 amino acids or nucleotides. In other embodiments, the first linker has a length of about 1 to about 1000 amino acids or nucleotides, about 10 to about 1000, about 50 to about 1000, about 100 to about 1000, about 200 to about 1000, about 300 to about 1000, about 400 to about 1000, about 500 to about 1000, about 600 to about 1000, about 700 to about 1000, about 800 to about 1000, or about 900 to about 1000 amino acids or nucleotides.

[0162] In the present application, the N-terminus of the polynucleotide ligase polypeptide is connected to the first linker, and the C-terminus of the polynucleotide binding polypeptide is connected to the first linker; or the C-terminus of the polynucleotide ligase polypeptide is connected to the first linker, and the N-terminus of the polynucleotide binding polypeptide is connected to the first linker.

[0163] In the present application, the first linker may include a G4S linker (eg, GGGGS (SEQ ID NO: 5)). In certain embodiments, the first linker may be a repeat of multiple G4S linkers, such as the sequence shown in SEQ ID 23.

[0164] In the present application, the N-terminus or C-terminus of the capping enzyme may have a tag for protein purification, for example, the protein purification tag includes one or more of 6xHis, GST, DYKDDDDK (SEQ ID NO: 45) (FLAG), c-myc, or HA.

[0165] In the present application, the polynucleotide ligase polypeptide includes the amino acid sequence shown in any one of SEQ ID NOs: 9-12 and 21-22.

[0166] In the present application, the polynucleotide ligase polypeptide is encoded by a nucleotide sequence as shown in any one of SEQ ID NOs: 33-35.

[0167] In the present application, the polynucleotide ligase polypeptide is encoded by the nucleotide sequence shown in SEQ ID NO: 47. In the present application, the polynucleotide binding polypeptide comprises the amino acid sequence shown in any one of SEQ ID NOs: 13-17.

[0168] In some embodiments, the polynucleotide-binding polypeptide may comprise an amino acid sequence as shown in SEQ ID NO:46.

[0169] In the present application, the polynucleotide binding polypeptide includes the nucleotide sequence encoded by any one of SEQ ID NOs: 30-32.

[0170] In the present application, the capping enzyme includes the amino acid sequence shown in any one of SEQ ID NOs: 18-20, 24-29.

[0171] In the present application, the capping enzyme may be an amino acid sequence as shown in SEQ ID NO: 48.

[0172] In the present application, the capping enzyme is encoded by the nucleotide sequence shown in any one of SEQ ID NOs: 36-44.

[0173] In the present application, the capping enzyme may be encoded by the nucleotide sequence shown in SEQ ID NO:49.

[0174] In the present application, the polynucleotide ligase polypeptide includes the amino acid sequence shown in any one of SEQ ID NOs: 9-12 and 21-22.

[0175] In the present application, the polynucleotide-binding polypeptide is the amino acid sequence shown in any one of SEQ ID NOs: 13-17.

[0176] In the present application, the capping enzyme is the amino acid sequence shown in any one of SEQ ID NOs: 18-20, 24-29.

[0177] The present application also includes conservative substitutions of one or more amino acids in the polypeptide sequence without significantly altering its biological activity. A skilled artisan will appreciate methods for making phenotypically silent amino acid substitutions (e.g., see Bowie et al., 1990, Science 247, 1306). Similarly, the present invention also includes functional variants produced by substitutions (including non-conservative substitutions) of one or more amino acids.

[0178] The term "variant" refers to polypeptides, including naturally occurring, recombinant, and synthetic polypeptides, that are at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 5%, at least 96%, at least 97%, at least 98%, or at least 99% identical to a sequence of the invention. Identity is found over a comparison window of at least 20 amino acid positions, preferably at least 50 amino acid positions, at least 100 amino acid positions, or over the entire length of the polypeptide of the invention.

[0179] Cap structure analogs

[0180] On the other hand, the present application provides a cap structure analogue, which may have the following general formula:

[0181] in:

[0182] R3 is selected from guanine, adenine, cytosine, uracil, a guanine analog, an adenine analog, a cytosine analog, a uracil analog; R4 is (N1p) x N2, wherein N1 and N2 are ribonucleosides, and N1 is the same as or different from N2; p1 is independently a phosphate group, a thiophosphate, a dithiophosphate, an alkylphosphonic acid, an arylphosphonic acid, or an N-phosphoramide bond at each position; x is an integer from 0 to 8, wherein if x ≥ 2, the ribonucleosides N1 in (N1-p)x are the same as or different from each other; R1 and R2 groups are independently selected from O-alkyl, halogen, acetamido (AcNH), hydrogen, or hydroxyl.

[0183] In the present application, the sugars in N1 and N2 are independently selected from ribose and deoxyribose for each position, and may contain modifications including 2'-O-alkyl, 2'-O-methoxyethyl, 2'-O allyl, 2'-O alkylamine, 2'-fluororibose and 2'-deoxyribose, for example, the bases in N1 and N2 are independently selected from adenine, uracil, guanine or cytosine for each position, or analogs of adenine, uracil, guanine or cytidine, and the nucleotide modified bases may be selected from xanthine, allylaminouracil, allylaminothymidine, hypoxanthine, dioxyadenine, dioxycytosine, dioxyguanine, dioxyuracil, 6-chloropurine nucleoside, N 6 -methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil, 5,6-dihydrouracil, 5-[(3-indolyl)propionamido-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxyuracil, 5-fluorouracil , 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyluracil, 5-propynylaminocytosine, 5-propynylaminouracil, 5-propynylcytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6- Thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaadenine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin-16-aminoallyluracil, 3-5-propynylaminocytosine, 3-6-propynyl Cyano-3-aminoallyl uracil, Cyano-3-aminoallyl cytosine, Cyano-3-aminoallyl uracil, Cyano-5-6-propynylamino cytosine, Cyano-5-6-propynylamino uracil, Cyano-5-aminoallyl cytosine, Cyano-5-aminoallyl uracil, Cyano-7-aminoallyl uracil, Dabcyl-5-3-aminoallyl uracil, Desthiobiotin-16-aminoallyl uracil, Desthiobiotin-6-aminoallyl cytosine, Isoguanine, N 1 -ethyl pseudouracil, N 1 -methoxymethyl pseudouracil, N1-methyladenine, N 1 -methyl pseudouracil, N 1-propyl pseudouracil, N2-methylguanine, N 4 -Biotin-OBEA-cytosine, N4-methylcytosine, N 6 -methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thienylcytosine, thienylguanine, thienyluracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6-macroaminoguanine, 5-formamidouracil, 5-ethynyluracil, N 6 -Isopentenyl adenine (i6A), 2-methylthio-N 6 -isopentenyl adenine (ms2i6A), 2-methylthio-N 6 -methyladenine (ms2m6A), N 6 -(cis-hydroxyisopentenyl)adenine (io6A), 2-methylthio-N 6 -(cis-hydroxyisopentenyl)adenine (ms2io6A), N 6 -glycylaminoformyladenine (g6A), N 6 -Threonylaminoformyladenine (t6A), 2-methylthio-N 6 -Threonylaminoformyladenine (ms2t6A), N 6 -methyl-N 6 -Threonylaminoformyladenine (m6t6A), N 6 -hydroxyvalylaminoformyl adenine (hn6A), 2-methylthio-N 6 -Hydroxyvalylcarbamoyladenine (ms2hn6A), N 6 ,N 6 -dimethyladenine (m62A) and N 6 -acetyl adenine (ac6A).

[0184] In the present application, the cap structure analogue may be a dinucleotide cap analogue or a trinucleotide cap analogue.

[0185] In the present application, the cap structure analogue may have the following general formula:

[0186] wherein R5 and R6 groups are independently selected from O-alkyl (O-methyl), halogen, tag, hydrogen or hydroxyl;

[0187] The R2 group is selected from -CH2NHAc, O-alkyl (O-methyl), halogen, tag, hydrogen or hydroxyl;

[0188] X1 and X2 are bases, and X1 and X2 are the same as or different from each other.

[0189] In the present application, X1 and / or X2 are guanine or adenine.

[0190] In the present application, the cap structure analog is selected from one of the following:

[0191] Capping method

[0192] In one aspect, the present application provides a method for capping a polynucleotide, comprising:

[0193] S1. Provide a composition comprising at least: an uncapped polynucleotide, the aforementioned capping enzyme, and the aforementioned cap structure analog, wherein the capping enzyme comprises a polynucleotide ligase polypeptide;

[0194] S2. Forming a capping polynucleotide according to the composition of S1.

[0195] In this application, an uncapped polynucleotide is DNA or RNA.

[0196] In the present application, uncapped polynucleotides are chemically synthesized.

[0197] In the present application, the length of the uncapped polynucleotide is 1 to 2000 nucleotides. For example, the length of the uncapped polynucleotide is 1 to 2000 nucleotides. In the present application, the length of the uncapped polynucleotide can be 1 to 2000 nucleotides, 1 to 1500 nucleotides, 1 to 1000 nucleotides, 1 to 500 nucleotides, or 1 to 100 nucleotides. For example, the number of uncapped polynucleotides can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83 3, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 1 19, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150.As another example, the number of nucleotides in an uncapped polynucleotide can range from 1-150, 1-140, 1-130, 1-120, 1-110, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, 10-150, 10-140, 10-130, 10-120, 10-110, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, 10-20, 20-150, 20-140, 20-130, 20-120, 20-110, 20-100, 20-90, 20-80 , 20-70, 20-60, 20-50, 20-40, 20-30, 30-150, 30-140, 30-130, 30-120, 30-110, 30-100, 30-90, 30-80, 30-70, 30-60, 30-50, 30-40, 40-150, 40-140, 40-130, 40-120, 40-110, 40-100, 40-90, 40-80, 40-70, 40-60, 40-50, 50-150, 50-140, 50-130, 50-120, 50-110, 50-100, 50-90, 50-80, 50-70, 50-60.

[0198] Due to different In the present application, in order to improve the connection efficiency, the 5' end of the uncapped polynucleotide can be a specific nucleotide, or a nucleotide containing a specific modification, or a nucleotide containing a specific sequence. For example, the base of the 5' end nucleotide of the uncapped polynucleotide is guanine. For another example, the base of the 5' end nucleotide of the uncapped polynucleotide is adenine. For another example, the base of the 5' end nucleotide of the uncapped polynucleotide is cytosine. For another example, the base of the 5' end nucleotide of the uncapped polynucleotide is guanine. The connection method disclosed in the present application can have different efficiencies for different bases of the 5' end nucleotide. Those skilled in the art can select the base of the 5' end nucleotide with higher connection efficiency according to the method disclosed in the present application.

[0199] In the present application, in order to prevent oligonucleotide self-ligation, the 5' terminal nucleotide of the uncapped polynucleotide is phosphorylated.

[0200] In the present application, the sugars of the nucleotides of the uncapped polynucleotide are independently selected for each position from ribose and deoxyribose, and may contain modifications including 2'-O-alkyl, 2'-O-methoxyethyl, 2'-O allyl, 2'-O alkylamine, 2'-fluororibose, and 2'-deoxyribose.

[0201] In the present application, the bases of the nucleotides of the uncapped polynucleotide are independently selected from adenine, uracil, guanine or cytosine for each position, or an analog of adenine, uracil, guanine or cytosine, and the nucleotide modifications may be selected from xanthine, allylaminouracil, allylaminothymidine, hypoxanthine, dioxyadenine, dioxycytosine, dioxyguanine, dioxyuracil, 6-chloropurine nucleoside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil, 5,6-dihydrouracil, 5-[(3-indolyl)propionamide-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxyuracil, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyluracil, 5-propynylaminocytosine, 5-propynylaminouracil, 5-propynylcytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deoxycytosine Azaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaadenine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin-16-aminoallyluracil, 3-5-propynylaminocytosine, 3-6-propynylaminouracil, cyano-3-aminoallylcytosine, cyano-3-aminoallyluracil, cyano-5-6-propynylaminocytosine, cyano-5-6-propynylaminouracil, cyano-5-aminoallylcytosine, cyano-5-amino Allyluracil, cyano 7-aminoallyluracil, Dabcyl-5-3-aminoallyluracil, desthiobiotin-16-aminoallyluracil, desthiobiotin-6-aminoallylcytosine, isoguanine, N1-ethylpseudouracil, N1-methoxymethylpseudouracil, N1-methyladenine, N1-methylpseudouracil, N1-propylpseudouracil, N2-methylguanine, N4-biotin-OBEA-cytosine, N4-methylcytosine, N6-methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thienylcytosine, thienylguanine, thienyluracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6-macroaminoguanine, 5-formamidouracil, 5-ethynyluracil, N6-isopentenyladenine (i6A), 2-methylthio-N6-isopentenyladenine (ms2i6A), 2-methylthio-N6-methyladenine (ms2m6A), N6-(cis-hydroxyisopentenyl)adenine (io6A), 2-methylthio-N6-(cis-hydroxyisopentenyl)adenine (ms2io6A), N6-glycylaminoformyladenine (g6A), N6-threonylaminoformyladenine (t6A), 2-methylthio-N6-threonylaminoformyladenine (ms2t6A), N6-methyl-N6-threonylaminoformyladenine (m6t6A), N6-hydroxyvalylaminoformyladenine (hn6A), 2-methylthio-N6-hydroxyvalylaminoformyladenine (ms2hn6A), N6,N6-dimethyladenine (m62A) and N6-acetyladenine (ac6A).

[0202] In the present application, the reaction temperature of step S2 is 5° C. to 80° C. For example, it can be 5° C. to 80° C., 5° C. to 70° C., 5° C. to 60° C., 5° C. to 50° C., 5° C. to 40° C., 5° C. to 30° C., 10° C. to 80° C., 10° C. to 70° C., 10° C. to 60° C., 10° C. to 50° C., 10° C. to 40° C., 10° C. to 30° C., 20° C. to 80° C., 20° C. to 70° C., 20° C. to 60° C., 20° C. to 50° C., 20° C. to 40° C., or 20° C. to 30° C.

[0203] In the present application, the reaction temperature of step S2 is 25°C to 60°C.

[0204] In the present application, the composition of step S1 further includes adenosine triphosphate.

[0205] In the present application, the composition of step S1 further includes a DNA and / or RNA buffer.

[0206] In the present application, the composition of step S1 further includes a DNase and / or RNase inhibitor.

[0207] The present application further provides a kit comprising the capping enzyme as described above.

[0208] The present application further provides use of the aforementioned capping enzyme in polynucleotide capping.

[0209] Without intending to be bound by any theory, the following embodiments are merely intended to illustrate various technical solutions of the present invention and are not intended to limit the scope of the present invention.

[0210] Tailing method

[0211] The present application further includes a method for synthesizing a polynucleotide, comprising a capping step and a tailing step of the polynucleotide, wherein the capping step is as described above, and the tailing step is as described below.

[0212] Component Design

[0213] The oligonucleotide sequence design of the present application can be shown in Figure 16, taking mRNA 100 as an example. mRNA 100 may include a 5'UTR 101, an ORF 102, a 3'UTR 103, and a PolyA region 104. Traditional chemical synthesis or in vitro transcription methods can obtain 5'UTR 101, ORF 102, and 3'UTR 103, but cannot obtain a complete and stable length PolyA region 104. One of the main problems addressed by the present application is how to obtain a complete and stable length PolyA region 104.

[0214] First, the target polynucleotide 100 (taking mRNA as an example) is divided into a polynucleotide to be connected 201, a first polynucleotide 202, and a second polynucleotide 203. The polynucleotide to be connected 201, the first polynucleotide 202, and the second polynucleotide 203 can constitute the complete target polynucleotide 100. Specifically, the first polynucleotide 202 can contain a non-PolyA region 2021 and a PolyA region 2022. The non-PolyA region 2021 can be a portion of the target polynucleotide sequence 103. The non-PolyA region 2021 can be connected to a portion of the polynucleotide to be connected 201 to form a complete 3'UTR 103. The second polynucleotide 203 can be entirely a PolyA region. The PolyA region 2022 can constitute the complete PolyA region 104 of the target polynucleotide 100 with the second polynucleotide 203. The first polynucleotide 202 and the second polynucleotide 203 can be chemically synthesized, and their lengths are stable and controllable. By dividing the target polynucleotide into multiple fragments, synthesizing them separately, and then connecting them with T4 RNA ligase, stable and controllable PolyA mRNA can be obtained.

[0215] Furthermore, to enhance the efficiency of fragment ligation, a splint oligonucleotide 300 can be designed. The splint oligonucleotide 300 can be divided into a first region 301 and a second region 302. The first region 301 corresponds to at least a portion of the polynucleotide 201 to be ligated (complementary or reverse complementary), and the first region 302 corresponds to at least a portion of the first polynucleotide 202 (complementary or reverse complementary). The second region 302 can include a second subregion 3021 and a first subregion 3022. The second subregion 3021 can correspond to the non-polyA region 2021 of the first polynucleotide 202 (complementary or reverse complementary), and the first subregion 3022 can correspond to the polyA region 2022 of the first polynucleotide 202 (complementary or reverse complementary). The splint oligonucleotide can be a DNA sequence or an RNA sequence.

[0216] In some embodiments, the first region 301 is reverse complementary to at least a portion of the polynucleotide to be connected 201, and the first region 302 is reverse complementary to at least a portion of the first polynucleotide 202. The second region 302 may include a second subregion 3021 and a first subregion 3022, wherein the second subregion 3021 may be reverse complementary to the non-polyA region 2021 of the first polynucleotide 202, and the first subregion 3022 may be reverse complementary to the polyA region 2022 of the first polynucleotide 202.

[0217] The above is a specific embodiment of the component design of the present application. The polynucleotide to be connected, the first polynucleotide and the second polynucleotide described in the present application are not limited to the above embodiment.

[0218] Multi-stage PolyA binding solution

[0219] To obtain a stable nucleic acid with a fixed-length Poly A tail, one approach of the present application is to divide the target polynucleotide into multiple fragments, obtain these fragments by chemical synthesis or in vitro transcription, at least two of which contain Poly A regions, and then synthesize these fragments into a complete target polynucleotide using a ligase. The target polynucleotide obtained by the above method has at least the following advantages:

[0220] (1) By obtaining fragments separately, a PolyA tail of controllable length can be obtained in the target polynucleotide;

[0221] (2) By obtaining the fragments separately, the PolyA tail can be modified in a controllable manner.

[0222] Based on the above scheme, on the one hand, the present application provides a method for preparing a target polynucleotide, which may include the following steps:

[0223] Step S1: providing polynucleotides to be connected, providing a first polynucleotide, and providing a second polynucleotide, wherein the first polynucleotide and / or the second polynucleotide includes a Poly A region;

[0224] Step S2: covalently linking the polynucleotide to be linked, the first polynucleotide, and the second polynucleotide to obtain the target polynucleotide.

[0225] In certain embodiments, the first polynucleotide, and / or the second polynucleotide, and / or the polynucleotide to be joined, and / or the target polynucleotide are ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). For example, when the target polynucleotide is RNA (e.g., mRNA), the first polynucleotide, the second polynucleotide, and the polynucleotide to be joined are all RNA. For another example, when the target polynucleotide is DNA, the first polynucleotide, the second polynucleotide, and the polynucleotide to be joined are all deoxyribonucleic acid (DNA).

[0226] In certain embodiments, the first polynucleotide, and / or the second polynucleotide, and / or the polynucleotide to be ligated, and / or the target polynucleotide is a combination of ribonucleic acid (RNA) and deoxyribonucleic acid (DNA).

[0227] In certain embodiments, the first polynucleotide, and / or the second polynucleotide, and / or the polynucleotide to be joined are chemically synthesized or in vitro transcribed. For example, the first polynucleotide and the second polynucleotide are chemically synthesized, and the polynucleotide to be joined is in vitro transcribed. In another example, the polynucleotide to be joined, the first polynucleotide, and the second polynucleotide are all chemically synthesized.

[0228] The advantages of chemical synthesis for obtaining the first polynucleotide, and / or the second polynucleotide, and / or the polynucleotide to be linked are: (1) chemical synthesis is more controllable, and the length of the obtained polynucleotides can be controlled, especially the first polynucleotide and / or the second polynucleotide, so that the length of the obtained poly A is easily controlled; (2) modification is convenient, and modified nucleotides can be added as required; (3) compared to traditional enzymatic polymerization methods, less biological raw materials are used, making it easier to meet GMP requirements. The synthesis of polynucleotides can be carried out using the method described in Example 5.

[0229] In certain embodiments, the first polynucleotide and / or the second polynucleotide may include a polyA region. For example, the first polynucleotide may include a polyA region and a non-polyA region, and the second polynucleotide may be polyA.

[0230] In certain embodiments, the first polynucleotide and the second polynucleotide may both be poly A. For another example, the first polynucleotide and the second polynucleotide each include a poly A region and a non-poly A region. The non-poly A region may include 5-100 nucleotides.

[0231] In certain embodiments, the length of the first polynucleotide and / or the second polynucleotide may be 1 to 150 nucleotides. The number of nucleotides in the first polynucleotide and / or the second polynucleotide may be 1 to 150 nucleotides.

[0232] In certain embodiments, the first polynucleotide and the second polynucleotide may be of the same length or different lengths, for example, the first polynucleotide is longer than the second polynucleotide, or the first polynucleotide is shorter than the second polynucleotide.

[0233] In certain embodiments, the polynucleotides to be linked may not contain a polyA region.

[0234] In certain embodiments, the target polynucleotide may include a first functional region. For example, the first functional region in mRNA may include one or more of a 3'UTR, an ORF, and a 5'UTR. For another example, the first functional region in DNA may include a coding region or a non-coding region.

[0235] In certain embodiments, at least a portion of the polynucleotide to be linked is covalently linked to the non-polyA region of the first polynucleotide and / or the second polynucleotide to form a first functional region. For example, the first polynucleotide includes a non-polyA region and a polyA region, the second polynucleotide includes a polyA region, wherein the non-polyA region includes a portion of the first functional region (e.g., a portion of the 3'UTR), and the polynucleotide to be linked includes the remaining portion of the first functional region (e.g., another portion of the 3'UTR). When the polynucleotide to be linked, the first polynucleotide, and the second polynucleotide are covalently linked, the first polynucleotide including the non-polyA region and the polynucleotide to be linked are linked to obtain the first functional region.

[0236] In certain embodiments, the method for preparing the target polynucleotide further comprises covalently linking the polynucleotide to be linked, the first polynucleotide, and the second polynucleotide in the presence of a ligase to obtain the target polynucleotide. The ligase may comprise a DNA ligase or an RNA ligase. For example, the ligase may comprise one or more of T4 DNA ligase and T4 RNA ligase. The T4 RNA ligase may comprise one or more of T4 RNA ligase 1 and T4 RNA ligase 2. The ligase may comprise a single-stranded ligase or a double-stranded ligase. For example, the single-stranded ligase may comprise one or more of T4 RNA ligase 1, RM378 RNA ligase, or TS2126 RNA ligase, and the double-stranded ligase may comprise one or more of T4 DNA ligase, T3 DNA ligase, and T4 RNA ligase 2. The ligase may further include one or more of RM378 RNA ligase, TS2126 RNA ligase, E. coli RNA ligase, Mth RNA ligase, RTCB RNA ligase, T3 DNA ligase, T7 DNA ligase, Taq DNA ligase, marine archaea Thermococcus sp DNA ligase, Chlorella virus DNA ligase, and RtcB DNA ligase.

[0237] In certain embodiments, the method for preparing a target polynucleotide further comprises covalently linking the polynucleotide to be linked, the first polynucleotide, and the second polynucleotide in the presence of a splint oligonucleotide to obtain the target polynucleotide. Furthermore, when a splint oligonucleotide is present, the method for preparing a target polynucleotide further comprises covalently linking the polynucleotide to be linked, the first polynucleotide, and the second polynucleotide in the presence of a splint oligonucleotide to obtain the target polynucleotide.

[0238] In certain embodiments, the method for preparing a target polynucleotide further comprises covalently ligating the polynucleotide to be ligated, the first polynucleotide, and the second polynucleotide in the presence of a splint oligonucleotide and a double-stranded ligase to obtain the target polynucleotide, wherein the double-stranded ligase may be T4 RNA ligase 2.

[0239] In certain embodiments, the splint oligonucleotide comprises a first region and a second region, wherein the first region corresponds to the polynucleotide to be connected, and the second region corresponds to the first polynucleotide and / or the second polynucleotide. The first region may be complementary or reverse complementary to the '3' end sequence of the polynucleotide to be connected, and the second region may be complementary or reverse complementary to the '5' end sequence of the first polynucleotide and / or the second polynucleotide. For example, when the first polynucleotide has a non-PolyA region, the second region may be complementary or reverse complementary to at least the non-PolyA region of the first polynucleotide. For another example, when the first polynucleotide has a non-PolyA region, the second region may only be complementary or reverse complementary to at least part of the non-PolyA region of the first polynucleotide, the second region may also be complementary or reverse complementary to at least part of the non-PolyA region and at least part of the polyA region of the first polynucleotide, the second region may also be complementary or reverse complementary to all non-PolyA regions and at least part of the polyA region of the first polynucleotide, or the second region may be complementary or reverse complementary to all non-PolyA regions and all polyA regions of the first polynucleotide.

[0240] In certain embodiments, step S2 of the method for preparing a target polynucleotide further comprises:

[0241] Step S.2.1.1. Covalently linking the first polynucleotide to the polynucleotide to be linked to obtain a first complex;

[0242] Step S.2.1.2. Covalently link the second polynucleotide to the first complex to obtain the target polynucleotide.

[0243] The first complex may refer to the product of covalently linking the polynucleotide to be linked and the first polynucleotide. For example, step S.2.1.1 may further include covalently linking the 3' end of the polynucleotide to be linked and the 5' end of the first polynucleotide to obtain the first complex.

[0244] In step S.2.1.1, the 3' and 5' end groups of the polynucleotides to be linked can be hydroxyl groups, the 3' end group of the first polynucleotide can be a hydroxyl group, and the 5' end group can be a phosphate group. In step S.2.1.2, the 3' and 5' end groups of the second polynucleotide can be phosphate groups.

[0245] During the ligation process, in order to prevent self-ligation of the polynucleotides, the linking polynucleotide and / or the first polynucleotide and / or the second polynucleotide may be phosphorylated or hydroxylated.

[0246] In step S.2.1.1., obtaining the first complex may further include performing a first modification on the 3' end of the polynucleotide to be linked and / or performing a second modification on the 5' end of the polynucleotide to be linked, and then covalently linking the polynucleotide to be linked to obtain the first complex. In step S.2.1.1., obtaining the first complex may further include performing a first modification on the 3' end of the first polynucleotide and / or performing a second modification on the 5' end of the first polynucleotide, and then covalently linking the polynucleotide to be linked to obtain the first complex. In step S.2.1.2., obtaining the target polynucleotide may further include performing a first modification on the 3' end of the first complex and / or performing a second modification on the 5' end of the first complex, and then covalently linking the polynucleotide to be linked to obtain the target polynucleotide. In step S.2.1.2., obtaining the target polynucleotide may further include performing a first modification on the 3' end of the second polynucleotide and / or performing a second modification on the 5' end of the second polynucleotide, and then covalently linking the polynucleotide to the first complex, and then obtaining the first complex.

[0247] For example, in step S.2.1.1., obtaining the first complex further comprises performing a first modification on the 3' end of the polynucleotide to be connected. For another example, in step S.2.1.1., obtaining the first complex further comprises performing a second modification on the 5' end of the polynucleotide to be connected. For example, in step S.2.1.1., obtaining the first complex further comprises performing a first modification on the 3' end of the first polynucleotide. For another example, in step S.2.1.1., obtaining the first complex further comprises performing a second modification on the 5' end of the first polynucleotide. The first modification and / or the second modification may be a hydroxylation modification or a phosphorylation modification. The phosphorylation modification may be performed using T4PNK enzyme.

[0248] For example, in step S.2.1.2, obtaining the target polynucleotide further includes performing a first modification on the 3' end of the first complex. For another example, in step S.2.1.2., obtaining the target polynucleotide further includes performing a second modification on the 5' end of the first complex. For example, in step S.2.1.2., obtaining the target polynucleotide further includes performing a first modification on the 3' end of the second polynucleotide. For another example, in step S.2.1.2., obtaining the target polynucleotide further includes performing a second modification on the 5' end of the second polynucleotide. The first modification and / or the second modification can be a hydroxylation modification or a phosphorylation modification. The phosphorylation modification can be performed using T4PNK enzyme.

[0249] When the first polynucleotide has a non-Poly A region, step S.2.1.1 of obtaining the first complex may further include covalently linking the first polynucleotide to the polynucleotide to be linked in the presence of a splint oligonucleotide to obtain the first complex.

[0250] In step S.2.1.1., the step may further include obtaining the first complex in the presence of a ligase, wherein the ligase may be a double-stranded ligase, for example, the double-stranded ligase may be T4 RNA ligase 2.

[0251] In step S.2.1.2., the step may further include obtaining the target polynucleotide in the presence of a ligase, wherein the ligase may be a single-stranded ligase, for example, the single-stranded ligase may be T4 RNA ligase 1.

[0252] In certain embodiments, the method of preparing a target polynucleotide further comprises:

[0253] Step S.2.2.1: connecting the first polynucleotide to the second polynucleotide to obtain a second complex;

[0254] Step S.2.2.2: Connect the polynucleotide to be connected with the second complex to obtain the target polynucleotide.

[0255] The second complex may refer to a product obtained by covalently linking the first polynucleotide and the second polynucleotide. For example, the 3' end of the first polynucleotide and the 5' end of the second polynucleotide may be covalently linked to obtain the second complex.

[0256] In step S.2.2.1, the 3' and 5' end groups of the polynucleotides to be linked can be hydroxyl groups, and the 3' and 5' end groups of the first polynucleotide can be hydroxyl groups. In step S.2.2.2, the 3' and 5' end groups of the second polynucleotide can be phosphate groups.

[0257] During the ligation process, in order to prevent self-ligation of the polynucleotides, phosphorylation modification or hydroxylation modification can be performed on the ligating polynucleotide and / or the first polynucleotide and / or the second polynucleotide. For example, in step S.2.2.1., obtaining the second complex further includes performing a first modification on the 3' end of the first polynucleotide and / or performing a second modification on the 5' end of the first polynucleotide, and then covalently linking it with the second polynucleotide to obtain the second complex. For another example, in step S.2.2.1., obtaining the second complex further includes performing a first modification on the 3' end of the second polynucleotide and / or performing a second modification on the 5' end of the second polynucleotide, and then covalently linking it with the first polynucleotide to obtain the second complex. For another example, in step S.2.2.2., obtaining the target polynucleotide further includes performing a first modification on the 3' end of the polynucleotide to be linked and / or performing a second modification on the 5' end of the polynucleotide to be linked, and then covalently linking it with the second complex to obtain the target polynucleotide. For another example, in step S.2.2.2., obtaining the target polynucleotide further includes performing a first modification on the 3' end of the second complex and / or performing a second modification on the 5' end of the second complex, and then covalently linking with the nucleotide to be linked to obtain the target polynucleotide.

[0258] For example, in step S.2.2.1., obtaining the second complex further comprises performing a first modification on the 3' end of the first polynucleotide. For another example, in step S.2.2.1., obtaining the second complex further comprises performing a second modification on the 5' end of the first polynucleotide. For another example, in step S.2.2.1., obtaining the second complex further comprises performing a first modification on the 3' end of the second polynucleotide. For another example, in step S.2.2.1., obtaining the second complex further comprises performing a second modification on the 5' end of the second polynucleotide. The first modification and / or the second modification may be a hydroxylation modification or a phosphorylation modification. The phosphorylation modification may be performed using T4PNK enzyme.

[0259] For another example, in step S.2.2.2., obtaining the target polynucleotide further includes performing a first modification on the 3' end of the polynucleotide to be connected. For another example, in step S.2.2.2., obtaining the target polynucleotide further includes performing a second modification on the 5' end of the polynucleotide to be connected. For another example, in step S.2.2.2., obtaining the target polynucleotide further includes performing a first modification on the 3' end of the second complex. For another example, in step S.2.2.2., obtaining the target polynucleotide further includes performing a second modification on the 5' end of the second complex. The first modification and / or the second modification can be a hydroxylation modification or a phosphorylation modification. The phosphorylation modification can be performed using T4PNK enzyme.

[0260] In some embodiments, when the first polynucleotide has a non-Poly A region, obtaining the target polynucleotide may further include covalently linking the second complex to the polynucleotide to be linked in the presence of a splint oligonucleotide to obtain the first complex.

[0261] In step S.2.2.1., the step may further include obtaining the second complex in the presence of a ligase, wherein the ligase may be a single-strand ligase, for example, the double-strand ligase may be T4 RNA ligase 1.

[0262] Step S.2.2.2. may further include obtaining the target polynucleotide in the presence of a ligase, wherein the ligase may be a double-stranded ligase, for example, T4 RNA ligase 2.

[0263] Non-PolyA region binding solution

[0264] To stabilize nucleic acids and enable the addition of a fixed-length Poly A tail, one approach of the present application is to divide the target polynucleotide into multiple fragments, obtain these fragments separately by chemical synthesis or in vitro transcription, at least two of which contain the sequence to be joined and the first nucleotide, and then synthesize these fragments into a complete target polynucleotide using a ligase. The target polynucleotide obtained in this manner has at least the following advantages:

[0265] (1) By obtaining fragments separately, a PolyA tail of controllable length can be obtained in the target polynucleotide;

[0266] (2) By obtaining the fragments separately, the PolyA tail can be modified in a controllable manner.

[0267] Based on the above scheme, on the one hand, the present application provides a method for preparing a target polynucleotide, which may include the following steps:

[0268] A1. Providing a first polynucleotide, providing a polynucleotide to be connected, wherein the first polynucleotide comprises a non-PolyA region and a PolyA region;

[0269] A2. Covalently linking the polynucleotide to be linked and the first polynucleotide to obtain the target polynucleotide.

[0270] In certain embodiments, step A1 further comprises providing a second polynucleotide, and step A2 further comprises covalently linking the polynucleotide to be linked, the first polynucleotide, and the second polynucleotide to obtain the target polynucleotide. In certain embodiments, the second polynucleotide is polyA.

[0271] In certain embodiments, the first polynucleotide, and / or the second polynucleotide, and / or the polynucleotide to be joined, and / or the target polynucleotide are ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). For example, when the target polynucleotide is RNA (e.g., mRNA), the first polynucleotide, the second polynucleotide, and the polynucleotide to be joined are all RNA. For another example, when the target polynucleotide is DNA, the first polynucleotide, the second polynucleotide, and the polynucleotide to be joined are all deoxyribonucleic acid (DNA).

[0272] In certain embodiments, the first polynucleotide, and / or the second polynucleotide, and / or the polynucleotide to be linked, and / or the target polynucleotide is a combination of ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). For example, the first polynucleotide and the second polynucleotide are RNA, the polynucleotide to be linked is DNA, and the target polynucleotide obtained according to the method of the present application is a combination of ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). For another example, the first polynucleotide includes a polyA region and a non-polyA region, wherein the non-polyA region is DNA and the polyA region is RNA. The first polynucleotide obtained in this way is a combination of DNA and RNA, and the target polynucleotide further obtained is a combination of ribonucleic acid (RNA) and deoxyribonucleic acid (DNA).

[0273] In certain embodiments, the first polynucleotide, and / or the second polynucleotide, and / or the polynucleotide to be joined are chemically synthesized or in vitro transcribed. For example, the first polynucleotide and the second polynucleotide are chemically synthesized, and the polynucleotide to be joined is in vitro transcribed. In another example, the polynucleotide to be joined, the first polynucleotide, and the second polynucleotide are all chemically synthesized.

[0274] The advantages of chemical synthesis for obtaining the first polynucleotide, and / or the second polynucleotide, and / or the polynucleotide to be linked are: (1) chemical synthesis is more controllable, and the length of the obtained polynucleotides can be controlled, especially the first polynucleotide and / or the second polynucleotide, so that the length of the obtained poly A is easily controlled; (2) modification is convenient, and modified nucleotides can be added as required; (3) compared to traditional enzymatic polymerization and in vitro transcription methods, less biological raw materials are used, making it easier to meet GMP requirements. Chemical synthesis of polynucleotides can be carried out using the method described in Example 5.

[0275] In certain embodiments, the first polynucleotide and / or the second polynucleotide may include a polyA region. For example, the first polynucleotide may include a polyA region and a non-polyA region, and the second polynucleotide may be polyA.

[0276] In certain embodiments, the first polynucleotide and the second polynucleotide may both be poly A. For another example, the first polynucleotide and the second polynucleotide each include a poly A region and a non-poly A region.

[0277] In certain embodiments, the first polynucleotide and the second polynucleotide may be of the same length or different lengths, for example, the first polynucleotide is longer than the second polynucleotide, or the first polynucleotide is shorter than the second polynucleotide.

[0278] In certain embodiments, the polynucleotides to be linked may not contain a polyA region.

[0279] In certain embodiments, the target polynucleotide may include a first functional region. For example, the first functional region in mRNA may include one or more of a 3'UTR, an ORF, and a 5'UTR. For another example, the first functional region in DNA may include a coding region or a non-coding region.

[0280] In certain embodiments, at least a portion of the polynucleotide to be linked is covalently linked to the non-polyA region of the first polynucleotide and / or the second polynucleotide to form a first functional region. For example, the first polynucleotide includes a non-polyA region and a polyA region, the second polynucleotide includes a polyA region, wherein the non-polyA region includes a portion of the first functional region (e.g., a portion of the 3'UTR), and the polynucleotide to be linked includes the remaining portion of the first functional region (e.g., another portion of the 3'UTR). When the polynucleotide to be linked, the first polynucleotide, and the second polynucleotide are covalently linked, the first polynucleotide including the non-polyA region and the polynucleotide to be linked are linked to obtain the first functional region.

[0281] In certain embodiments, the method for preparing the target polynucleotide further comprises covalently linking the polynucleotide to be linked, the first polynucleotide, and the second polynucleotide in the presence of a ligase to obtain the target polynucleotide. The ligase may comprise a DNA ligase or an RNA ligase. For example, the ligase may comprise one or more of T4 DNA ligase and T4 RNA ligase. The T4 RNA ligase may comprise one or more of T4 RNA ligase 1 and T4 RNA ligase 2. The ligase may comprise a single-stranded ligase or a double-stranded ligase. For example, the single-stranded ligase may comprise one or more of T4 RNA ligase 1, RM378 RNA ligase, and TS2126 RNA ligase, and the double-stranded ligase may comprise one or more of T4 DNA ligase, T3 DNA ligase, and T4 RNA ligase 2. The ligase may further include one or more of RM378 RNA ligase, TS2126 RNA ligase, E. coli RNA ligase, Mth RNA ligase, RTCB RNA ligase, T3 DNA ligase, T7 DNA ligase, Taq DNA ligase, marine archaea Thermococcus sp DNA ligase, Chlorella virus DNA ligase, and RtcB DNA ligase.

[0282] In certain embodiments, the method for preparing a target polynucleotide further comprises covalently linking the polynucleotide to be linked, the first polynucleotide, and the second polynucleotide in the presence of a splint oligonucleotide to obtain the target polynucleotide. Furthermore, when a splint oligonucleotide is present, the method for preparing a target polynucleotide further comprises covalently linking the polynucleotide to be linked, the first polynucleotide, and the second polynucleotide in the presence of a splint oligonucleotide to obtain the target polynucleotide.

[0283] In certain embodiments, the splint oligonucleotide comprises a first region and a second region, wherein the first region corresponds to the polynucleotide to be connected, and the second region corresponds to the first polynucleotide and / or the second polynucleotide. The first region may be complementary or reverse complementary to the '3' end sequence of the polynucleotide to be connected, and the second region may be complementary or reverse complementary to the '5' end sequence of the first polynucleotide and / or the second polynucleotide. For example, when the first polynucleotide has a non-PolyA region, the second region may be complementary or reverse complementary to at least the non-PolyA region of the first polynucleotide. For another example, when the first polynucleotide has a non-PolyA region, the second region may only be complementary or reverse complementary to at least part of the non-PolyA region of the first polynucleotide, the second region may also be complementary or reverse complementary to at least part of the non-PolyA region and at least part of the polyA region of the first polynucleotide, the second region may also be complementary or reverse complementary to all non-PolyA regions and at least part of the polyA region of the first polynucleotide, or the second region may be complementary or reverse complementary to all non-PolyA regions and all polyA regions of the first polynucleotide.

[0284] In certain embodiments, the second region of the splint oligonucleotide may include a first subregion and a second subregion, wherein the first subregion corresponds to a polyA region and the second subregion corresponds to a non-polyA region. For example, when the first polynucleotide has a non-polyA region, the first subregion is complementary or reverse complementary to at least a portion of the polyA region of the first polynucleotide, and the second subregion is complementary or reverse complementary to at least a portion of the non-polyA region of the first polynucleotide. For another example, when the first polynucleotide has a non-polyA region, the first subregion is complementary or reverse complementary to at least a portion of the polyA region of the first polynucleotide, and the second subregion is complementary or reverse complementary to the entire non-polyA region of the first polynucleotide.

[0285] In certain embodiments, step A2 of the method for preparing a target polynucleotide further comprises:

[0286] Step A.2.1.1. Covalently linking the first polynucleotide to the polynucleotide to be linked to obtain a first complex;

[0287] Step A.2.1.2. Covalently link the second polynucleotide to the first complex to obtain the target polynucleotide.

[0288] The first complex may refer to the product of covalently linking the polynucleotide to be linked and the first polynucleotide. For example, step A.2.1.1 may further include covalently linking the 3' end of the polynucleotide to be linked and the 5' end of the first polynucleotide to obtain the first complex.

[0289] In step A.2.1.1, the 3' and 5' end groups of the polynucleotides to be linked can be hydroxyl groups, the 3' end group of the first polynucleotide can be hydroxyl groups, and the 5' end group can be a phosphate group. In step A.2.1.2, the 3' and 5' end groups of the second polynucleotide can be phosphate groups.

[0290] During the ligation process, in order to prevent self-ligation of the polynucleotides, the linking polynucleotide and / or the first polynucleotide and / or the second polynucleotide may be phosphorylated or hydroxylated.

[0291] In step A.2.1.1., obtaining the first complex may further include performing a first modification on the 3' end of the polynucleotide to be linked and / or performing a second modification on the 5' end of the polynucleotide to be linked, and then covalently linking the polynucleotide to be linked to obtain the first complex. In step A.2.1.1., obtaining the first complex may further include performing a first modification on the 3' end of the first polynucleotide and / or performing a second modification on the 5' end of the first polynucleotide, and then covalently linking the polynucleotide to be linked to obtain the first complex. In step A.2.1.2., obtaining the target polynucleotide may further include performing a first modification on the 3' end of the first complex and / or performing a second modification on the 5' end of the first complex, and then covalently linking the polynucleotide to be linked to obtain the target polynucleotide. In step A.2.1.2., obtaining the target polynucleotide may further include performing a first modification on the 3' end of the second polynucleotide and / or performing a second modification on the 5' end of the second polynucleotide, and then covalently linking the polynucleotide to the first complex, and then obtaining the first complex.

[0292] For example, in step A.2.1.1., obtaining the first complex further comprises performing a first modification on the 3' end of the polynucleotide to be connected. For another example, in step A.2.1.1., obtaining the first complex further comprises performing a second modification on the 5' end of the polynucleotide to be connected. For example, in step A.2.1.1., obtaining the first complex further comprises performing a first modification on the 3' end of the first polynucleotide. For another example, in step A.2.1.1., obtaining the first complex further comprises performing a second modification on the 5' end of the first polynucleotide. The first modification and / or the second modification may be a hydroxylation modification or a phosphorylation modification. The phosphorylation modification may be performed using T4PNK enzyme.

[0293] For example, in step A.2.1.2, obtaining the target polynucleotide further includes performing a first modification on the 3' end of the first complex. For another example, in step A.2.1.2., obtaining the target polynucleotide further includes performing a second modification on the 5' end of the first complex. For example, in step A.2.1.2., obtaining the target polynucleotide further includes performing a first modification on the 3' end of the second polynucleotide. For another example, in step A.2.1.2., obtaining the target polynucleotide further includes performing a second modification on the 5' end of the second polynucleotide. The first modification and / or the second modification can be a hydroxylation modification or a phosphorylation modification. The phosphorylation modification can be performed using T4PNK enzyme.

[0294] When the first polynucleotide has a non-Poly A region, step A.2.1.1 of obtaining the first complex may further include covalently linking the first polynucleotide to the polynucleotide to be linked in the presence of a splint oligonucleotide to obtain the first complex.

[0295] In step A.2.1.1., the step may further include obtaining the first complex in the presence of a ligase, wherein the ligase may be a double-stranded ligase, for example, the double-stranded ligase may be T4 RNA ligase 2.

[0296] In step A.2.1.2., the step may further include obtaining the target polynucleotide in the presence of a ligase, wherein the ligase may be a single-stranded ligase, for example, the single-stranded ligase may be T4 RNA ligase 1.

[0297] In certain embodiments, step A2 of the method for preparing a target polynucleotide further comprises:

[0298] Step A.2.2.1: connecting the first polynucleotide to the second polynucleotide to obtain a second complex;

[0299] Step A.2.2.2: Connect the polynucleotide to be connected with the second complex to obtain the target polynucleotide.

[0300] The second complex may refer to a product obtained by covalently linking the first polynucleotide and the second polynucleotide. For example, the 3' end of the first polynucleotide and the 5' end of the second polynucleotide may be covalently linked to obtain the second complex.

[0301] In order to prevent the polynucleotides from self-ligating during the ligation process, the ligating polynucleotide and / or the first polynucleotide and / or the second polynucleotide may be phosphorylated or hydroxylated.

[0302] During the ligation process, in order to prevent self-ligation of the polynucleotides, phosphorylation modification or hydroxylation modification can be performed on the ligating polynucleotide and / or the first polynucleotide and / or the second polynucleotide. For example, in step A.2.2.1., obtaining the second complex further includes performing a first modification on the 3' end of the first polynucleotide and / or performing a second modification on the 5' end of the first polynucleotide, and then covalently linking it with the second polynucleotide to obtain the second complex. For another example, in step A.2.2.1., obtaining the second complex further includes performing a first modification on the 3' end of the second polynucleotide and / or performing a second modification on the 5' end of the second polynucleotide, and then covalently linking it with the first polynucleotide to obtain the second complex. For another example, in step A.2.2.2., obtaining the target polynucleotide further includes performing a first modification on the 3' end of the polynucleotide to be linked and / or performing a second modification on the 5' end of the polynucleotide to be linked, and then covalently linking it with the second complex to obtain the target polynucleotide. For another example, in step A.2.2.2., obtaining the target polynucleotide further includes performing a first modification on the 3' end of the second complex and / or performing a second modification on the 5' end of the second complex, and then covalently linking with the nucleotide to be linked to obtain the target polynucleotide.

[0303] For example, in step A.2.2.1., obtaining the second complex further includes performing a first modification on the 3' end of the first polynucleotide. For another example, in step A.2.2.1., obtaining the second complex further includes performing a second modification on the 5' end of the first polynucleotide. For another example, in step A.2.2.1., obtaining the second complex further includes performing a first modification on the 3' end of the second polynucleotide. For another example, in step A.2.2.1., obtaining the second complex further includes performing a second modification on the 5' end of the second polynucleotide. The first modification and / or the second modification may be a hydroxylation modification or a phosphorylation modification. The phosphorylation modification may be performed using the T4PNK enzyme.

[0304] For another example, in step A.2.2.2., obtaining the target polynucleotide further includes performing a first modification on the 3' end of the polynucleotide to be connected. For another example, in step A.2.2.2., obtaining the target polynucleotide further includes performing a second modification on the 5' end of the polynucleotide to be connected. For another example, in step A.2.2.2., obtaining the target polynucleotide further includes performing a first modification on the 3' end of the second complex. For another example, in step A.2.2.2., obtaining the target polynucleotide further includes performing a second modification on the 5' end of the second complex. The first modification and / or the second modification can be a hydroxylation modification or a phosphorylation modification. The phosphorylation modification can be performed using T4PNK enzyme.

[0305] In some embodiments, when the first polynucleotide has a non-Poly A region, obtaining the target polynucleotide may further include covalently linking the second complex to the polynucleotide to be linked in the presence of a splint oligonucleotide to obtain the first complex.

[0306] In step A.2.2.1., the step may further include obtaining the second complex in the presence of a ligase, wherein the ligase may be a single-stranded ligase, for example, the single-stranded ligase may be T4 RNA ligase 1.

[0307] In step A.2.2.2., the step may further include obtaining the target polynucleotide in the presence of a ligase, wherein the ligase may be a double-stranded ligase, for example, the double-stranded ligase may be T4 RNA ligase 2.

[0308] In some embodiments, the first polynucleotide and / or the second polynucleotide comprises one or more modified nucleotides.

[0309] wherein a first modified nucleotide is present in the first polynucleotide and a second modified nucleotide is present in the second polynucleotide, wherein the distance of the first modified nucleotide relative to the 3' end of the first polynucleotide on the first polynucleotide is equal to the distance of the second modified nucleotide relative to the 3' end of the second polynucleotide on the second polynucleotide.

[0310] In some embodiments, the modified nucleotides may include base-modified nucleotides, sugar-modified nucleotides, or phosphate-modified nucleotides, or a combination thereof.

[0311] This application also provides the following implementation methods:

[0312] 1. A method for capping a polynucleotide, comprising:

[0313] S1. Provide a composition comprising at least: an uncapped polynucleotide, a capping enzyme and a cap structure analog, wherein:

[0314] Capping enzymes include polynucleotide ligase polypeptides;

[0315] S2. Forming a capping polynucleotide according to the composition of S1.

[0316] 2. The capping method according to embodiment 1, wherein the polynucleotide ligase polypeptide comprises a DNA ligase polypeptide.

[0317] 3. The capping method according to embodiment 2, wherein the DNA ligase polypeptide is a prokaryotic DNA ligase, a prokaryotic DNA ligase variant, or a functional fragment thereof.

[0318] 4. The capping method according to any one of embodiments 2-3, wherein the DNA ligase polypeptide is a bacterial DNA ligase, a bacterial DNA ligase variant, or a functional fragment thereof.

[0319] 5. The capping method according to any one of embodiment 2, wherein the DNA ligase polypeptide is a viral DNA ligase, a viral DNA ligase variant, or a functional fragment thereof.

[0320] 6. The capping method according to any one of embodiment 5, wherein the DNA ligase polypeptide is T4 DNA ligase, a variant thereof, or a functional fragment thereof.

[0321] 7. The capping method of embodiment 1, wherein the polynucleotide ligase polypeptide comprises an RNA ligase polypeptide.

[0322] 8. The capping method according to embodiment 7, wherein the RNA ligase polypeptide is a prokaryotic RNA ligase, a prokaryotic RNA ligase variant, or a functional fragment thereof.

[0323] 9. The capping method according to any one of embodiments 7-8, wherein the RNA ligase polypeptide is a bacterial RNA ligase, a bacterial RNA ligase variant, or a functional fragment thereof.

[0324] 10. The capping method according to any one of embodiment 7, wherein the RNA ligase polypeptide is a viral RNA ligase, a viral RNA ligase variant, or a functional fragment thereof.

[0325] 11. The capping method according to embodiment 10, wherein the RNA ligase polypeptide is T4 RNA ligase, a variant thereof, or a functional fragment thereof.

[0326] 12. The capping method of embodiment 11, wherein the RNA ligase polypeptide is T4 RNA ligase.

[0327] 13. The capping method of embodiment 12, wherein the RNA ligase polypeptide is T4 RNA ligase 1.

[0328] 14. The capping method of any one of embodiments 1 to 13, wherein the capping enzyme is a fusion polypeptide comprising a polynucleotide binding polypeptide fused to a polynucleotide ligase polypeptide.

[0329] 15. The capping method of embodiment 14, wherein the polynucleotide binding polypeptide comprises a DNA binding polypeptide.

[0330] 16. The capping method of embodiment 14, wherein the polynucleotide binding polypeptide comprises an RNA binding polypeptide.

[0331] 17. The capping method according to embodiment 14, wherein the polynucleotide binding polypeptide is selected from one or more of a DNA double-strand binding domain, a DNA single-strand binding domain, an RNA / DNA composite chain binding domain, an RNA single-strand or RNA double-strand binding protein, wherein

[0332] The DNA double-strand binding domain includes Sso7d and / or NF-kappaB p50;

[0333] The RNA / DNA complex chain binding domain includes one or more of ScFV, RNaseH1 (D210N), and HBD of the monoclonal antibody S9.6;

[0334] RNA single-strand binding domains include: Sso7d, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4.

[0335] 18. The capping method of any one of embodiments 14-17, wherein the polynucleotide-binding polypeptide comprises one or more of Sso7d, NF-kappaB p50, ScFV of monoclonal antibody S9.6, RNaseH1 (D210N), HBD, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1, Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4.

[0336] 19. The capping method of embodiment 18, wherein the polynucleotide binding polypeptide is Sso7d.

[0337] 20. The capping method of any one of embodiments 14-19, wherein the C-terminus of the polynucleotide ligase polypeptide is linked to the N-terminus of the polynucleotide binding polypeptide.

[0338] 21. The capping method of any one of embodiments 14-20, wherein the N-terminus of the polynucleotide ligase polypeptide is linked to the C-terminus of the polynucleotide binding polypeptide.

[0339] 22. The capping method of any one of embodiments 14-21, wherein the polynucleotide ligase polypeptide and the polynucleotide binding polypeptide are linked via a first linker.

[0340] 23. The capping method as described in embodiment 22, wherein the N-terminus of the polynucleotide ligase polypeptide is connected to the first linker, and the C-terminus of the polynucleotide binding polypeptide is connected to the first linker.

[0341] 24. The capping method as described in embodiment 22, wherein the C-terminus of the polynucleotide ligase polypeptide is connected to the first linker, and the N-terminus of the polynucleotide binding polypeptide is connected to the first linker.

[0342] 25. The capping method of any one of embodiments 22-24, wherein the first linker is a G4S linker.

[0343] 26. The capping method according to any one of embodiments 1 to 25, wherein the polynucleotide ligase polypeptide comprises the amino acid sequence shown in SEQ ID NOs: 9-12, 21-22.

[0344] 27. The capping method according to any one of embodiments 1 to 26, wherein the polynucleotide-binding polypeptide comprises an amino acid sequence as shown in SEQ ID NOs: 13-17.

[0345] 28. The capping method according to any one of embodiments 1 to 27, wherein the capping enzyme comprises the amino acid sequence shown in SEQ ID NOs: 18-20, 24-29.

[0346] 29. The capping method of any one of embodiments 1-28, wherein the uncapped polynucleotide is DNA or RNA.

[0347] 30. The capping method of any one of embodiments 1-29, wherein the uncapped polynucleotide is chemically synthesized.

[0348] 31. The capping method of any one of embodiments 1-30, wherein the uncapped polynucleotide is 1 to 150 nucleotides in length.

[0349] 32. The capping method according to any one of embodiments 1 to 31, wherein the base of the 5'-terminal nucleotide of the uncapped polynucleotide is guanine.

[0350] 33. The capping method according to any one of embodiments 1 to 32, wherein the 5' terminal nucleotide of the uncapped polynucleotide is phosphorylated.

[0351] 34. The capping method of any one of embodiments 1 to 33, wherein the sugars of the nucleotides of the uncapped polynucleotide are independently selected from ribose and deoxyribose for each position and may contain modifications comprising 2'-O-alkyl, 2'-O-methoxyethyl, 2'-O allyl, 2'-O alkylamine, 2'-fluororibose, and 2'-deoxyribose;

[0352] and / or the bases of the nucleotides of the uncapped polynucleotide are independently selected for each position from adenine, uracil, guanine or cytosine, or an analog of adenine, uracil, guanine or cytosine,

[0353] And the nucleotide modified base can be selected from xanthine, allylaminouracil, allylaminothymidine, hypoxanthine, dioxyadenine, dioxycytosine, dioxyguanine, dioxyuracil, 6-chloropurine nucleoside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil, 5,6-dihydrouracil, 5-[(3-indolyl)propionamide-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxyuracil, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyluracil, 5-propynylaminocytosine, 5-propynylaminouracil, 5-propynyl cytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaadenine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin-16-aminoallyluracil, 3-5-Propynylaminocytosine, 3-6-Propynylaminouracil, cyano-3-aminoallylcytosine, cyano-3-aminoallyllurazolopyrimidine, cyano-5-6-propynylaminocytosine, cyano-5-6-propynylaminouracil, cyano-5-aminoallylcytosine, cyano-5-aminoallyluracil, cyano-7-aminoallyluracil, Dabcyl-5-3-aminoallyluracil, desthiobiotin-16-aminoallyluracil, desthiobiotin-6-aminoallylcytosine, isoguanine, N 1 -ethyl pseudouracil, N 1 -methoxymethyl pseudouracil, N1-methyladenine, N 1 -methyl pseudouracil, N 1 -propyl pseudouracil, N2-methylguanine, N 4 -Biotin-OBEA-cytosine, N4-methylcytosine, N 6-methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thienylcytosine, thienylguanine, thienyluracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6-macroaminoguanine, 5-formamidouracil, 5-ethynyluracil, N 6 -Isopentenyl adenine (i6A), 2-methylthio-N 6 -isopentenyl adenine (ms2i6A), 2-methylthio-N 6 -methyladenine (ms2m6A), N 6 -(cis-hydroxyisopentenyl)adenine (io6A), 2-methylthio-N 6 -(cis-hydroxyisopentenyl)adenine (ms2io6A), N 6 -glycylaminoformyladenine (g6A), N 6 -Threonylaminoformyladenine (t6A), 2-methylthio-N 6 -Threonylaminoformyladenine (ms2t6A), N 6 -methyl-N 6 -Threonylaminoformyladenine (m6t6A), N 6 -hydroxyvalylaminoformyl adenine (hn6A), 2-methylthio-N 6 -Hydroxyvalylcarbamoyladenine (ms2hn6A), N 6 ,N 6 -dimethyladenine (m62A) and N 6 -acetyl adenine (ac6A).

[0354] 35. The capping method according to any one of embodiments 1 to 34, wherein the cap structure analog has the following general formula:

[0355] in:

[0356] R3 is selected from guanine, adenine, cytosine, uracil, a guanine analog, an adenine analog, a cytosine analog, and a uracil analog;

[0357] R4 is (N1p) x N2, wherein N1 and N2 are ribonucleosides, and N1 is the same as or different from N2;

[0358] Each position of p1 is independently a phosphate group, a phosphorothioate, a phosphorodithioate, an alkylphosphonic acid, an arylphosphonic acid, or an N-phosphoramide bond;

[0359] X is an integer from 0 to 8, wherein if X ≥ 2, the ribonucleosides N1 in (N1p)x are identical to or different from each other;

[0360] The R1 and R2 groups are independently selected from -O-alkyl, halogen, acetylamino (AcNH), hydrogen or hydroxy.

[0361] 36. The method of embodiment 35, wherein the sugars in N1 and N2 are independently selected from ribose and deoxyribose for each position and may contain modifications including 2'-O-alkyl, 2'-O-methoxyethyl, 2'-O allyl, 2'-O alkylamine, 2'-fluororibose, and 2'-deoxyribose;

[0362] And / or the bases in N1 and N2 are independently selected from adenine, uracil, guanine or cytosine for each position, or analogs of adenine, uracil, guanine or cytosine, and the nucleotide modified base may be selected from xanthine, allylaminouracil, allylaminothymidine, hypoxanthine, dioxyadenine, dioxycytosine, dioxyguanine, dioxyuracil, 6-chloropurine nucleoside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil , 5,6-dihydrouracil, 5-[(3-indolyl)propionamido-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyl Uracil, 5-propynylaminocytosine, 5-propynylaminouracil, 5-propynylcytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaadenine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin Biotin-16-aminoallyl uracil, 3-5-propynylaminocytosine, 3-6-propynylaminouracil, cyano-3-aminoallyl cytosine, cyano-3-aminoallyl uracil, cyano-5-6-propynylaminocytosine, cyano-5-6-propynylaminouracil, cyano-5-aminoallyl cytosine, cyano-5-aminoallyl uracil, cyano-7-aminoallyl uracil, Dabcyl-5-3-aminoallyl uracil, desthiobiotin-16-aminoallyl uracil, desthiobiotin-6-aminoallyl cytosine, isoguanine, N 1 -ethyl pseudouracil, N 1 -methoxymethyl pseudouracil, N1-methyladenine, N 1 -methyl pseudouracil, N 1 -propyl pseudouracil, N2-methylguanine, N 4 -Biotin-OBEA-cytosine, N4-methylcytosine, N 6-methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thienylcytosine, thienylguanine, thienyluracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6-macroaminoguanine, 5-formamidouracil, 5-ethynyluracil, N 6 -Isopentenyl adenine (i6A), 2-methylthio-N 6 -isopentenyl adenine (ms2i6A), 2-methylthio-N 6 -methyladenine (ms2m6A), N 6 -(cis-hydroxyisopentenyl)adenine (io6A), 2-methylthio-N 6 -(cis-hydroxyisopentenyl)adenine (ms2io6A), N 6 -glycylaminoformyladenine (g6A), N 6 -Threonylaminoformyladenine (t6A), 2-methylthio-N 6 -Threonylaminoformyladenine (ms2t6A), N 6 -methyl-N 6 -Threonylaminoformyladenine (m6t6A), N 6 -hydroxyvalylaminoformyl adenine (hn6A), 2-methylthio-N 6 -Hydroxyvalylcarbamoyladenine (ms2hn6A), N 6 ,N 6 -dimethyladenine (m62A) and N 6 -acetyl adenine (ac6A).

[0363] 37. The capping method according to any one of embodiments 35-36, wherein the cap structure analog is a dinucleotide cap analog or a trinucleotide cap analog.

[0364] 38. The capping method according to any one of embodiments 35-37, wherein the cap structure analog has the following general formula:

[0365] wherein R5 and R6 groups are independently selected from O-alkyl (O-methyl), halogen, tag, hydrogen or hydroxyl;

[0366] X1 and X2 are bases, and X1 and X2 are the same as or different from each other.

[0367] 39. The method of embodiment 38, wherein X1 and / or X2 is guanine or adenine.

[0368] 40. The method of embodiment 39, wherein the cap analog is one of:

[0369] 41. The capping method according to any one of embodiments 1-40, wherein the reaction temperature of step S2 is 5°C to 80°C.

[0370] 42. The method according to embodiment 41, wherein the reaction temperature of step S2 is 25°C to 60°C.

[0371] 43. The capping method according to any one of embodiments 1-42, wherein the composition of step S1 further comprises adenosine triphosphate.

[0372] 44. The capping method according to any one of embodiments 1-43, wherein the composition of step S1 further comprises a DNA / RNA buffer.

[0373] 45. The capping method according to any one of embodiments 1-44, wherein the composition of step S1 further comprises a DNase and / or RNase inhibitor.

[0374] 46. ​​A capping enzyme for capping a polynucleotide, wherein the capping enzyme is a fusion polypeptide comprising a polynucleotide binding polypeptide fused to a polynucleotide ligase polypeptide.

[0375] 47. The capping enzyme of embodiment 46, wherein the polynucleotide ligase polypeptide comprises a DNA ligase polypeptide.

[0376] 48. The capping enzyme of embodiment 47, wherein the DNA ligase polypeptide is a prokaryotic DNA ligase, a prokaryotic DNA ligase variant, or a functional fragment thereof.

[0377] 49. The capping enzyme of embodiment 47, wherein the DNA ligase polypeptide is a bacterial DNA ligase, a bacterial DNA ligase variant, or a functional fragment thereof.

[0378] 50. The capping enzyme of embodiment 47, wherein the DNA ligase polypeptide is a viral DNA ligase, a viral DNA ligase variant, or a functional fragment thereof.

[0379] 51. The capping enzyme of embodiment 50, wherein the DNA ligase polypeptide is T4 DNA ligase, a variant thereof, or a functional fragment thereof.

[0380] 52. The capping enzyme of embodiment 46, wherein the polynucleotide ligase polypeptide comprises an RNA ligase polypeptide.

[0381] 53. The capping enzyme of embodiment 52, wherein the RNA ligase polypeptide is a prokaryotic RNA ligase, a prokaryotic RNA ligase variant, or a functional fragment thereof.

[0382] 54. The capping enzyme of embodiment 52, wherein the RNA ligase polypeptide is a bacterial RNA ligase, a bacterial RNA ligase variant, or a functional fragment thereof.

[0383] 55. The capping enzyme of embodiment 52, wherein the RNA ligase polypeptide is a viral RNA ligase, a viral RNA ligase variant, or a functional fragment thereof.

[0384] 56. The capping enzyme of embodiment 55, wherein the RNA ligase polypeptide is T4 RNA ligase, a variant thereof, or a functional fragment thereof.

[0385] 57. The capping enzyme of embodiment 56, wherein the RNA ligase polypeptide is T4 RNA ligase.

[0386] 58. The capping enzyme of embodiment 57, wherein the RNA ligase polypeptide is T4 RNA ligase 1.

[0387] 59. The capping enzyme of any one of embodiments 46-58, wherein the capping enzyme is a fusion polypeptide comprising a polynucleotide binding polypeptide fused to a polynucleotide ligase polypeptide.

[0388] 60. The capping enzyme of any one of embodiments 46-59, wherein the polynucleotide binding polypeptide comprises a DNA binding polypeptide.

[0389] 61. The capping enzyme of any one of embodiments 46-59, wherein the polynucleotide binding polypeptide comprises an RNA binding polypeptide.

[0390] 62. The capping enzyme of any one of embodiments 46-60, wherein the polynucleotide binding polypeptide is selected from one or more of a DNA double-strand binding domain, a DNA single-strand binding domain, an RNA / DNA composite strand binding domain, an RNA single-strand or RNA double-strand binding protein, wherein

[0391] DNA double-strand binding domains include Sso7d and / or NF-kappaB p50;

[0392] The RNA / DNA complex chain binding domain includes one or more of ScFV, RNaseH1 (D210N), and HBD of the monoclonal antibody S9.6;

[0393] RNA single-strand binding domains include: Sso7d, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4.

[0394] 63. The capping enzyme of any one of embodiments 46-62, wherein the polynucleotide binding polypeptide comprises one or more of Sso7d, NF-kappaB p50, ScFV of monoclonal antibody S9.6, RNaseH1 (D210N), HBD, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4.

[0395] 64. The capping enzyme of embodiment 63, wherein the polynucleotide binding polypeptide is Sso7d.

[0396] 65. The capping enzyme of any one of embodiments 46-64, wherein the C-terminus of the polynucleotide ligase polypeptide is linked to the N-terminus of the polynucleotide binding polypeptide.

[0397] 66. The capping enzyme of any one of embodiments 46-64, wherein the N-terminus of the polynucleotide ligase polypeptide is linked to the C-terminus of the polynucleotide binding polypeptide.

[0398] 67. The capping enzyme of any one of embodiments 46-66, wherein the polynucleotide ligase polypeptide and the polynucleotide binding polypeptide are linked by a first linker.

[0399] 68. The capping enzyme of embodiment 67, wherein the first linker is a G4S linker.

[0400] 69. The capping enzyme of any one of embodiments 46-68, wherein the polynucleotide ligase polypeptide comprises the amino acid sequence shown in SEQ ID NOs: 9-12, 21-22.

[0401] 70. The capping enzyme of any one of embodiments 46-69, wherein the polynucleotide binding polypeptide comprises the amino acid sequence shown in SEQ ID NOs: 13-17.

[0402] 71. The capping enzyme of any one of embodiments 46-70, wherein the capping enzyme comprises the amino acid sequence shown in SEQ ID NOs: 18-20, 24-29.

[0403] 72. A kit comprising the capping enzyme of embodiments 46-71.

[0404] 73. Use of the capping enzyme described in embodiments 46-71 in capping a polynucleotide.

[0405] Example

[0406] Example 1 Polynucleotide ligase capping strategy

[0407] This application utilizes RNA ligase for efficient sequence ligation, allowing RNA fragments to be ligated in a short period of time. Therefore, we selected RNA oligonucleotide sequences, added a phosphate group to the 5' end, and chemically synthesized them. Then, we enzymatically ligated them using RNA ligase 1 to add a cap structure. A schematic diagram of cap ligation is shown in Figure 1.

[0408] Example 2 Construction of capping enzyme

[0409] This application selected RNA ligase 1 derived from T4 phage and RM378 phage for capping test. In order to further improve the enzymatic reaction activity, we added the domain Sso7d at the N-terminus of T4 RNA ligase 1 (T4RL1) that can enhance the affinity with the DNA / RNA hybrid chain substrate to construct the chimeric protein Sso7d-T4RL1. In order to facilitate purification, a 6×His tag was added to the N-terminus of the above protein to construct a fusion protein. Since the tag is close to the core region of the protein, it affects the conformation and function and is not conducive to purification. A flexible amino acid sequence G4S linker (GGGGSGGGGSGGGGS, SEQ ID NO: 23) was added after the His tag to express the protein by fusion.

[0410] We performed capping tests with different RNA ligases 1 (wild-type T4 RNAligase 1, T4 RNA ligase 1 purchased from NEB, RM378 RNA ligase 1, and Sso7d-T4RL1) to compare their performance. Because different enzymes have different optimal reaction temperatures, a temperature gradient was established to screen for the optimal capping temperature for each enzyme.

[0411] Example 3 Ligation Effect of Wild-Type T4 RNA Ligase 1

[0412] First, we tested whether wild-type T4 RNA ligase 1 could ligate RNA and cap analogs. The oligonucleotide (oligo) sequences used are shown in Table 1. A phosphate group was added to the 5' end of Oligo 1, allowing the ligase to form a ligation with the -OH group at the 3' end of the cap structure. The addition of a phosphate group to the 3' end of Oligo 1 prevented self-ligation of the oligonucleotide. The cap analog, LzCap@AG(3'Acm), was selected. Its molecular structure is shown below. We attempted to ligate it to the 5' end of the oligo using an enzymatic reaction.

[0413] The ligation system is shown in Table 2. Mix all components in a centrifuge tube and evenly divide into several tubes. Incubate at different temperatures (25°C, 37°C, 50°C, and 60°C) for 1 hour. Terminate the reaction with 2× RNA loading buffer and denature at 60°C for 5 minutes. Analyze by Urea-PAGE gel electrophoresis. Use wild-type T4 RNA ligase 1 at a final concentration of 6 μM.

[0414] Table 1 RNA sequences used in capping temperature gradient test

[0415] Table 2 Connection system for capping temperature gradient test

[0416] The results of Urea-PAGE detection are shown in Figure 2, where the product represents the linker of Oligo 1 and the cap structure analog, indicating that wild-type T4 RNA ligase 1 can connect the linker of Oligo 1 and the cap structure analog.

[0417] Example 4 Ligation Effect of Wild-Type RM378 RNA Ligase 1

[0418] First, we tested whether wild-type RM378 RNA ligase 1 could ligate RNA and cap analogs. The experimental procedure was the same as in Example 3, except that wild-type T4 RNA ligase 1 was replaced with wild-type RM378 RNA ligase 1. The results of the Urea-PAGE assay are shown in Figure 3. The absence of distinct "Production" bands on the Urea-PAGE gel image of RM378RL1 indicates that capping was not effective within the established temperature gradient.

[0419] Example 5: Ligation Effects of Different T4 RNA Ligase 1

[0420] We performed capping tests with different T4 RNA ligase 1s (T4 RNA ligase 1 purchased from NEB and Sso7d-T4RL1) to compare their differences. The experimental method was the same as in Example 3, except that wild-type T4 RNA ligase 1 was replaced with T4 RNA ligase 1 purchased from NEB (NEB, M0204S) and Sso7d-T4 RNA ligase 1. NEB-T4RL1 was purchased commercially at a concentration of 10,000 units / ml. According to the instructions, 10 units were added to a 10 μl reaction system.

[0421] The Urea-PAGE detection results are shown in FIG4 . The grayscale analysis of the image was performed using Image J software, and the ligation efficiency was calculated as follows: Production grayscale value / (Oligo 1 grayscale value+Production grayscale value). The results are shown in Table 3 .

[0422] Table 3 Results of different enzyme capping temperature gradient test ligation efficiency

[0423] From the analysis of the ligation efficiency results, the optimal temperature for capping of Sso7d-T4RL1 was 25°C, and the highest ligation efficiency was 32.4%, which was most suitable for capping reaction.

[0424] Example 6 Comparison of RNA ligase ligation with conventional oligonucleotides

[0425] To further verify whether each enzyme has the same effect on ligating conventional oligonucleotides and cap analogs, we continued to test the effects of the four enzymes on oligonucleotide ligation at their respective optimal temperatures. The oligonucleotide sequences used are shown in Table 4. A phosphate group was added to the 5' end of Oligo 3, which can be connected to the 3' end -OH of Oligo 2 under the action of the ligase. A phosphate group was also added to the 3' end of the Oligo 1 sequence to prevent the Oligo from self-ligating. A FAM group was added to the 5' end of Oligo 2, which can display bands on Urea-PAGE gel without additional staining. The connection diagram is shown in Figure 5. When Oligo 2 and Oligo 3 are successfully connected, the 5' end of the connection product is FAM-modified, which can be judged by fluorescence imaging and molecular weight changes.

[0426] The ligation system is shown in Table 5. All components were mixed in a centrifuge tube and evenly divided into several tubes. The mixture was incubated at different temperatures (25°C, 37°C, 45°C, 50°C, and 60°C) for 1 hour. The reaction was terminated with 2× RNA loading buffer and denatured at 60°C for 5 minutes. The samples were analyzed by Urea-PAGE gel electrophoresis.

[0427] Table 4 RNA sequences used for Oligo temperature gradient ligation test

[0428] Table 5 Connection system for Oligo temperature gradient connection test

[0429] The Urea-PAGE detection results are shown in FIG6 . The grayscale analysis of the image was performed using Image J software, and the ligation efficiency was calculated as follows: Production grayscale value / (Oligo 2 grayscale value+Production grayscale value). The results are shown in Table 6 .

[0430] Table 6 Results of oligo ligation efficiency test with different enzymes at temperature gradient

[0431] Ligation efficiency analysis revealed that when ligating conventional oligonucleotides, Sso7d-T4RL1 had an optimal temperature of 37°C, with a maximum ligation efficiency of 87.6%. Activity gradually decreased with increasing temperature. NEB-T4RL1 had an optimal temperature of 43°C, with a maximum ligation efficiency of 86.7%. However, activity decreased rapidly with increasing temperature, reaching only 40% of Sso7d-T4RL1 at 60°C. T4RL1 also had an optimal temperature of 43°C, with a maximum ligation efficiency of only 34.6%. RM378RL1 had an optimal temperature of 60°C, with a maximum ligation efficiency of 95.7%. In conventional oligonucleotide ligation testing, RM378RL1 performed best, achieving the highest ligation efficiency, followed by Sso7d-T4RL1 and NEB-T4RL1. These two enzymes exhibited similar ligation efficiencies at their optimal temperatures. Wild-type T4RL1 performed the worst, achieving the lowest ligation efficiency.

[0432] The above tests show that different RNA ligase 1 enzymes have different preferences when connecting conventional oligonucleotides and connecting cap analogs. Sso7d-T4RL1 is most suitable for connecting cap analogs, and RM378 is most suitable for connecting conventional oligonucleotides.

[0433] Example 7 RNA ligase ligates different oligonucleotides and different cap structure analogs

[0434] Sso7d-T4RL1 was selected as the enzyme for the capping reaction, and the reaction conditions were set to incubate at 25°C for 1 hour. In order to further improve the connection efficiency, we optimized the capping reaction conditions. Other oligo sequences and 3' modification groups were selected to test the universality of the capping reaction. The oligo sequences used are shown in Table 7. A phosphate group was added to the 5' of Oligo 4 to facilitate connection with the cap, and a biotin group was added to the 3' to prevent self-connection. The final concentration of Oligo was controlled to 10μM, LzCap @ The final concentration gradient of AG (3′Acm) was set from 0 to 1000 μM, and the reaction conditions were shown in Table 8.

[0435] Table 7 RNA sequences used for cap analog concentration testing

[0436] Table 8 Connection system for hat analog concentration test

[0437] Incubate at 25°C for 1 hour. Terminate the reaction with 2× RNA loading buffer and denature at 60°C for 5 minutes. Analyze by Urea-PAGE gel electrophoresis.

[0438] The Urea-PAGE detection results are shown in Figure 7. The grayscale analysis of the image was performed using Image J software, and the ligation efficiency was calculated as Production grayscale value / (Oligo 4 grayscale value + Production grayscale value). The results are shown in Table 9. The results were analyzed using Prism 8 software, and the analysis curve was drawn, as shown in Figure 8.

[0439] Table 9 Cap analog concentration test connection efficiency results

[0440] According to the analytical curve, the connection efficiency has passed the inflection point when the final concentration of the cap analog is below 200 μM. Above 200 μM, the connection efficiency slowly increases with the increase of the cap analog concentration, and reaches the maximum connection efficiency above 500 μM. The best connection effect can be obtained when the final concentration of the cap analog in the capping enzymatic reaction is above 500 μM and the oligo concentration is 10 μM. In addition, compared with the previous experiment of connecting Oligo 1 and LzCap using Sso7d-T4RL1, when the final concentration of LzCap was 750 μM, the connection efficiency of this experiment was significantly higher than the previous one. Since the conditions are the same except for the oligo sequence and modification used, it is speculated that the reason for the difference in the two connection efficiencies is the different oligo sequences and structures used.

[0441] To further verify the connection effect of Sso7d-T4RL1 on different oligos and different cap structure analogs, we selected two receptor RNAs and 6 different cap structure analogs for connection tests. The specific information is shown in Table 10.

[0442] Table 10 Different cap structure analogs and oligo information

[0443] The ligation system is shown in Table 11. Incubate at 25°C for 1 hour. Terminate the reaction with 2× RNA loading buffer and denature at 60°C for 5 minutes. Analyze by Urea-PAGE gel electrophoresis.

[0444] Table 11 RNA capping and ligation system

[0445] The Urea-PAGE assay results are shown in Figure 8. Based on the size of the electrophoretic bands, it can be determined that different cap structure analogs and different receptor sequences can be successfully ligated, and the ligation efficiency of Oligo 4 is significantly higher than that of Oligo 1. Grayscale analysis of the images was performed using Image J software, and ligation efficiency was calculated as: Production grayscale value / (Oligo grayscale value + Production grayscale value). The results are shown in Table 12. Oligo 4 achieved a ligation efficiency of approximately 90% for various cap analogs, while Oligo 1 had a ligation efficiency of only approximately 30%. Different 5' bases have a significant impact on ligation efficiency, with G having a higher ligation efficiency than A.

[0446] Table 12 Capping efficiency of different cap structure analogs and receptor RNA

[0447] The above tests show that Sso7d-T4RL1 enzyme can be applied to the connection between different cap structure analogs and different oligonucleotide receptors.

[0448] Example 8 HiBiT mRNA ligation preparation

[0449] Luciferase NanoBiT can be separated into two fragments: a short fragment, HiBiT, and a long fragment, LgBiT. Separately, the two fragments lack luciferase activity. When mixed, they spontaneously assemble into the complete luciferase, producing fluorescence in the presence of substrate. We leveraged this property to validate our developed capping method and synthesize HiBiT mRNA. When the HiBiT mRNA is properly translated, the HiBiT fragments are obtained. Detectable chemiluminescence in the presence of LgBiT protein and substrate confirms that we have synthesized fully functional HiBiT mRNA.

[0450] HiBiT mRNA synthesis method:

[0451] First, the HiBiT sequence was split into three segments, namely Oligo H1 (28 nt), Oligo H2 (38 nt), and Oligo H3 (39 nt). At the same time, a DNA sequence that was reverse complementary to the above Oligo 1 and Oligo 2 sequences was designed and named adaptor H. The required sequence is shown in Table 13.

[0452] Table 13 Sequences required for HiBiT

[0453] First, add caps, Oligo H1 and LzCap @ AG (3'Acm) are connected together, and the connection system is shown in Table 14:

[0454] Table 14 HiBiT capping system

[0455] The mixture prepared according to the above system was placed in a metal bath and incubated at 25°C for 60 minutes. After the reaction, the reaction was terminated with 2× RNA loading buffer, denatured at 60°C for 5 minutes, and analyzed by Urea-PAGE gel electrophoresis. The results of Urea-PAGE are shown in Figure 10.

[0456] A portion of the ligation mixture was taken for mass spectrometry detection, and the results are shown in Figure 11. The predicted molecular weight of Oligo H1 is 9222.57, and the predicted molecular weight of the ligation product of Oligo H1 and LzCap is 10405.32. The mass spectrometry results show that the detected molecular weights of the main components contained in the ligation mixture are 9223.0 and 10405.7, corresponding to Oligo H1 and the ligation product of Oligo H1 and LzCap, respectively, with differences of 0.05‰ and 0.04‰, indicating that Oligo H1 was successfully capped. @The AG (3′Acm) ligation product was recovered and diluted to 100 μM with enzyme-free sterile water.

[0457] The next step is tailing ligation, using T4 RNA ligase 1 to connect Oligo H2 and Oligo H3 together. The ligation system is shown in Table 15:

[0458] Table 15 Hibit tailing system

[0459] The mixture was prepared and reacted at 25°C for 2 hours. The reaction was terminated with 2× RNA loading buffer and denatured at 60°C for 5 minutes. The resulting mixture was then analyzed by Urea-PAGE gel electrophoresis.

[0460] The results of Urea-PAGE detection are shown in Figure 12, which show that Oligo H2 and Oligo H3 can be effectively connected. The grayscale analysis of the image was performed using Image J software, and the connection efficiency was calculated as Production grayscale value / (Oligo grayscale value + Production grayscale value), which is about 50%. We recovered the 77nt connection product and diluted it to 100 μM with enzyme-free sterile water.

[0461] Then, T4PNK was used to add a phosphate group to the 5' end of the Oligo H2-Oligo H3 ligation product and to bind to the LzCap @ The phosphate group was removed from the 3' end of AG(3'Acm)-Oligo 1. The reaction system is shown in Table 16:

[0462] Table 16 PNK reaction system

[0463] The above system was prepared into a mixed solution and reacted at 37°C for 30 minutes. After the reaction, it was transferred to 65°C for 15 minutes to inactivate the T4PNK enzyme. The entire reaction mixture was put into the next ligation system.

[0464] Next step is to connect LzCap @ AG (3'Acm) -Oligo H1 and Oligo H2-Oligo H3 connection products, the connection system is shown in Table 17:

[0465] Table 17 HiBiT connection system

[0466] The mixture prepared according to the above system was placed in a PCR instrument and programmed to repeat 5 cycles of (52°C for 30 seconds, 37°C for 5 minutes). After the reaction, 1 / 10 volume of DNase I and 10× buffer were added, and the mixture was incubated at 37°C for 15 minutes. The reaction was terminated with 2× RNA loading buffer and denatured at 60°C for 5 minutes. Analysis was then performed using Urea-PAGE gel electrophoresis. As shown in Figure 13, the target product of approximately 108 nt was obtained, and the HiBiT mRNA target band was recovered.

[0467] The recovered product was subjected to mass spectrometry, and the test results are shown in Figure 14. The predicted molecular weight of the recovered product was 35575.14, and the detected molecular weight was 35582.9, with a difference of 0.22‰, indicating that the recovered product is the desired Hibit mRNA. The recovered product was subjected to Urea-PAGE, and the test results are shown in Figure 15. The recovered product has a single band and a purity of greater than 95%. The cells will be transfected for subsequent activity testing.

[0468] Example 9 Synthesis of oligonucleotides

[0469] The oligonucleotides in the aforementioned examples can be synthesized using the method in this example.

[0470] Specifically, RNA synthesis uses the solid-phase phosphoramidite method, which goes through a four-step cycle of "deprotection, coupling reaction, capping reaction, and oxidation reaction" (the synthesis efficiency of each cycle is ≥98%). Synthesis is carried out at room temperature according to the set parameters. The synthesis equipment is sealed, the synthesis humidity is ≤30%, and the temperature is 15-25°C. The first base at the 3' end of the oligonucleotide is bound to the solid-phase carrier CPG (Controlled Pore Glass), and then synthesized in the direction from 3'→5'. Adjacent nucleotides are connected by 3'→5' phosphate bonds. Each cycle requires the highly efficient and chemically active nucleotide 3'-phosphite tetrazolium, which becomes a stable pentavalent phosphate triester after oxidation, thereby forming a more stable structure. After multiple cycles of reaction, a nucleic acid chain with a specific sequence is obtained, which is the oligonucleotide.

[0471] The specific method is as follows:

[0472] 1. Synthesis

[0473] Prepare the deprotecting agent, activating coupling agent, capping agent, and oxidizing agent used in the synthesis. The deprotecting agent is a 3% (w / v) dichloroacetic acid solution in dichloromethane, used in a volume of 200 μl each time; the activating coupling agent is a 0.3M ethylmercaptotetrazole solution in acetonitrile, used in a volume of 60 μl each time; the capping agent CAPA is a 10% (v / v) acetic anhydride solution in tetrahydrofuran, used in a volume of 120 μl each time; the capping agent CAPB is a 16% (v / v) 1-methylimidazole solution in tetrahydrofuran, used in a volume of 120 μl each time. CAPA and CAPB are automatically added simultaneously by the instrument. The oxidizing agent is a mixture of 0.05M iodine in tetrahydrofuran / pyridine / ultrapure water (v / v / v = 7 / 2 / 1), used in a volume of 160 μl each time. Dissolve 10g of RNA phosphoramidite monomer in 150ml of anhydrous acetonitrile and protect with argon gas, using a volume of 55 μl each time.

[0474] Load all the above synthesis reagents onto the RNA nucleic acid synthesizer and set the following parameters.

[0475] A 500 nmol solid phase Frit support was loaded onto the column of the synthesizer.

[0476] Check the equipment pressure, reagent bottle pressure, reagent dosage and other parameters. After confirming that they are correct, click the "Start Synthesis" button to start the synthesis. During the synthesis process, check the reagent dosage and equipment operating status until the synthesis is completed.

[0477] 2. Aminolysis and deprotection

[0478] After the synthesis, the Frit vector connected to the RNA nucleic acid was placed in 1.5 ml of APA ammoniolysis solution (ammonia water: n-propylamine: water v / v / v = 2 / 2 / 1) and heated to 65°C for 70 minutes. After the reaction, it was cooled to room temperature.

[0479] The turbid liquid after the above aminolysis was filtered to obtain a concentrated ammonia solution containing RNA nucleic acid, which was then concentrated to a dry powder state. 125 μl of dimethyl sulfoxide (DMSO) was then added to completely dissolve the RNA nucleic acid. After dissolution, 125 μl of triethylamine trihydrofluoride was added and heated to 65°C for 150 minutes.

[0480] Slowly add 500 μl of 2 M ammonium bicarbonate solution in batches to the crude nucleic acid obtained after deprotection to quench the unreacted triethylamine trihydrofluoride. When no obvious bubbles are generated during the process, shake it with a vortex mixer. Once the solution is clear, it can be used for subsequent purification.

[0481] 3. Purification

[0482] The white solid was dissolved in enzyme-free sterile water and purified using high-performance liquid chromatography (HPLC) with a mobile phase consisting of acetonitrile and a 0.1 M triethylamine carbonate (TEAB) solution. Using a proportional valve, the acetonitrile ratio was increased from 8% to 15% while the TEAB solution ratio was decreased from 92% to 85% over 30 minutes, resulting in highly purified long-chain RNA nucleic acid.

[0483] The foregoing detailed description is provided by way of explanation and example and is not intended to limit the scope of the appended claims. Various changes to the embodiments listed in the present application are obvious to those skilled in the art and are intended to fall within the scope of the appended claims and their equivalents.

Claims

1. A method for capping a polynucleotide, comprising: S1. Provide a composition comprising at least: an uncapped polynucleotide, a capping enzyme, and a cap structure analog, wherein the capping enzyme comprises a polynucleotide ligase polypeptide; S2. Forming a capping polynucleotide according to the composition of S1.

2. The capping method of claim 1, wherein the polynucleotide ligase polypeptide comprises a DNA ligase polypeptide. 3 . The capping method according to claim 2 , wherein the DNA ligase polypeptide is a prokaryotic DNA ligase, a prokaryotic DNA ligase variant, or a functional fragment thereof.

4. The capping method according to any one of claims 2 to 3, wherein the DNA ligase polypeptide is a bacterial DNA ligase, a bacterial DNA ligase variant or a functional fragment thereof.

5. The capping method according to claim 2, wherein the DNA ligase polypeptide is a viral DNA ligase, a viral DNA ligase variant or a functional fragment thereof.

6. The capping method according to claim 5, wherein the DNA ligase polypeptide is T4 DNA ligase, a variant thereof or a functional fragment thereof.

7. The capping method of claim 1, wherein the polynucleotide ligase polypeptide comprises an RNA ligase polypeptide.

8. The capping method according to claim 7, wherein the RNA ligase polypeptide is a prokaryotic RNA ligase, a prokaryotic RNA ligase variant or a functional fragment thereof.

9. The capping method according to any one of claims 7 to 8, wherein the RNA ligase polypeptide is a bacterial RNA ligase, a bacterial RNA ligase variant or a functional fragment thereof. 10 . The capping method according to claim 7 , wherein the RNA ligase polypeptide is a viral RNA ligase, a viral RNA ligase variant, or a functional fragment thereof. The capping method according to claim 10 , wherein the RNA ligase polypeptide is T4 RNA ligase, a variant thereof, or a functional fragment thereof. The capping method according to claim 11 , wherein the RNA ligase polypeptide is T4 RNA ligase. The capping method according to claim 12 , wherein the RNA ligase polypeptide is T4 RNA ligase 1.

14. The capping method according to any one of claims 1 to 13, wherein the capping enzyme is a fusion polypeptide comprising a polynucleotide binding polypeptide fused to a polynucleotide ligase polypeptide.

15. The capping method of claim 14, wherein the polynucleotide binding polypeptide comprises a DNA binding polypeptide.

16. The capping method of claim 14, wherein the polynucleotide binding polypeptide comprises an RNA binding polypeptide.

17. The capping method according to claim 14, wherein the polynucleotide binding polypeptide is selected from one or more of a DNA double-strand binding domain, a DNA single-strand binding domain, an RNA / DNA composite chain binding domain, an RNA single-strand or RNA double-strand binding protein, wherein The DNA double-strand binding domain includes Sso7d and / or NF-kappaB p50; The RNA / DNA complex chain binding domain includes one or more of ScFV, RNaseH1 (D210N), and HBD of the monoclonal antibody S9.6; The RNA single-strand binding domain includes: Sso7d, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, RDE-4.

18. The capping method according to any one of claims 14 to 17, wherein the polynucleotide binding polypeptide comprises one or more of Sso7d, NF-kappaB p50, ScFV of monoclonal antibody S9.6, RNaseH1 (D210N), HBD, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4. The capping method of claim 18 , wherein the polynucleotide binding polypeptide is Sso7d.

20. The capping method of any one of claims 14 to 19, wherein the C-terminus of the polynucleotide ligase polypeptide is linked to the N-terminus of the polynucleotide binding polypeptide.

21. The capping method of any one of claims 14 to 20, wherein the N-terminus of the polynucleotide ligase polypeptide is linked to the C-terminus of the polynucleotide binding polypeptide.

22. The capping method of any one of claims 14 to 21, wherein the polynucleotide ligase polypeptide and the polynucleotide binding polypeptide are linked via a first linker.

23. The capping method of claim 22, wherein the N-terminus of the polynucleotide ligase polypeptide is linked to the first linker, and the C-terminus of the polynucleotide binding polypeptide is linked to the first linker.

24. The capping method of claim 22, wherein the C-terminus of the polynucleotide ligase polypeptide is linked to the first linker, and the N-terminus of the polynucleotide binding polypeptide is linked to the first linker.

25. The capping method according to any one of claims 22 to 24, wherein the first linker is a G4S linker.

26. The capping method according to any one of claims 1 to 25, wherein the polynucleotide ligase polypeptide comprises the amino acid sequence shown in SEQ ID NOs: 9 to 12, 21 to 22.

27. The capping method according to any one of claims 1 to 26, wherein the polynucleotide binding polypeptide comprises the amino acid sequence shown in SEQ ID NO: 13-17.

28. The capping method according to any one of claims 1 to 27, wherein the capping enzyme comprises the amino acid sequence shown in SEQ ID NOs: 18-20, 24-29.

29. The capping method of any one of claims 1 to 28, wherein the uncapped polynucleotide is DNA or RNA.

30. The capping method of any one of claims 1 to 29, wherein the uncapped polynucleotide is chemically synthesized.

31. The capping method according to any one of claims 1 to 30, wherein the uncapped polynucleotide is 1 to 150 nucleotides in length.

32. The capping method according to any one of claims 1 to 31, wherein the base of the 5' terminal nucleotide of the uncapped polynucleotide is guanine.

33. The capping method according to any one of claims 1 to 32, wherein the 5' terminal nucleotide of the uncapped polynucleotide is modified by phosphorylation.

34. The capping method of any one of claims 1 to 33, wherein the sugars of the nucleotides of the uncapped polynucleotide are independently selected from ribose and deoxyribose for each position and may contain modifications comprising 2'-O-alkyl, 2'-O-methoxyethyl, 2'-O allyl, 2'-O alkylamine, 2'-fluororibose, and 2'-deoxyribose; and / or the bases of the nucleotides of the uncapped polynucleotide are independently selected for each position from adenine, uracil, guanine or cytosine, or an analog of adenine, uracil, guanine or cytosine, And the nucleotide modified base can be selected from xanthine, allylaminouracil, allylaminothymidine, hypoxanthine, dioxyadenine, dioxycytosine, dioxyguanine, dioxyuracil, 6-chloropurine nucleoside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil, 5,6-dihydrouracil, 5-[(3-indolyl)propionamide-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxyuracil, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyluracil, 5-propynylaminocytosine, 5-propynylaminouracil, 5-propynyl cytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaadenine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin-16-aminoallyluracil, 3-5-Propynylaminocytosine, 3-6-Propynylaminouracil, cyano-3-aminoallylcytosine, cyano-3-aminoallyllurazolopyrimidine, cyano-5-6-propynylaminocytosine, cyano-5-6-propynylaminouracil, cyano-5-aminoallylcytosine, cyano-5-aminoallyluracil, cyano-7-aminoallyluracil, Dabcyl-5-3-aminoallyluracil, desthiobiotin-16-aminoallyluracil, desthiobiotin-6-aminoallylcytosine, isoguanine, N 1 -ethyl pseudouracil, N 1 -methoxymethyl pseudouracil, N1-methyladenine, N 1 -methyl pseudouracil, N 1 -propyl pseudouracil, N2-methylguanine, N 4 -Biotin-OBEA-cytosine, N4-methylcytosine, N 6 -methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thienylcytosine, thienylguanine, thienyluracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6-macroaminoguanine, 5-formamidouracil, 5-ethynyluracil, N 6 -Isopentenyl adenine (i6A), 2-methylthio-N 6 -isopentenyl adenine (ms2i6A), 2-methylthio-N 6 -methyladenine (ms2m6A), N 6 -(cis-hydroxyisopentenyl)adenine (io6A), 2-methylthio-N 6 -(cis-hydroxyisopentenyl)adenine (ms2io6A), N 6 -glycylaminoformyladenine (g6A), N 6 -Threonylaminoformyladenine (t6A), 2-methylthio-N 6 -Threonylaminoformyladenine (ms2t6A), N 6 -methyl-N 6 -Threonylaminoformyladenine (m6t6A), N 6 -hydroxyvalylaminoformyl adenine (hn6A), 2-methylthio-N 6 -Hydroxyvalylcarbamoyladenine (ms2hn6A), N 6 ,N 6 -dimethyladenine (m62A) and N 6 -acetyl adenine (ac6A).

35. The capping method according to any one of claims 1 to 34, wherein the cap structure analog has the following general formula: in: R3 is selected from guanine, adenine, cytosine, uracil, a guanine analog, an adenine analog, a cytosine analog, and a uracil analog; R4 is (N1p) x N2, wherein N1 and N2 are ribonucleosides, and N1 is the same as or different from N2; Each position of p1 is independently a phosphate group, a phosphorothioate, a phosphorodithioate, an alkylphosphonic acid, an arylphosphonic acid, or an N-phosphoramide bond; X is an integer from 0 to 8, wherein if X ≥ 2, the ribonucleosides N1 in (N1p1)x are identical to or different from each other; The R1 and R2 groups are independently selected from -O-alkyl, halogen, acetylamino (AcNH), hydrogen or hydroxy.

36. The method of claim 35, wherein the sugars in N1 and N2 are independently selected from ribose and deoxyribose for each position and may contain modifications including 2'-O-alkyl, 2'-O-methoxyethyl, 2'-O allyl, 2'-O alkylamine, 2'-fluororibose, and 2'-deoxyribose; And / or the bases in N1 and N2 are independently selected from adenine, uracil, guanine or cytosine for each position, or analogs of adenine, uracil, guanine or cytosine, and the nucleotide modified base may be selected from xanthine, allylaminouracil, allylaminothymidine, hypoxanthine, dioxyadenine, dioxycytosine, dioxyguanine, dioxyuracil, 6-chloropurine nucleoside, N6-methyladenine, methylpseudouracil, 2-thiocytosine, 2-thiouracil, 5-methyluracil, 4-thiothymidine, 4-thiouracil, 5,6-dihydro-5-methyluracil , 5,6-dihydrouracil, 5-[(3-indolyl)propionamido-N-allyl]uracil, 5-aminoallylcytosine, 5-aminoallyluracil, 5-bromouracil, 5-bromocytosine, 5-carboxycytosine, 5-carboxyuracil, 5-fluorouracil, 5-formylcytosine, 5-formyluracil, 5-hydroxycytosine, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5-hydroxyuracil, 5-iodocytosine, 5-iodouracil, 5-methoxycytosine, 5-methoxyuracil, 5-methylcytosine, 5-methyl Uracil, 5-propynylaminocytosine, 5-propynylaminouracil, 5-propynylcytosine, 5-propynyluracil, 6-azacytosine, 6-azauracil, 6-chloropurine, 6-thioguanine, 7-deazaadenine, 7-deazaguanine, 7-deaza-7-propylaminoadenine, 7-deaza-7-propynylaminoadenine, 8-azaadenine, 8-azidoadenine, 8-chloroadenine, 8-oxoadenine, 8-oxoguanine, biotin-16-7-deaza-7-propynylaminoguanine, biotin-16-aminoallylcytosine, biotin Biotin-16-aminoallyl uracil, 3-5-propynylaminocytosine, 3-6-propynylaminouracil, cyano-3-aminoallyl cytosine, cyano-3-aminoallyl uracil, cyano-5-6-propynylaminocytosine, cyano-5-6-propynylaminouracil, cyano-5-aminoallyl cytosine, cyano-5-aminoallyl uracil, cyano-7-aminoallyl uracil, Dabcyl-5-3-aminoallyl uracil, desthiobiotin-16-aminoallyl uracil, desthiobiotin-6-aminoallyl cytosine, isoguanine, N 1 -ethyl pseudouracil, N 1 -methoxymethyl pseudouracil, N1-methyladenine, N 1 -methyl pseudouracil, N 1 -propyl pseudouracil, N2-methylguanine, N 4 -Biotin-OBEA-cytosine, N4-methylcytosine, N 6 -methyladenine, 06-methylguanine, pseudoisocytosine, pseudouracil, thienylcytosine, thienylguanine, thienyluracil, xanthine, 3-deazaadenine, 2,6-diaminoadenine, 2,6-macroaminoguanine, 5-formamidouracil, 5-ethynyluracil, N 6 -Isopentenyl adenine (i6A), 2-methylthio-N 6 -isopentenyl adenine (ms2i6A), 2-methylthio-N 6 -methyladenine (ms2m6A), N 6 -(cis-hydroxyisopentenyl)adenine (io6A), 2-methylthio-N 6 -(cis-hydroxyisopentenyl)adenine (ms2io6A), N 6 -glycylaminoformyladenine (g6A), N 6 -Threonylaminoformyladenine (t6A), 2-methylthio-N 6 -Threonylaminoformyladenine (ms2t6A), N 6 -methyl-N 6 -Threonylaminoformyladenine (m6t6A), N 6 -hydroxyvalylaminoformyl adenine (hn6A), 2-methylthio-N 6 -Hydroxyvalylcarbamoyladenine (ms2hn6A), N 6 ,N 6 -dimethyladenine (m62A) and N 6 -acetyl adenine (ac6A).

37. The capping method according to any one of claims 35-36, wherein the cap structure analog is a dinucleotide cap analog or a trinucleotide cap analog.

38. The capping method according to any one of claims 35 to 37, wherein the cap structure analog has the following general formula: in, R5 and R6 groups are independently selected from O-alkyl (O-methyl), halogen, tag, hydrogen or hydroxyl; X1 and X2 are bases, and X1 and X2 are the same as or different from each other.

39. The method of claim 38, wherein X1 and / or X2 is guanine or adenine.

40. The method of claim 39, wherein the cap structure analog is selected from one of the following: The capping method according to any one of claims 1 to 40, wherein the reaction temperature in step S2 is 5°C to 80°C. The method according to claim 41 , wherein the reaction temperature of step S2 is 25° C. to 60° C.

43. The capping method according to any one of claims 1 to 42, wherein the composition of step S1 further comprises adenosine triphosphate. The capping method according to any one of claims 1 to 43, wherein the composition of step S1 further comprises a DNA / RNA buffer. The capping method according to any one of claims 1 to 44, wherein the composition of step S1 further comprises a DNase and / or RNase inhibitor.

46. ​​A capping enzyme for capping a polynucleotide, wherein the capping enzyme is a fusion polypeptide comprising a polynucleotide binding polypeptide fused to a polynucleotide ligase polypeptide.

47. The capping enzyme of claim 46, wherein the polynucleotide ligase polypeptide comprises a DNA ligase polypeptide. The capping enzyme according to claim 47 , wherein the DNA ligase polypeptide is a prokaryotic DNA ligase, a prokaryotic DNA ligase variant or a functional fragment thereof. The capping enzyme of claim 47 , wherein the DNA ligase polypeptide is a bacterial DNA ligase, a bacterial DNA ligase variant, or a functional fragment thereof.

50. The capping enzyme of claim 47, wherein the DNA ligase polypeptide is a viral DNA ligase, a viral DNA ligase variant, or a functional fragment thereof.

51. The capping enzyme of claim 50, wherein the DNA ligase polypeptide is T4 DNA ligase, a variant thereof, or a functional fragment thereof.

52. The capping enzyme of claim 46, wherein the polynucleotide ligase polypeptide comprises an RNA ligase polypeptide. The capping enzyme of claim 52 , wherein the RNA ligase polypeptide is a prokaryotic RNA ligase, a prokaryotic RNA ligase variant, or a functional fragment thereof.

54. The capping enzyme of claim 52, wherein the RNA ligase polypeptide is a bacterial RNA ligase, a bacterial RNA ligase variant, or a functional fragment thereof.

55. The capping enzyme of claim 52, wherein the RNA ligase polypeptide is a viral RNA ligase, a viral RNA ligase variant, or a functional fragment thereof.

56. The capping enzyme of claim 55, wherein the RNA ligase polypeptide is T4 RNA ligase, a variant thereof, or a functional fragment thereof.

57. The capping enzyme of claim 56, wherein the RNA ligase polypeptide is T4 RNA ligase.

58. The capping enzyme of claim 57, wherein the RNA ligase polypeptide is T4 RNA ligase 1.

59. The capping enzyme of any one of claims 46-58, wherein the capping enzyme is a fusion polypeptide comprising a polynucleotide binding polypeptide fused to a polynucleotide ligase polypeptide.

60. The capping enzyme of any one of claims 46-59, wherein the polynucleotide binding polypeptide comprises a DNA binding polypeptide.

61. The capping enzyme of any one of claims 46-59, wherein the polynucleotide binding polypeptide comprises an RNA binding polypeptide.

62. The capping enzyme according to any one of claims 46 to 60, wherein the polynucleotide binding polypeptide is selected from one or more of a DNA double-strand binding domain, a DNA single-strand binding domain, an RNA / DNA composite strand binding domain, an RNA single-strand or RNA double-strand binding protein, wherein The DNA double-strand binding domain includes Sso7d and / or NF-kappaB p50; The RNA / DNA complex chain binding domain includes one or more of ScFV, RNaseH1 (D210N), and HBD of the monoclonal antibody S9.6; The RNA single-strand binding domain includes: Sso7d, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, RDE-4.

63. The capping enzyme of any one of claims 46-62, wherein the polynucleotide binding polypeptide comprises one or more of Sso7d, NF-kappaB p50, ScFV of monoclonal antibody S9.6, RNaseH1 (D210N), HBD, PKR, TRBP, PACT, Staufen, NFAR1, NFAR2, SPNR, RHA, NREBP, Kanadaptin, HYL1 Hyponastic leaves, ADAR1, ADAR2, ADAR3, TENR, RNaseIII, Dicer, and RDE-4.

64. The capping enzyme of claim 63, wherein the polynucleotide binding polypeptide is Sso7d.

65. The capping enzyme of any one of claims 46-64, wherein the C-terminus of the polynucleotide ligase polypeptide is linked to the N-terminus of the polynucleotide binding polypeptide.

66. The capping enzyme of any one of claims 46-64, wherein the N-terminus of the polynucleotide ligase polypeptide is linked to the C-terminus of the polynucleotide binding polypeptide.

67. The capping enzyme of any one of claims 46-66, wherein the polynucleotide ligase polypeptide and the polynucleotide binding polypeptide are linked by a first linker.

68. The capping enzyme of claim 67, wherein the first linker is a G4S linker.

69. The capping enzyme of any one of claims 46-68, wherein the polynucleotide ligase polypeptide comprises the amino acid sequence shown in SEQ ID NOs: 9-12, 21-22.

70. The capping enzyme of any one of claims 46-69, wherein the polynucleotide binding polypeptide comprises the amino acid sequence shown in SEQ ID NOs: 13-17.

71. The capping enzyme of any one of claims 46-70, wherein the capping enzyme comprises the amino acid sequence shown in SEQ ID NOs: 18-20, 24-29.

72. A kit comprising the capping enzyme according to any one of claims 46 to 71.

73. Use of the capping enzyme according to any one of claims 46 to 71 in capping polynucleotides.

Citation Information

Patent Citations

  • Fusion polypeptides and uses thereof

    CN102597006A

  • Taq DNA ligase fusion protein

    CN108129571A

  • Recombinant T4 DNA ligase mutant, fusion protein and application thereof

    CN115896047A

  • Method for connecting nucleic acid fragment and linker

    CN116218953A

  • RNA ligase enzymes and methods of preparation and use thereof

    WO2023085955A1