Nucleic acid ligase and application thereof

CN120476200APending Publication Date: 2025-08-12BGI HANGZHOU CYCLONESEQ TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280102693.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing nucleic acid ligases are less efficient in catalyzing the ligation of nucleic acid fragments, especially for nucleic acid substrates with terminal modifications, which affects the efficiency of molecular cloning and sequencing.

Method used

Through protein structure analysis and mutant design of T4 DNA ligase, a mutant nucleic acid ligase with high ligation efficiency was constructed. The mutation points include steric hindrance regulation and charge regulation of amino acid residues, which improves the ligation efficiency between nucleic acid fragments. Includes ligation of unmodified and modified fragments.

Benefits of technology

The connection efficiency between nucleic acid fragments is significantly improved, especially in nanopore sequencing, the effective library construction and sequencing time is improved, and the connection ability of modified nucleic acid fragments is enhanced.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A nucleic acid ligase is provided comprising a mutant sequence having at least 80% identity as compared to SEQ ID NO: 1, the mutation comprising at least one of substitution, deletion and insertion.
Need to check novelty before this filing date? Find Prior Art

Description

Nucleic acid ligase and its application Technical Field

[0001] The present application relates to the field of molecular biology technology, and in particular to a nucleic acid ligase and its application. Background Art

[0002] Nucleic acid ligases are important enzymes that catalyze the formation of phosphodiester bonds between two nucleic acid fragments (such as DNA or RNA fragments), thereby joining the two fragments. Nucleic acid ligases come from a variety of sources, with T4 ligase, derived from bacteriophage T4, being widely used in applications such as molecular cloning.

[0003] Depending on the conformation of the nucleic acid substrate catalyzed by the ligase, ligation reactions include nick ligation, cohesive end ligation, TA ligation, blunt end ligation, and branch ligation. TA ligation and blunt end ligation are widely used in molecular cloning and library sequencing, but their efficiency is relatively low, especially for nucleic acid substrates with terminal modifications.

[0004] Therefore, it is urgent to provide a T4 nucleic acid ligase with improved ligation efficiency.

[0005] Summary of the Invention

[0006] In order to solve the problem of low ligation efficiency of existing nucleic acid ligases, the embodiments of the present application provide a nucleic acid ligase with high ligation efficiency, which is particularly suitable for catalyzing nucleic acid substrates with terminal modifications.

[0007] In a first aspect, the present application provides a nucleic acid ligase or a biologically active fragment thereof, wherein the nucleic acid ligase or the biologically active fragment thereof comprises a mutant sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identity to SEQ ID NO: 1, wherein the mutation comprises at least one of a substitution, a deletion and an insertion.

[0008] In some embodiments, the nucleic acid ligase or a biologically active fragment thereof comprises a mutant sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 amino acid mutations compared to SEQ ID NO: 1.

[0009] In some embodiments, the mutant sequence has at least one amino acid mutation compared to SEQ ID NO: 1, and the at least one amino acid mutation includes a mutation at position 159 and / or position 164 of SEQ ID NO: 1. In some embodiments, the mutation is a substitution.

[0010] In some embodiments, the mutations at position 159 and / or position 164 of SEQ ID NO: 1 include: the amino acid at position 159 of SEQ ID NO: 1 is substituted with G, H, N, A, V, Q, C, S or T; the amino acid at position 164 of SEQ ID NO: 1 is substituted with G, H, N, A, V, Q, C, S or T.

[0011] In some embodiments, the amino acid at position 159 of SEQ ID NO: 1 is substituted with G or H and / or the amino acid at position 164 of SEQ ID NO: 1 is substituted with G or H.

[0012] In some embodiments, the amino acid at position 159 of SEQ ID NO: 1 is substituted with G and / or the amino acid at position 164 of SEQ ID NO: 1 is substituted with H.

[0013] In some embodiments, the mutant sequence is the sequence shown in SEQ ID NO: 3 or SEQ ID NO: 7.

[0014] In some embodiments, the nucleic acid ligase is T4 DNA ligase.

[0015] The second aspect of the present application provides a polynucleotide encoding a nucleic acid ligase or a biologically active fragment thereof or a complementary sequence thereof as provided in any embodiment of the first aspect of the present application.

[0016] The third aspect of the present application provides a vector comprising the polynucleotide provided in the embodiment of the second aspect of the present application.

[0017] The fourth aspect of the present application provides a cell, which comprises a nucleic acid ligase or a biologically active fragment thereof as provided in any embodiment of the first aspect of the present application or a polynucleotide as provided in any embodiment of the second aspect of the present application.

[0018] The fifth aspect of the present application proposes a kit, which comprises at least one of the following: a nucleic acid ligase or a biologically active fragment thereof as proposed in any embodiment of the first aspect of the present application, a polynucleotide as proposed in an embodiment of the second aspect of the present application, a vector as proposed in an embodiment of the third aspect of the present application, and a cell as proposed in an embodiment of the fourth aspect of the present application.

[0019] The sixth aspect of the present application provides a kit comprising a nucleic acid ligase or a biologically active fragment thereof and a reaction buffer as provided in any embodiment of the first aspect of the present application.

[0020] The seventh aspect of the present application proposes the use of a nucleic acid ligase or a biologically active fragment thereof as proposed in any embodiment of the first aspect of the present application, a polynucleotide as proposed in an embodiment of the second aspect of the present application, a vector as proposed in an embodiment of the third aspect of the present application, a cell as proposed in an embodiment of the fourth aspect of the present application, or a kit as proposed in an embodiment of the fifth or sixth aspect of the present application in polynucleotide ligation.

[0021] In the eighth aspect of the present application, a method for connecting polynucleotides is proposed, comprising: mixing a nucleic acid ligase or a biologically active fragment thereof as proposed in any embodiment of the first aspect of the present application with a first nucleotide fragment and a second nucleotide fragment to obtain a reaction mixture; incubating the reaction mixture so that the 3' end of the first nucleotide fragment and the 5' end of the second nucleotide fragment are connected through a phosphodiester bond under the action of the nucleic acid ligase or the biologically active fragment thereof.

[0022] In some embodiments, the 3' end of the first nucleotide fragment and the 5' end of the second nucleotide fragment independently comprise blunt ends or sticky ends. In some embodiments, the sticky ends are sticky T ends or sticky A ends.

[0023] In some embodiments, the first nucleotide segment and / or the second nucleotide segment comprises a modified mononucleotide or polynucleotide. In some embodiments, the first nucleotide segment and / or the second nucleotide segment comprises a modified mononucleotide. In some embodiments, the modification is a substitution modification; in some embodiments, the substitution modification is an alkyl substitution modification. In some embodiments, the alkyl substitution modification occurs on a phosphate group of the mononucleotide or polynucleotide, and the alkyl group is preferably a methyl group.

[0024] In some embodiments, the modified single or multinucleotides are located at the blunt ends or sticky ends of the first nucleotide segment and the second nucleotide segment.

[0025] In a ninth aspect, the present application provides a method for preparing a sequencing library, comprising: mixing a nucleic acid ligase or a biologically active fragment thereof as provided in any embodiment of the first aspect of the present application with a test fragment and a linker fragment to obtain a reaction mixture; and incubating the reaction mixture so that the linker fragment and the test fragment are linked via a phosphodiester bond under the action of the nucleic acid ligase or the biologically active fragment thereof. In some embodiments, the test fragment comprises a modified mononucleotide or polynucleotide.

[0026] In some embodiments, the sequencing library is used for NGS sequencing or nanopore sequencing.

[0027] In some embodiments, the linker fragment comprises a spacer sequence, and the spacer sequence is selected from at least one of iSp18, iSpC3, or a modified single nucleotide or polynucleotide.

[0028] In some embodiments, the modification is a substitution modification. In some embodiments, the substitution modification is an alkyl substitution modification. In some embodiments, the alkyl substitution modification occurs on the phosphate group of the mononucleotide or polynucleotide, and the alkyl group is preferably a methyl group.

[0029] The technical solution of this application achieves the following technical effects:

[0030] Compared to wild-type T4 DNA ligase, the nucleic acid ligase of the present application demonstrates higher catalytic ligation efficiency in ligating conventional nucleic acid fragments, ligating nucleic acid fragments with terminal modifications, and ligating conventional nucleic acid fragments with nucleic acid fragments with terminal modifications. Furthermore, in sequencing library construction, the adapter ligation catalyzed by the nucleic acid ligase of the present application is highly efficient, enabling the generation of more effective libraries. Furthermore, libraries constructed using the nucleic acid ligase of the present application exhibit higher effective sequencing times in nanopore sequencing. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0032] FIG1 is an electrophoresis detection diagram of a purified protein according to an embodiment of the present application;

[0033] FIG2 is an electrophoresis detection diagram of the purified protein according to an embodiment of the present application;

[0034] FIG3 is an electrophoresis detection diagram of the ligation product of the test fragment and the adapter Adp according to an embodiment of the present application;

[0035] FIG4 is a diagram of a typical sequencing current signal according to an embodiment of the present application;

[0036] FIG5 shows the effective sequencing time obtained by nanopore sequencing of sequencing libraries constructed using wild-type and mutant T4 DNA ligase according to an embodiment of the present application;

[0037] FIG6 shows an α-alkyl substituted deoxyribonucleotide according to an embodiment of the present application;

[0038] FIG7 is a structural diagram of iSpC3 according to an embodiment of the present application;

[0039] FIG8 is a structural diagram of iSp18 according to an embodiment of the present application. DETAILED DESCRIPTION

[0040] The present invention will be further described in detail below in conjunction with specific embodiments. The examples provided are only for illustrating the present invention and are not intended to limit the scope of the present invention. The examples provided below can serve as a guide for further improvements by those skilled in the art and are not intended to limit the present invention in any way.

[0041] This application is made based on the following knowledge of the inventors:

[0042] In related technologies, T4 nucleic acid ligase, derived from T4 bacteriophage, possesses blunt-end or TA-terminal ligation activity and is widely used in molecular cloning, nucleic acid modification, and library sequencing. In library sequencing, ligation of blunt-end or TA-terminal adapters is a key step in library construction. Its ligation efficiency directly affects the ligation reaction time, the coverage, and efficiency of subsequent sequencing results. This is especially true for sequencing of longer fragments, where the efficiency of T4 nucleic acid ligase-based adapter ligation is a key factor limiting sequencing efficiency.

[0043] Nanopore sequencing is a third-generation sequencing technology that has emerged in recent years. It has the advantages of long read length, high throughput, low cost and portability. Nanopore sequencing is a sequencing technology based on electrical signal recognition, in which different molecules entering the nanopore will hinder the flow of ions, which is called the nanopore signal. When the fragment to be tested, such as ssDNA, passes through the nanopore, the difference in nucleotides will cause different degrees of current obstruction. By detecting the current fluctuation signal of the nanopore and building a model to analyze the current signal through computer deep learning, the sequence of the perforated fragment can be determined.

[0044] The library construction process of nanopore sequencing technology includes end repair and terminal (3' end) A tail (dATP tail) steps similar to the second-generation sequencing library construction process, and then adapter connection is performed through T4 ligase-mediated TA ligation. During this process, the ligation catalytic activity of T4 ligase has an important influence on the speed and efficiency of adapter connection.

[0045] In addition, in nanopore sequencing, when no electric potential is applied, it is usually necessary to stop the polynucleotide binding protein on the target polynucleotide to prevent the polynucleotide binding protein from moving further along the target polynucleotide. Therefore, the linker sequence in nanopore sequencing not only contains conventional nucleic acid molecules such as tag sequences and tether sequences, but also contains spacer sequences, which can be used to stop the polynucleotide binding protein to block further unwinding of the polynucleotide binding protein before sequencing.

[0046] Typically, a spacer sequence containing a base-free group such as iSp18 or iSpC3 is used to halt the polynucleotide binding protein. In addition to iSp18 or iSpC3, other nucleotides with base modifications can also act as spacer sequences, such as deoxyribonucleotides with alkyl substitutions on α-phosphates (as shown in Figure 6). T4 ligase can be used to connect this sequence containing modified nucleotides (such as a linker sequence) to a nucleic acid fragment (such as a fragment to be tested) and used to prepare a sequencing library. In this step, T4 nucleic acid ligase also has an important influence on the connection speed and efficiency.

[0047] The inventors of this application conducted extensive protein structure analysis and rational mutant design of T4 DNA ligase, targeted steric and charge regulation of the amino acid residues in the reactive center, and constructed and prepared a large number of T4 DNA ligase mutants using protein engineering methods. These mutants were then subjected to extensive performance testing for further screening. Through extensive screening and experimentation, the inventors were pleasantly surprised to find that, compared to wild-type T4 DNA ligase, the T4 DNA ligase mutants proposed in the examples of this application had improved ligation efficiency for ligation between nucleic acid fragments, including ligation between conventional unmodified nucleic acid fragments, ligation between conventional nucleic acid fragments and modified nucleic acid fragments, and ligation between modified nucleic acid fragments. Furthermore, they also significantly improved the efficiency of nanopore sequencing.

[0048] In the examples of the present application, the term "nucleic acid ligase" refers to an enzyme that catalyzes the ligation of DNA or RNA fragments via phosphodiester bonds. In some embodiments, the "nucleic acid ligase" is DNA ligase. In other embodiments, the "nucleic acid ligase" is RNA ligase, specifically T4 DNA ligase or T4 RNA ligase. In the present application, "T4 ligase" can be T4 DNA ligase or T4 RNA ligase.

[0049] In the examples herein, the term "biologically active fragment" refers to any fragment, derivative, homolog, or analog of a T4 DNA ligase mutant that possesses in vivo or in vitro activity specific to a biomolecule, including, for example, ligase activity or repair of mismatches in nicked DNA via ligation. In some embodiments, the biologically active fragment, derivative, homolog, or analog of a mutant T4 DNA ligase possesses any degree of biological activity of the mutant T4 DNA ligase in any in vivo or in vitro assay of interest.

[0050] In some embodiments, the biologically active fragment may optionally include any number of consecutive amino acid residues of a mutant T4 DNA ligase. The present invention also includes polynucleotides encoding any such biologically active fragment.

[0051] In the examples of the present application, "naturally occurring" or "wild type" refers to the form found in nature. For example, a naturally occurring or wild type polypeptide or polynucleotide sequence is a sequence present in an organism that has not been intentionally modified by human manipulation. "Mutant" means a sequence that has at least one amino acid change relative to a natural or wild type amino acid sequence. In some embodiments, the change (mutation) includes at least one of a substitution, a deletion, and an insertion.

[0052] In the present application embodiment, the term " identity percentage " about nucleic acid or peptide sequence is defined as after arranging sequence to obtain maximum identity percentage and introducing breach (if necessary) to realize maximum homology percentage, the percentage of nucleotide or amino acid residue identical with known polypeptide in candidate sequence.N-terminal or C-terminal insertion or deletion should not be interpreted as affecting homology.Homology or identity on nucleotide or amino acid sequence level can be determined by BLAST (basic local alignment search tool, Basic Local Alignment Search Tool) analysis, described analysis uses the algorithm (Altschul (1997) adopted by program blastp, blastn, blastx, tblastn and tblastx, Nucleic Acids Res [nucleic acid research] 25,3389-3402 and Karlin (1990), Proc.Natl.Acad.Sci.USA [U.S. National Academy of Sciences] 87,2264-2268), described program is customized for sequence similarity search.

[0053] The nucleic acid ligase or biologically active fragment thereof provided in the embodiments of the present application can comprise a sequence having at least 80%, at least 85%, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.91%, 99.92%, 99.93%, 99.94%, 99.95%, 99.96%, 99.97%, 99.98%, 99.99% but less than 100% identity to SEQ ID NO: 1, wherein the sequence has one or more amino acid mutations compared to SEQ ID NO: 1.

[0054] Specifically, the nucleic acid ligase or its biologically active fragment proposed in the first aspect of the present application comprises a mutant sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identity with SEQ ID NO: 1, wherein the mutation comprises at least one of substitution, deletion and insertion.

[0055] In some embodiments, the nucleic acid ligase or biologically active fragment thereof comprises a mutant sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 amino acid mutations compared to SEQ ID NO: 1. In the examples of the present application, unless otherwise specified, except for the specified mutations, the remaining amino acids of the mutant fragment of the nucleic acid ligase or biologically active fragment thereof are identical to those of SEQ ID NO: 1.

[0056] In some embodiments, the mutant sequence has at least one amino acid mutation compared to SEQ ID NO: 1, and the at least one amino acid mutation includes a mutation at position 159 of SEQ ID NO: 1. In other embodiments, the at least one amino acid mutation includes a mutation at position 164 of SEQ ID NO: 1. In other embodiments, the at least one amino acid mutation includes a mutation at positions 159 and 164 of SEQ ID NO: 1. In some embodiments, the mutation at positions 159 and / or 164 of the mutant sequence is a substitution.

[0057] In some embodiments, the mutant form of the mutant sequence at position 159 and / or position 164 of SEQ ID NO: 1 can be: the amino acid K at position 159 of SEQ ID NO: 1 is substituted with G, H, N, A, V, Q, C, S or T, and / or the amino acid R at position 164 of SEQ ID NO: 1 is substituted with G, H, N, A, V, Q, C, S or T. It will be appreciated that the mutant sequence may have any of the above-mentioned single-point substitution forms or a combination thereof, such as K159G, K159H, K159N, K159A, K159V, K159Q, K159C, K159S, K159T, R164G, R164H, R164N, K164A, K164V, K164Q, K164C, K164S, K164T or any combination thereof, such as K159G-R164G, K159G-R164H, K159G-R164N, K159H-R164G, K159H-R164H, K159H-R164N, K159N-R164G, K159N-R164H or K159N-R164N, and the like.

[0058] In some embodiments, the mutant sequence is any one of SEQ ID NOs: 3-9. In some embodiments, the amino acid K at position 159 of SEQ ID NO: 1 is substituted with G, and / or the amino acid R at position 164 of SEQ ID NO: 1 is substituted with H, i.e., the mutant forms are K159G, R164H, or K159G-R164H, and the corresponding amino acid sequences are SEQ ID NO: 3, SEQ ID NO: 7, and SEQ ID NO: 9, respectively. In some embodiments, the mutant sequence is the sequence set forth in SEQ ID NO: 3 or SEQ ID NO: 7.

[0059] The present application embodiment also relates to a polynucleotide encoding a mutant nucleic acid ligase or its biologically active fragment or its complementary sequence in any of the above embodiments, a vector comprising the polynucleotide, or a cell expressing the mutant nucleic acid ligase or its biologically active fragment in any of the above embodiments, and a kit, the kit comprising any of the mutant nucleic acid ligase or its biologically active fragment in any of the above embodiments, a polynucleotide encoding a mutant nucleic acid ligase or its biologically active fragment or its complementary sequence in any of the above embodiments, a vector comprising the polynucleotide, or a cell expressing the mutant nucleic acid ligase or its biologically active fragment in any of the above embodiments, or a combination thereof. The present application embodiment also proposes a kit, the kit comprising the mutant nucleic acid ligase or its biologically active fragment and a reaction buffer in any of the above embodiments. It is understandable that, as long as the reaction buffer is capable of providing a reaction background for the mutant ligase or its biologically active fragment, the present application does not limit the type of reaction buffer.

[0060] The mutant nucleic acid ligases or biologically active fragments thereof proposed in the examples of this application have improved catalytic ligation efficiency of nucleic acid fragment substrates compared to the wild-type (SEQ ID NO: 1). When the substrate contains terminally modified nucleotides, the mutant nucleic acid ligases or biologically active fragments proposed in the examples of this application also exhibit highly efficient catalytic ligation ability compared to the wild-type (SEQ ID NO: 1).

[0061] In a second aspect, the present application proposes a use of a nucleic acid ligase or a biologically active fragment thereof, a polynucleotide, a vector, a cell or a kit proposed in any of the above embodiments in polynucleotide ligation.

[0062] In a third aspect, the present application provides a method for ligating polynucleotides, comprising: mixing the nucleic acid ligase or its biologically active fragment provided in any of the above embodiments with a first nucleotide fragment and a second nucleotide fragment to obtain a reaction mixture; and incubating the reaction mixture so that the 3' end of the first nucleotide fragment and the 5' end of the second nucleotide fragment are linked via a phosphodiester bond under the action of the nucleic acid ligase or its biologically active fragment. It is understood that the first nucleotide fragment and the second nucleotide fragment form a phosphodiester bond under the catalysis of the nucleic acid ligase or its biologically active fragment provided in any of the above embodiments, thereby achieving the ligation of the two. In the embodiments of the present application, the first nucleotide fragment can be a DNA fragment or an RNA fragment, and the second nucleotide fragment can be a DNA fragment or an RNA fragment.

[0063] In some embodiments, the 3' end of the first nucleotide fragment and the 5' end of the second nucleotide fragment independently comprise a blunt end or a sticky end. It is understood that the T4 DNA ligase mutant proposed in the examples of the present application can mediate the connection between the blunt end of the first nucleotide fragment and the blunt end of the second nucleotide fragment, and can also mediate and catalyze the connection between the sticky complementary ends of the first nucleotide fragment and the second nucleotide fragment. In some embodiments, the sticky complementary ends are a sticky T end and a sticky A end. The mutant nucleic acid ligase or its biologically active fragment proposed in the examples of the present application shows a high catalytic connection efficiency in both the blunt end connection and the sticky end connection of the nucleotide fragments.

[0064] In some embodiments, the first nucleotide fragment and / or the second nucleotide fragment may comprise a mononucleotide or polynucleotide with a modification. "Modification" refers to the presence of a modifying group in its phosphate group, pentose moiety or base moiety based on the basic structure of the nucleotide. In some embodiments, the modification may be a substitution modification, i.e., the presence of a substitution of an atom or a group of atoms in the nucleotide, such as a carbon atom, an oxygen atom, a nitrogen atom, a sulfur atom, an alkyl group, a hydroxyl group, an amino group, a carboxyl group, a phosphate group, etc. replacing an atom or group in the original nucleotide. In some embodiments, the modification may be an alkyl substitution modification. In some embodiments, the alkyl substitution modification may occur on the phosphate group of a mononucleotide or polynucleotide. In some embodiments, the alkyl group may be a methyl group, an ethyl group, a propyl group, a butyl group, etc., and its modified product is an α-alkyl substituted-(deoxy)ribonucleotide. Figure 6 shows a specific modified structure of a deoxyribonucleotide in an embodiment of the present application, wherein R represents an alkyl group and base (base) may be A, T, C or G.

[0065] In some embodiments, the modified mononucleotide or polynucleotide can be located in the middle portion of the first nucleotide segment and the second nucleotide segment, or at a blunt end or a sticky end. In some embodiments, the first nucleotide segment and / or the second nucleotide segment comprises a modified mononucleotide. The mutant nucleic acid ligase or its biologically active fragment proposed in the embodiments of the present application exhibits high catalytic ligation efficiency in both blunt-end ligation and sticky-end ligation of modified nucleotide segments.

[0066] Specifically, the mutant nucleic acid ligase or its biologically active fragment proposed in any embodiment of the first aspect of the present application can be used for the connection between nucleic acid fragments, for example, in molecular cloning, it can be used for the connection between a vector and an insert sequence; in modification introduction, it can be used for the connection between a modified nucleotide fragment and other fragments, where the other fragments can be a modified nucleotide fragment or a conventional nucleotide fragment without modification; in sequencing library construction, it can be used for the connection between the fragment to be tested and the linker.

[0067] In a fourth aspect, the present application proposes a method for preparing a sequencing library, comprising: mixing the nucleic acid ligase or its biologically active fragment proposed in any of the above embodiments with a fragment to be tested and a linker fragment to obtain a reaction mixture; incubating the reaction mixture so that the linker fragment and the fragment to be tested are linked through a phosphodiester bond under the action of the nucleic acid ligase or its biologically active fragment.

[0068] In some embodiments, the fragment to be detected comprises a single nucleotide or a polynucleotide with a modification.

[0069] In certain embodiments, the sequencing library prepared by the sequencing library preparation method proposed in any embodiment of the fourth aspect of the application can be used for next generation sequencing (NGS) or nanopore sequencing.In nanopore sequencing, the mutant nucleic acid ligase proposed in the embodiment of the application or its biologically active fragment can be used to connect the sequence to be tested with the joint fragment with modified fragment, wherein the modified fragment in the joint fragment can be a mononucleotide or a polynucleotide, and as a spacer sequence.Compared to wild type, the sequencing library prepared by the mutant nucleic acid ligase proposed in the embodiment of the application or its biologically active fragment has higher effective library content, and significantly improves effective sequencing time in nanopore sequencing.

[0070] Unless otherwise specified, the experimental methods in the following examples are conventional methods and were performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials and reagents used in the following examples, unless otherwise specified, were all commercially available.

[0071] Unless otherwise specified, the quantitative tests in the following examples were performed three times, and the results were averaged.

[0072] Example 1: Preparation of wild-type T4 DNA ligase (T4L-WT)

[0073] 1.1 Cloning and expression of wild-type T4 DNA ligase (T4L-WT)

[0074] The amino acid sequence of wild-type T4 DNA ligase (T4L-WT) is shown in SEQ ID NO: 1, and the coding region sequence is shown in SEQ ID NO: 2. The DNA sequence shown in SEQ ID NO: 2 was synthesized (synthesis was commissioned by Liuhe BGI, and all subsequent DNA sequence syntheses were performed in the same manner and will not be repeated here). Based on the dual restriction sites NdeI and XhoI, this sequence was inserted into the PET-28a(+) plasmid using conventional methods. The corresponding T4L-WT protein expressed has a 6×His-tag tag and a thrombin (thrombin) restriction site at the N-terminus.

[0075] T4L-WT amino acid sequence:

[0076]

[0077] T4L-WT coding region sequence:

[0078]

[0079] The connected PET-28a(+)-T4L-WT plasmid was transformed into the Escherichia coli expression bacteria BL21 (DE3). The transformation product was spread on solid LB medium containing kanamycin resistance and cultured at 37°C overnight. The next day, a single colony was picked and inoculated into 5 mL of LB liquid medium containing kanamycin resistance and cultured at 37°C overnight. Then, 10 μL of the overnight cultured bacterial solution was transferred into 1 L of LB liquid medium and cultured at 37°C with shaking until the OD 600 =0.6 to 0.8, and then IPTG was added to a final concentration of 500 μM, and cultured overnight at 16°C to induce protein expression.

[0080] 1.2 Purification

[0081] Prepare the following solution:

[0082] Buffer A: 20mM Tris-HCl pH 7.5, 250mM NaCl, 20mM imidazole

[0083] Buffer B: 20 ​​mM Tris-HCl pH 7.5, 250 mM NaCl, 300 mM imidazole

[0084] Buffer C: 20mM Tris-HCl pH 7.5, 50mM NaCl

[0085] Buffer D: 20mM Tris-HCl pH 7.5, 100mM NaCl

[0086] Collect the cells after IPTG induction in step 1.1, resuspend them in Buffer A, disrupt them using a cell disruptor, and centrifuge the supernatant. Mix the supernatant with Ni-NTA medium pre-equilibrated with Buffer A and allow to bind for 1 hour. Collect the medium and wash it with Buffer A until all contaminants are removed. Then, add Buffer B to the medium to elute the T4L-WT. Pass the eluted T4L-WT protein through a desalting column pre-equilibrated with Buffer C. Desalt the protein using Buffer C, then add thrombin (Yisheng Bio, 20402ES05) and digest it overnight at 4°C. The next day, purify the digested protein using Superdex 200 (Sigma, GE28-9909-44) using Buffer D. Collect the target protein peak, concentrate it using a 10 kD ultrafiltration tube (Millipore MRCPRT010), and freeze it at -80°C until needed. The purified and concentrated protein was detected using SDS-PAGE (sodium dodecyl sulfate polyacrylamide gel electrophoresis).

[0087] The test results are shown in Figure 1, where lane 2 is the purified T4L-WT protein. As shown in Figure 1, a large amount of high-purity T4L-WT protein was prepared in this example after the purification steps.

[0088] Example 2: Preparation and purification of T4 DNA ligase mutants

[0089] The specific mutation sites and amino acid sequences of the T4 DNA ligase mutants in this example are shown in Table 1.

[0090] Table 1

[0091]

[0092] Note: Taking T4L-Mut1 (i.e., mutant 1) as an example, based on the sequence shown in SEQ ID NO: 1, K at position 159 of SEQ ID NO: 1 is replaced by G, and the remaining positions are identical to SEQ ID NO: 1; the same applies to T4L-Mut2 to T4L-Mut6. T4L-Mut7, based on the sequence shown in SEQ ID NO: 1, K at position 159 of SEQ ID NO: 1 is replaced by G, and R at position 164 is replaced by G, and the remaining positions are identical to SEQ ID NO: 1.

[0093] T4L-Mut1 amino acid sequence:

[0094]

[0095] T4L-Mut2 amino acid sequence:

[0096]

[0097] T4L-Mut3 amino acid sequence:

[0098]

[0099] T4L-Mut4 amino acid sequence:

[0100]

[0101]

[0102] T4L-Mut5 amino acid sequence:

[0103]

[0104] T4L-Mut6 amino acid sequence:

[0105]

[0106] T4L-Mut7 amino acid sequence:

[0107]

[0108] DNA encoding the corresponding amino acid sequences of the T4 DNA ligase mutants in Table 1 were synthesized, and the DNA sequences of these mutants were inserted into the PET-28a(+) plasmid according to step 1.1 of Example 1 and expressed respectively. The expressed proteins were then purified according to step 1.2 of Example 1.

[0109] The results are shown in Figure 2. Lanes 1-7 are electrophoretic images of the purified proteins T4L-Mut1 to T4L-Mut7, respectively. As shown in Figure 2, after the purification steps, this example produced a large amount of high-purity mutant proteins with purification efficiency comparable to that of the T4L-WT protein.

[0110] Example 3: Preparation of connector

[0111] The sequences shown in SEQ ID NO: 10 and SEQ ID NO: 11 were synthesized, and 30 μL of each of the sequences shown in SEQ ID NO: 10 and SEQ ID NO: 11 at a concentration of 20 μM were mixed. After mixing, the mixture was placed in a PCR instrument and incubated at 70°C for 10 minutes. The temperature was then lowered to 25°C at a rate of 0.1°C / s and incubated for another half hour to obtain a 10 μM linker solution. The linker was named Adp.

[0112] SEQ ID NO: 10:

[0113] 5'-XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXTTTTTTTTTTYYYYGGTTGTTTCTGTTGGTGCTGATATTGCT-3' (X=iSpC3, whose structure is shown in FIG7 ; Y=iSp18, whose structure is shown in FIG8 )

[0114] SEQ ID NO: 11:

[0115] 5'pho-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3'

[0116] Example 4: Preparation of fragments to be tested

[0117] 4.1 Large-scale extraction of plasmids

[0118] The empty pUC57 plasmid was transformed into DH5alpha (Vazyme Biotech, C502-02) competent bacteria and cultured in large quantities. The empty pUC57 plasmid in DH5alpha was extracted using a plasmid extraction kit (Tiangen, DP117). The concentration of the extracted plasmid was determined using the Qubit dsDNA BR kit (Thermofisher, Q32853).

[0119] 4.2 Enzyme Digestion

[0120] The plasmid obtained in step 4.1 was digested with HindIII (NEB, R0104L) and EcoRI (NEB, R0101L). The specific digestion reaction system is shown in Table 2.

[0121] Table 2

[0122]

[0123] 4.3 Purification

[0124] The enzyme digestion products were purified using Ampure XP magnetic beads (Beckman Coulter, A63882) as follows:

[0125] Add 50 μL of magnetic beads equilibrated at room temperature to the above enzyme digestion system, shake to mix, centrifuge briefly, and let stand at room temperature for 10 minutes. Place the reaction system on a magnetic stand for 10 minutes and remove the supernatant. Wash the magnetic beads twice with 200 μL of 75% ethanol solution. After washing, remove the supernatant, open the tube cap, and after the magnetic beads are dry, add 22 μL TE buffer (pH = 8) to resuspend the magnetic beads and let stand at room temperature for 10 minutes. Place the centrifuge tube on a magnetic stand, and after the liquid is clarified, transfer the supernatant to a new centrifuge tube to obtain a purified enzyme digestion product with a length of nearly 2Kb (i.e., the fragment to be tested), whose sequence is shown in SEQ ID NO: 12.

[0126] DNA sequence of the fragment to be tested:

[0127]

[0128]

[0129] Example 5: Adding an A tail at the end

[0130] 5.1 End repair, A-tailing and phosphorylation

[0131] The 2Kb fragment obtained in Example 4 was end-repaired, A-tailed, and 5'-end phosphorylated using ordinary unmodified dATP and α-methylphosphate-modified dATP (Me-dATP). The reaction system is shown in Table 3, and the reaction procedure is shown in Table 4.

[0132] Table 3

[0133]

[0134] Table 4

[0135] Temperature time 37℃30min72℃30min

[0136] After this step, the fragment to be tested has completed end repair, end A tailing and phosphorylation, and the following two groups of products are obtained: ① A dATP is added to the 3' end of the fragment to be tested and the 5' end is phosphorylated; ② A Me-dATP is added to the 3' end of the fragment to be tested and the 5' end is phosphorylated.

[0137] 5.2 Purification of the product

[0138] The two groups of products obtained in step 5.1 were purified using Ampure XP magnetic beads (Beckman Coulter, A63882) as follows:

[0139] Add 50 μL of magnetic beads equilibrated at room temperature to the end repair system in step 5.1, shake to mix, centrifuge briefly, and let stand at room temperature for 10 minutes. Place the reaction system on a magnetic stand for 10 minutes and remove the supernatant. Wash the magnetic beads twice with 200 μL of 75% ethanol solution. After washing, remove the supernatant, open the tube cap, and resuspend the magnetic beads after the magnetic beads are dry, add 21 μL of molecular grade water, and let stand at room temperature for 10 minutes. Place the centrifuge tube on a magnetic stand, wait for the liquid to clarify, and transfer the supernatant to a new centrifuge tube to obtain the purified product. Take 1 μL of the purified product and quantify it using the Qubit dsDNA HS kit (Thermofisher, Q32854). The remaining product is subjected to the next reaction.

[0140] Example 6: Ligation reaction between the fragment to be tested and the linker Adp

[0141] 6.1 Ligation of the fragment to be tested and the adapter Adp

[0142] Prepare ligation reaction solution A as shown in Table 5, and prepare ligation reaction system A as shown in Table 6 based on this reaction solution. Ligate the test fragment obtained in step 5.2 with the adapter Adp under the action of T4 DNA ligase wild type (T4L-WT) or mutant (T4L-Mut1-7). The ligation conditions are 25°C for 10 minutes.

[0143] Table 5

[0144]

[0145] Table 6

[0146]

[0147] 6.2 Purification of ligation products

[0148] The ligation product obtained in step 6.1 was purified using Ampure XP magnetic beads (Beckman Coulter, A63882) as follows:

[0149] Add 60 μL of magnetic beads equilibrated at room temperature to the connection system in step 6.1, shake to mix, centrifuge briefly, and let stand at room temperature for 10 minutes. Place the reaction system on a magnetic stand for 10 minutes and remove the supernatant. Wash the magnetic beads twice with 200 μL of 75% ethanol solution. After washing, remove the supernatant, open the tube cap, and after the magnetic beads are dry, add 22 μL TE buffer (pH = 8) to resuspend the magnetic beads and let stand at room temperature for 10 minutes. Place the centrifuge tube on a magnetic stand, wait for the liquid to clarify, and transfer the supernatant to a new centrifuge tube to obtain the purified connection product of the test fragment 1 and Adp.

[0150] 6.3 Detection of ligation products

[0151] The purified ligation products obtained in 6.2 were analyzed using native polyacrylamide gel electrophoresis (PAGE), and the results are shown in Figure 3. As shown in Figure 3, when conventional dATP was added to the termini of the test fragments, T4L-Mut1 (lane 3) and T4L-Mut5 (lane 7) exhibited higher ligation efficiency in the conventional A-tailed ligation reaction compared to T4L-WT (lane 2). Specifically, the catalytic effects of T4L-Mut1 and T4L-Mut5 yielded more ligation products and less residual substrate. Furthermore, when modified Me-dATP was added to the termini of the test fragments, the catalytic ligation efficiencies of the two mutants, T4L-Mut1 (lane 3) and T4L-Mut5 (lane 7), were also slightly higher than those of T4L-WT (lane 2). Similarly, the effects of T4L-Mut1 and T4L-Mut5 yielded more ligation products and less residual substrate. This indicates that the T4 DNA ligase mutant proposed in the examples of the present application has a higher efficiency in catalyzing the ligation reaction between the linker and the fragment to be tested than the wild type.

[0152] Example 7: Cloning, expression and purification of helicase Dda

[0153] In this example, helicase Dda was produced by recombinant expression in E. coli, and the helicase was used as a motor protein. The amino acid sequence of helicase Dda is shown in SEQ ID NO: 13, and the coding region sequence is shown in SEQ ID NO: 14.

[0154] Amino acid sequence of helicase Dda:

[0155]

[0156] Coding region sequence of helicase Dda:

[0157]

[0158]

[0159] The specific experimental steps are as follows.

[0160] 7.1 Cloning and Expression of Helicase Dda

[0161] The DNA sequence shown in SEQ ID NO: 14 was synthesized and inserted into the PET-28a(+) plasmid based on the double restriction sites NdeI and XhoI. The corresponding helicase Dda protein expressed had a 6×His-tag tag and a thrombin restriction site at the N-terminus.

[0162] The ligated PET-28a(+)-Dda plasmid was transformed into Arctic Express (DE3) competent bacteria (Tolo Biotech, 96183-02). The transformation product was spread on solid LB medium containing kanamycin resistance and cultured at 37°C overnight. The next day, a single colony was picked and inoculated into 5 mL of LB liquid medium containing kanamycin resistance and cultured at 37°C with shaking overnight. Then, 10 μL of the overnight cultured bacterial solution was transferred into 1 L of LB liquid medium and cultured at 37°C with shaking until the OD 600 =0.6 to 0.8, and then IPTG was added at a final concentration of 500 μM, and cultured overnight at 16°C to induce Dda protein expression.

[0163] 7.2 Purification

[0164] Prepare the following solution:

[0165] Buffer E: 20 mM Tris-HCl pH 7.5, 100 mM NaCl

[0166] Collect the cells after IPTG induction in step 7.1, resuspend the cells in Buffer A prepared in step 1.2 of Example 1, disrupt the cells with a cell disruptor, and collect the supernatant after centrifugation. Mix the resulting supernatant with Ni-NTA filler pre-equilibrated with Buffer A and bind for 1 hour. Collect the filler and wash the filler with Buffer A until no impurities are washed out. Then add Buffer B to the filler to elute Dda. The eluted Dda protein is passed through a desalting column pre-equilibrated with Buffer C. Desalt with Buffer C, then add thrombin (Yisheng Bio, 20402ES05) and add it to ssDNA cellulose (Sigma, D8273-10G) filler pre-equilibrated with Buffer C. Digest overnight at 4°C. The next day, collect the ssDNA cellulose filler, wash 3-4 times with Buffer C, and elute the purified protein with Buffer D. The purified Dda protein solution was applied to a molecular sieve Superdex 200 (Sigma, GE28-9909-44), wherein the molecular sieve buffer used was Buffer E. The target protein peak was collected and concentrated to obtain the purified helicase Dda.

[0167] Example 8: Binding of Sequencing Library to Motor Protein

[0168] In this example, the purified ligation product (ie, sequencing library) obtained in Example 6 was incubated with the helicase Dda in Example 7 to prepare a conjugate for nanopore sequencing. The specific experimental steps are as follows.

[0169] Prepare 2× binding buffer: Mix 100 mL of 1 M Tris-HCl (pH 7.5) with 100 mL of 1 M KCl solution, then add ultrapure water to 1 L. Prepare the binding system of helicase Dda and ligation product on ice according to Table 7, then incubate at 30°C for one hour.

[0170] Table 7

[0171]

[0172]

[0173] The concentration of the bound product was determined using the Qubit DNA HS kit.

[0174] Example 9: Nanopore sequencing

[0175] In this example, a nanopore sequencing platform based on a patch clamp platform was constructed, and nanopore sequencing was performed on the sequencing library obtained in Example 8 and the motor protein binding product to verify the performance of the sequencing library constructed by the T4 DNA ligase mutant in the disclosed example in nanopore sequencing.

[0176] The single-channel electrophysiological detection system disclosed in the reference (Geng Jia, Guo Peixuan, "Application of bacteriophage phi29 DNA packaging motor phospholipid membrane chimera in single molecule detection and nanomedicine". Life Science, 2011, 23(11): 1114-1129) was used to build a nanopore detection platform based on a patch clamp platform, and porin (Sigma-Aldrich, H9395-5mg) was inserted into the phospholipid bilayer membrane to form a single-channel nanopore.

[0177] The binding product obtained in Example 8 was added to the single-channel nanopore system, and the current amplitude change was detected and recorded using a patch clamp system. A typical sequencing current signal of the test fragment (SEQ ID NO: 12) is shown in FIG4 .

[0178] The sequencing results were statistically analyzed, and the total duration of all sequencing signals accumulated in each hour was counted to determine the content of the effective library in the sequencing library sample. Figure 5 shows the effective sequencing time obtained for the sequencing libraries constructed based on the wild type and mutant of T4 DNA ligase in this embodiment during nanopore sequencing. As can be seen from Figure 5, compared with the wild type T4L-WT, the ligation reaction catalyzed by T4L-Mut1 showed a higher effective sequencing time in nanopore sequencing, indicating that the ligation reaction catalyzed by the T4 DNA ligase mutant of the embodiment of the present application obtained more effective libraries, proving that the T4 DNA ligase mutant proposed in the embodiment of the present application has a higher catalytic ligation efficiency than the wild type, and that the mutant can effectively increase the content of the effective library in nanopore sequencing, significantly improving the effective sequencing time.

[0179] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0180] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

[0181]

[0182]

[0183]

[0184]

[0185]

[0186]

[0187]

Claims

1. A nucleic acid ligase or a biologically active fragment thereof, comprising a mutant sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identity to SEQ ID NO: 1, wherein the mutation comprises at least one of a substitution, a deletion and an insertion.

2. The nucleic acid ligase or a biologically active fragment thereof according to claim 1, comprising a mutant sequence having at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9 or at least 10 amino acid mutations compared to SEQ ID NO:

1.

3. The nucleic acid ligase or biologically active fragment thereof according to claim 1 or 2, wherein the mutant sequence has at least one amino acid mutation compared to SEQ ID NO: 1, and the at least one amino acid mutation comprises a mutation at position 159 and / or position 164 of SEQ ID NO: 1, Preferably, the mutation is a substitution.

4. The nucleic acid ligase or biologically active fragment thereof according to claim 3, wherein the mutation at position 159 and / or position 164 of SEQ ID NO: 1 include: The amino acid at position 159 of SEQ ID NO: 1 is substituted with G, H, N, A, V, Q, C, S or T, The amino acid at position 164 of SEQ ID NO: 1 is substituted with G, H, N, A, V, Q, C, S or T.

5. The nucleic acid ligase or biologically active fragment thereof according to any one of claims 1 to 4, wherein the amino acid at position 159 of SEQ ID NO: 1 is substituted with G or H, and / or the amino acid at position 164 of SEQ ID NO: 1 is substituted with G or H, Preferably, the amino acid at position 159 of SEQ ID NO: 1 is substituted with G, and / or the amino acid at position 164 of SEQ ID NO: 1 is substituted with H. 6 . The nucleic acid ligase or biologically active fragment thereof according to claim 5 , wherein the mutant sequence is the sequence shown in SEQ ID NO: 3 or SEQ ID NO:

7. 7 . The nucleic acid ligase or a biologically active fragment thereof according to claim 1 , wherein the nucleic acid ligase is T4 DNA ligase. 8 . A polynucleotide encoding the nucleic acid ligase according to claim 1 , or a biologically active fragment thereof, or a complementary sequence thereof.

9. A vector comprising the polynucleotide according to claim 8.

10. A cell comprising the polynucleotide of claim 8 or expressing the nucleic acid ligase or a biologically active fragment thereof of any one of claims 1 to 7. 11 . A kit comprising at least one of the following: the nucleic acid ligase or biologically active fragment thereof according to any one of claims 1 to 7 , the polynucleotide according to claim 8 , the vector according to claim 9 , and the cell according to claim 10 . 12 . A kit comprising the nucleic acid ligase or a biologically active fragment thereof according to claim 1 and a reaction buffer.

13. Use of the nucleic acid ligase or biologically active fragment thereof according to any one of claims 1 to 7, the polynucleotide according to claim 8, the vector according to claim 9, the cell according to claim 10 or the kit according to claim 11 or 12 in polynucleotide ligation.

14. A method for connecting polynucleotides, include: mixing the nucleic acid ligase or the biologically active fragment thereof according to any one of claims 1 to 7 with the first nucleotide fragment and the second nucleotide fragment to obtain a reaction mixture; Incubating the reaction mixture so that the 3' end of the first nucleotide fragment and the 5' end of the second nucleotide fragment are linked by a phosphodiester bond under the action of the nucleic acid ligase or a biologically active fragment thereof, Optionally, the 3' end of the first nucleotide fragment and the 5' end of the second nucleotide fragment independently comprise a blunt end or a sticky end, preferably, the sticky end is a sticky T end or a sticky A end.

15. The method according to claim 14, wherein the first nucleotide fragment and / or the second nucleotide fragment comprises a modified mononucleotide or polynucleotide, Preferably, the first nucleotide segment and / or the second nucleotide segment comprises a single nucleotide with a modification, Optionally, the modification is a substitution modification, optionally, the substitution modification is an alkyl substitution modification, Preferably, the alkyl substitution modification occurs on the phosphate group of the mononucleotide or polynucleotide, and the alkyl group is preferably a methyl group. 16 . The method according to claim 14 or 15 , wherein the modified single or polynucleotide is located at the blunt end or sticky end of the first nucleotide fragment and / or the second nucleotide fragment.

17. A method for preparing a sequencing library, include: mixing the nucleic acid ligase or the biologically active fragment thereof according to any one of claims 1 to 7 with the fragment to be tested and the linker fragment to obtain a reaction mixture; Incubating the reaction mixture so that the linker fragment and the fragment to be detected are linked through a phosphodiester bond under the action of the nucleic acid ligase or its biologically active fragment, Optionally, the fragment to be detected comprises a modified single nucleotide or polynucleotide.

18. The method according to claim 17, wherein the sequencing library is used for NGS sequencing or nanopore sequencing.

19. The method according to claim 18, wherein the linker fragment comprises a spacer sequence, and the spacer sequence is selected from at least one of iSp18, iSpC3, or a modified single nucleotide or polynucleotide.

20. The method according to claim 19, wherein the modification is a substitution modification, preferably, the substitution modification is an alkyl substitution modification, Preferably, the alkyl substitution modification occurs on the phosphate group of the mononucleotide or polynucleotide, and the alkyl group is preferably a methyl group.