T7 DNA ligase variants with increased ligation efficiency

By performing specific amino acid mutations on T7 DNA ligase, its ligation activity on blunt-terminal DNA substrates is enhanced, and the problem of insufficient activity of existing T7 DNA ligases is solved, achieving more efficient DNA amplification and sequencing applications.

CN120349979APending Publication Date: 2025-07-22WUHAN AIBO TAIKE BIOTECH CO LTD
View PDF 32 Cites 0 Cited by

Patent Information

Application Number
CN202510069610.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-20
Filing Date
2025-01-16
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing T7 DNA ligase has low activity when ligating blunt-terminal DNA substrates, making it difficult to meet the efficient ligation needs of molecular biology research.

Method used

By engineering T7 DNA ligase, specific amino acid mutations (E63K, D132R, E243K, D245R, E272K, D288R, E289K, E292K) were introduced to enhance its ligation activity on the blunt-terminal substrate.

Benefits of technology

It improves the amplification activity of T7 DNA ligase on DNA sequences at low concentrations, is suitable for sequencing methods and plasmid replication, and improves ligation efficiency and speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120349979A_ABST
    Figure CN120349979A_ABST
Patent Text Reader

Abstract

The invention comprises a mutant T7DNA ligase or a bioactive fragment thereof. The activity of the mutant T7DNA ligase to a blunt end dsDNA substrate is higher than that of a wild T7DNA ligase. The mutant T7DNA ligase or the bioactive fragment thereof has one or more substitutions different from the wild type, and the substitutions are E63K (SEQ ID NO: 3; sEQ ID NO: 4), D132R (SEQ ID NO: 5; sEQ ID NO: 6), E243K (SEQ ID NO: 7; sEQ ID NO: 8), D245R (SEQ ID NO: 9; sEQ ID NO: 10), E272K (SEQ ID NO: 11; sEQ ID NO: 12), D288R (SEQ ID NO: 13; sEQ ID NO: 14), E289K (SEQ ID NO: 15; sEQ ID NO: 16) and E292K (SEQ ID NO: 17; sEQ ID NO: 18).
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Ligases are commonly used in molecular biology to form phosphodiester bonds between double-stranded nucleic acid fragments at the junction of juxtaposed 5'-phosphate and 3'-hydroxyl termini. By designing complementary overhangs between each double-stranded fragment, ligation can be directed to be position-specific and orientationally oriented. This allows for the specific integration of DNA or RNA fragments into larger vectors to meet the needs of molecular biology research. T7 DNA ligase is a thermostable ATP-dependent enzyme that efficiently catalyzes the ligation of double-stranded DNA polynucleotides that exhibit sticky ends longer than a single overhanging base. Additionally, it can also repair mismatches present in nicked DNA via ligation.

[0002] Ligases are the backbone of many molecular biology protocols, enabling users to design and ligate various nucleotide molecules for uses such as cloning, protein expression, and sequencing. Depending on the user's goals, increasing the inherent ligation activity of T7 DNA ligase increases the usefulness of the enzyme, thereby allowing for ligation to require less enzyme or less total time to achieve complete ligation of the substrate.

[0003] Conventionally, ligases are a molecular cloning tool used to insert specific DNA fragments into vectors prior to transformation into competent cells in order to achieve the ligation of sequencing adapters as part of an NGS workflow, or as part of various other molecular biology protocols that require the ligation of two or more fragments of double-stranded DNA polynucleotides. When DNA polynucleotides do not contain unpaired bases at their 3' or 5' ends, they are considered to have blunt ends, in contrast to when overhanging bases extend from the fragment ends, in which case they are considered to have sticky ends, and these overhangs can provide sites for complementary DNA binding and ligation.

[0004] T7 DNA ligase is an ATP-dependent enzyme from a bacteriophage that can catalyze the phosphodiester bond between two complementary overhanging ends of double-stranded DNA fragments, thereby effectively joining two separate DNA fragments together. T7 DNA ligase can also repair mismatches present in nicked DNA by phosphorylating the 5'-phosphate to the 3'-hydroxyl. However, in addition, the ligation of double-stranded DNA substrates with blunt ends is possible in the presence of crowding agents such as polyethylene glycol (PEG). PEG is also used to increase the overall activity of T7 DNA ligase up to 100-fold the standard activity and is a common additive in ligase buffers. Summary of the Invention

[0005] The present invention relates to engineered T7 DNA ligase mutants which exhibit enhanced ligation activity on blunt-ended substrates as compared to wild-type ligase. The following T7 DNA ligase mutants have been identified as having such enhanced ligation activity (for each mutant, the odd-numbered sequence is the DNA sequence and the even-numbered sequence is the amino acid sequence): E63K (SEQ ID NO: 3; SEQ ID NO: 4), D132R (SEQ ID NO: 5; SEQ ID NO: 6), E243K (SEQ ID NO: 7; SEQ ID NO: 8), D245R (SEQ ID NO: 9; SEQ ID NO: 10), E272K (SEQ ID NO: 11; SEQ ID NO: 12), D288R (SEQ ID NO: 13; SEQ ID NO: 14), E289K (SEQ ID NO: 15; SEQ ID NO: 16), and E292K (SEQ ID NO: 17; SEQ ID NO: 18).

[0006] The present invention also includes T7 DNA ligase mutant amino acid sequences having at least one of the above mutations, provided that the remainder of the T7 DNA ligase mutant amino acid sequence has only conservative substitutions such that the molecule has at least 70%, 80%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the corresponding T7 DNA ligase mutant amino acid sequence in the sequence listing (hereinafter referred to as the "variant sequence").

[0007] The present invention also includes the coding DNA sequences (i.e., the odd-numbered SEQ ID NOs: 1 to 17, respectively) preceding each amino acid sequence of the above mutants, which may or may not encode a C-terminal histidine tag for purification (e.g., a hexameric histidine tag), or other sequences including a linker before the histidine tag at the C-terminus, such as alternating glycine and serine residues; and also includes the foregoing DNA sequences and other degenerate nucleic acid sequences (collectively referred to as "degenerate nucleic acid sequences") which encode (i) each of the above T7 DNA ligase mutants (including the even-numbered SEQ ID NOs: 2 to 18) with or without added sequences encoding tags or linkers, and (ii) the amino acid sequences of any variant sequences.

[0008] The present invention also includes vectors that bind to any degenerate nucleic acid sequence; and cells transformed with any such vector or degenerate nucleic acid sequence and capable of expressing any one of the above T7 DNA ligase mutant amino acid sequences or variant sequences.

[0009] The present invention also includes a composition or a kit, which comprises any one of the above T7 DNA ligase mutant amino acid sequences or variant sequences, degenerate nucleic acid sequences or vectors binding to such degenerate nucleic acid sequences. The present invention also includes a method for amplifying a target nucleic acid, wherein any one of the above T7 DNA ligase mutants or variant sequences is used in a reaction mixture designed to amplify the target nucleic acid, and the reagent mixture is subjected to conditions for amplifying the target nucleic acid.

[0010] Compared with the wild type, the above mutant T7 DNA ligase mutants have greater activity in amplifying target DNA sequences at lower concentrations, and it is expected that the variant sequences also have this greater activity. The mutant ligase mixture and / or mutant T7 DNA ligase can be used in sequencing methods, including ligating adapters to library fragments for subsequent sequencing; or generating oligonucleotides with regions for sequencing after replication in a plasmid.

[0011] Based on the following detailed description and the drawings, additional aspects and advantages of the present disclosure will become clear to those skilled in the art. Only illustrative embodiments of the present disclosure are shown and described in the following specific embodiments. The present disclosure can have other and different embodiments, and several details can be modified in various obvious aspects, all of which do not depart from the present disclosure. Therefore, the description and examples in this summary of the invention are only for illustrative purposes and are not intended to show the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 A series of gel electrophoresis results showing the activity assays comparing wild-type T7 DNA ligase ("WT" in the upper left) with each of the T7 DNA ligase mutants are presented. There are 10 lanes in each gel, and each enzyme has replicated results. The substrates before ligation are included in the figure in the lower left and are labeled "no enzyme (-)". It is known that treatment with T4 DNA ligase has increased activity at blunt ends compared to T7 DNA ligase, which is also included in the second figure from the left in the lower left. Replicate examples of WT T7 DNA ligase activity are in the upper left. The remaining 8 mutant enzymes showing replicate reaction results are as follows: E63K, D132R, E243K, D245R (upper figure, from left to right) and E272K, D288R, E289K, and E292K (lower figure, from left to right). DETAILED DESCRIPTION

[0013] Non-limiting embodiments shown in the accompanying drawings and described in detail below more fully explain the embodiments herein and their various features and advantageous details. Without departing from the present invention, various variations, changes, and alternatives can be conceived by those skilled in the art. It should be understood that various alternatives to the embodiments of the present disclosure can be adopted.

[0014] First, for ease of reference, certain terms used in this application and their meanings as used in context are set forth. If a term used herein is not defined below, the broadest definition given by persons skilled in the relevant art shall be given, as reflected in at least one printed publication or issued patent. Additionally, the present technology is not limited by the use of the terms shown below, as all equivalents, synonyms, newly developed, and terms or techniques used for the same or similar purposes are considered to be within the scope of the claims of the present invention.

[0015] When applied to any feature in the embodiments of the present invention described in the specification and claims, as used herein, the articles "a" and "an" mean one or more. The use of "a" and "an" does not limit their meaning to a single feature unless such a limitation is expressly stated. The article "the" before a singular or plural noun or noun phrase refers to one or more specifically designated features and may have a singular or plural meaning depending on the context in which it is used. The adjective "any" means one, some, or all, regardless of quantity.

[0016] The term "bioactive fragment" refers to any fragment, derivative, homolog, or analogue of a T7 DNA ligase mutant that has in vivo or in vitro activity characteristic of a biomolecule; including, for example, ligase activity, or the repair of mismatches present in nicked DNA via ligation. In some embodiments, a bioactive fragment, derivative, homolog, or analogue of the mutant T7 DNA ligase has any degree of the bioactivity of the mutant T7 DNA ligase in any in vivo or in vitro assay of interest.

[0017] In some embodiments, a bioactive fragment can optionally include any number of contiguous amino acid residues of the mutant T7 DNA ligase. The present invention also includes polynucleotides encoding any such bioactive fragment.

[0018] A bioactive fragment can result from post-transcriptional processing or translation of alternatively spliced RNA, or alternatively can be produced by engineering, bulk synthesis, or other suitable operations. Bioactive fragments include fragments expressed in natural or endogenous cells, as well as fragments produced in expression systems such as bacteria, yeast, plants, insects, or mammalian cells.

[0019] As used herein, the phrase "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid with another amino acid having common properties. A functional way to define common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz (1979) Principles of Protein Structure, Springer-Verlag). Based on such analysis, amino acid groups can be defined such that amino acids within a group preferentially exchange with each other and are thus most similar to each other in terms of their effect on the overall protein structure (Schulz (1979) supra). Examples of amino acid groups defined in this way can include: "charged / polar group", including Glu, Asp, Asn, Gln, Lys, Arg, and His; "aromatic or cyclic group", including Pro, Phe, Tyr, and Trp; and "aliphatic group", including Gly, Ala, Val, Leu, Ile, Met, Ser, Thr, and Cys. Within each group, subgroups can also be identified. For example, the group of charged / polar amino acids can be subdivided into multiple subgroups, including: "positively charged subgroup", including Lys, Arg, and His; "negatively charged subgroup", including Glu and Asp; and "polar subgroup", including Asn and Gln. In another example, the aromatic or cyclic group can be subdivided into multiple subgroups, including: "nitrogenous ring subgroup", including Pro, His, and Trp; "phenyl subgroup", including Phe and Tyr. In yet another further example, the aliphatic group can be subdivided into multiple subgroups, including: "large aliphatic non-polar subgroup", including Val, Leu, and Ile; "aliphatic weakly polar subgroup", including Met, Ser, Thr, and Cys; and "small residue subgroup", including Gly and Ala. Examples of conservative mutations include amino acid substitutions within the above-described subgroups, such as but not limited to: Lys substituting for Arg and vice versa, such that the positive charge can be maintained; Glu substituting for Asp and vice versa, such that the negative charge can be maintained; Ser substituting for Thr and vice versa, such that the free - OH can be maintained; and Gln substituting for Asn and vice versa, such that the free - NH2 can be maintained. A "conservative variant" is a polypeptide that includes one or more amino acids that have been substituted to replace one or more amino acids of a reference polypeptide (e.g., a polypeptide whose sequence is disclosed in a publication or sequence database, or whose sequence has been determined by nucleic acid sequencing) with amino acids having common properties (e.g., belonging to the same amino acid group or subgroup as described above).

[0020] When referring to a gene, "mutant" means that the gene has at least one base (nucleotide) change, deletion, or insertion relative to the native or wild-type gene. The mutation (change, deletion, and / or insertion of one or more nucleotides) can be in the coding region of the gene, or can be in an intron, 3'UTR, 5'UTR, or promoter region. As a non-limiting example, a mutant gene can be a gene with an insertion in the promoter region, which insertion can increase or decrease the expression of the gene; it can be a gene with a deletion, resulting in the production of a non-functional protein, a truncated protein, a dominant negative protein, or no protein; or, it can be a gene with one or more point mutations, resulting in an amino acid change in the encoded protein or abnormal splicing of the gene transcript.

[0021] The terms "mutant T7 DNA ligase of the present invention" and "mutant T7 DNA ligase" when used in this detailed description section, depending on the context, refer jointly or individually to mutant T7 DNA ligase polypeptides that have been tested and exhibit enhanced ligation activity, which are: E63K (SEQ ID NO: 3; SEQ ID NO: 4), D132R (SEQ ID NO: 5; SEQ ID NO: 6), E243K (SEQ ID NO: 7; SEQ ID NO: 8), D245R (SEQ ID NO: 9; SEQ ID NO: 10), E272K (SEQ ID NO: 11; SEQ ID NO: 12), D288R (SEQ ID NO: 13; SEQ ID NO: 14), E289K (SEQ ID NO: 15; SEQ ID NO: 16), and E292K (SEQ ID NO: 17; SEQ ID NO: 18) and / or variant sequences and / or degenerate nucleic acid sequences, as defined in the Summary of the Invention section.

[0022] "Naturally occurring" or "wild-type" refers to the form found in nature. For example, a naturally occurring or wild-type polypeptide or polynucleotide sequence is a sequence that exists in an organism and has not been deliberately modified by human manipulation.

[0023] The term "percent identity" or "homology" with respect to a nucleic acid or polypeptide sequence is defined as the percentage of nucleotide or amino acid residues in a candidate sequence that are identical to a known polypeptide after aligning the sequences to obtain the maximum percentage identity and introducing gaps (if necessary) to achieve the maximum percentage homology. N-terminal or C-terminal insertions or deletions should not be construed as affecting homology. Homology or identity at the nucleotide or amino acid sequence level can be determined by BLAST (Basic Local Alignment Search Tool) analysis using the algorithms employed by the programs blastp, blastn, blastx, tblastn, and tblastx (Altschul (1997), Nucleic Acids Res. 25, 3389-3402 and Karlin (1990), Proc. Natl. Acad. Sci. USA 87, 2264-2268), which are customized for sequence similarity searching. The method used by the BLAST programs is to first consider similar segments (with or without gaps) between the query sequence and the database sequences, then evaluate the statistical significance of all the identified matches, and finally summarize only those matches that meet a preselected significance threshold. For a discussion of the basic issues in sequence database similarity searching, see Altschul (1994), Nature Genetics 6, 119-129. The search parameters for histograms, descriptions, alignments, expectations (i.e., the statistical significance threshold for reporting matches to database sequences), cutoffs, matrices, and filters (low complexity) can be the default settings. The default scoring matrix used by blastp, blastx, tblastn, and tblastx is the BLOSUM62 matrix (Henikoff (1992), Proc. Natl. Acad. Sci. USA 89, 10915-10919), which is recommended for query sequences longer than 85 units (nucleotide bases or amino acids).

[0024] In some embodiments, the invention relates to methods (and related kits, systems, devices, and compositions) for performing ligation reactions, the methods comprising or consisting of contacting a mutant T7 DNA ligase or a biologically active fragment thereof with a nucleic acid template in the presence of one or more nucleotides and ligating at least one of the one or more nucleotides using the mutant T7 DNA ligase or a biologically active fragment thereof.

[0025] In some embodiments, the method can include ligating a double-stranded RNA or DNA polynucleotide chain into a circular molecule. In some embodiments, the method can further include detecting a signal indicative of the ligation by using a sensor. In some embodiments, the sensor is an ISFET. In some embodiments, the sensor can include a detectable label or detectable reagent in the ligation reaction.

[0026] In some embodiments, the present invention relates to methods (and related kits, systems, devices, and compositions) for performing rolling circle amplification of nucleic acids (see U.S. Patent No. 5,714,320, incorporated by reference), which methods use a mutant T7 DNA ligase as the enzyme in the ligation step of the amplification process. Amplification includes amplifying nucleic acids in solution and clonally amplifying nucleic acids on a solid support such as nucleic acid beads, flow cells, nucleic acid arrays, or wells present on the surface of a solid support.

[0027] Preparation of mutant T7 DNA ligase

[0028] The mutant T7 DNA ligase of the present invention can be expressed in any suitable host system, including bacterial, yeast, fungal, baculovirus, plant, or mammalian host cells. For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure include promoters obtained from the Escherichia coli lactose operon, the Streptomyces coelicolor agarase gene (dagA), the Bacillus subtilis levansucrase gene (sacB), the Bacillus licheniformis α-amylase gene (amyL), the Bacillus stearothermophilus maltogenic amylase gene (amyM), the Bacillus amyloliquefaciens α-amylase gene (amyQ), the Bacillus licheniformis penicillinase gene (penP), the Bacillus subtilis xylA and xylB genes, and the prokaryotic β-lactamase gene (Villa-Kamaroff et al., 1978, Proc. Natl. Acad. Sci. USA 75:3727-3731), as well as the tac promoter (DeBoer et al., 1983, Proc. Natl. Acad. Sci. USA 80:21-25).

[0029] For filamentous fungal host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure include promoters obtained from the genes of the following enzymes: Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral α-amylase, Aspergillus niger acid-stable α-amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans acetamidase, and Fusarium oxysporum trypsin-like protease (WO 96 / 00787), as well as the NA2-tpi promoter (a hybrid of the promoters from the Aspergillus niger neutral α-amylase and Aspergillus oryzae triose phosphate isomerase genes), and their mutant, truncated, and hybrid promoters.

[0030] In yeast hosts, useful promoters may be from the genes of the following enzymes: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3-phosphoglycerate kinase. Other useful promoters for yeast host cells are described by Romanos et al., 1992, Yeast 8:423-488.

[0031] For baculovirus expression, insect cell lines derived from Lepidoptera (moths and butterflies), such as Spodoptera frugiperda, are used as hosts. Gene expression is controlled by strong promoters (e.g., pPolh).

[0032] Plant expression vectors are based on the Ti plasmid of Agrobacterium tumefaciens, or on tobacco mosaic virus (TMV), potato virus X, or cowpea mosaic virus. A commonly used constitutive promoter in plant expression vectors is the cauliflower mosaic virus (CaMV) 35S promoter.

[0033] For mammalian expression, cultured mammalian cell lines such as Chinese hamster ovary (CHO), COS, including human cell lines such as HEK and HeLa, can be used to produce mutant T7 DNA ligase. Examples of mammalian expression vectors include adenovirus vectors, pSV and pCMV series plasmid vectors, vaccinia virus, and retroviral vectors, as well as baculovirus. Promoters of cytomegalovirus (CMV) and SV40 are commonly used in mammalian expression vectors to drive gene expression. Non-viral promoters, such as the elongation factor (EF)-1 promoter, are also known.

[0034] The control sequence for expression may also be a suitable transcription terminator sequence, i.e., a sequence recognized by the host cell to terminate transcription. The terminator sequence is operably linked to the 3'-end of the nucleic acid sequence encoding the polypeptide. Any terminator functional in the selected host cell can be used.

[0035] For example, exemplary transcription terminators for filamentous fungal host cells can be obtained from the genes of the following enzymes: Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger α-glucosidase, and Fusarium oxysporum trypsin-like protease.

[0036] Exemplary terminators for yeast host cells can be obtained from the genes of the following enzymes: Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Terminators for insect, plant, and mammalian host cells are also well-known.

[0037] The control sequence may also be a suitable leader sequence, i.e., an untranslated region of the mRNA that is important for translation by the host cell. The leader sequence is operably linked to the 5'-end of the nucleic acid sequence encoding the polypeptide. Any leader sequence functional in the selected host cell can be used. Exemplary leader sequences for filamentous fungal host cells are obtained from the genes of Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase. Suitable leader sequences for yeast host cells are obtained from the genes of the following: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae α-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).

[0038] The control sequence may also be a polyadenylation sequence, a sequence operably linked to the 3'-end of the nucleic acid sequence and recognized by the host cell when transcribed as a signal to add polyadenylate residues to the transcribed mRNA. Any polyadenylation sequence functional in the selected host cell can be used in the present invention. Exemplary polyadenylation sequences for filamentous fungal host cells can be from the genes of the following enzymes: Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsin-like protease, and Aspergillus niger α-glucosidase.

[0039] The control sequence may also be a signal peptide coding region that encodes an amino acid sequence linked to the amino terminus of a polypeptide and directs the encoded polypeptide into the secretory pathway of the cell. The 5' end of the coding sequence of the nucleic acid sequence may itself contain a signal peptide coding region that is naturally linked in the translation reading frame to a segment of the coding region that encodes a secreted polypeptide. Alternatively, the 5' end of the coding sequence may contain a signal peptide coding region that is foreign to the coding sequence. In cases where the coding sequence does not naturally contain a signal peptide coding region, a foreign signal peptide coding region may be required.

[0040] Alternatively, the foreign signal peptide coding region may simply replace the native signal peptide coding region in order to enhance polypeptide secretion. However, any signal peptide coding region that directs the expressed polypeptide into the secretory pathway of the selected host cell may be used.

[0041] An effective signal peptide coding region for bacterial host cells is a signal peptide coding region obtained from the genes of the following enzymes: Bacillus NCIB 11837 maltogenic amylase, Bacillus stearothermophilus α-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis β-lactamase, Bacillus stearothermophilus neutral protease (nprT, nprS, nprM), and Bacillus subtilis prsA. Further signal peptides are described by Simonen and Palva, 1993, Microbiol Rev [Microbiology Reviews] 57:109-137.

[0042] An effective signal peptide coding region for filamentous fungal host cells may be a signal peptide coding region obtained from the genes of the following enzymes: Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase.

[0043] Useful signal peptides for yeast host cells may be from the genes of Saccharomyces cerevisiae α-factor and Saccharomyces cerevisiae invertase. Signal peptides for other host cell systems are also well known.

[0044] The control sequence may also be a propeptide encoding region encoding an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is termed a proenzyme or pro-polypeptide (or in some cases zymogen). The pro-polypeptide is generally inactive and can be converted to the mature, active polypeptide by catalytic or autocatalytic cleavage of the propeptide from the pro-polypeptide. The propeptide encoding region may be obtained from the genes of the following enzymes: Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Saccharomyces cerevisiae α-factor, Rhizomucor miehei aspartic proteinase, and Myceliophthora thermophila lactase (WO 95 / 33836).

[0045] In the case where both a signal peptide and a propeptide region are present at the amino terminus of the polypeptide, the propeptide region is located immediately adjacent to the amino terminus of the polypeptide, and the signal peptide region is located immediately adjacent to the amino terminus of the propeptide region.

[0046] It may also be desirable to add regulatory sequences which allow the expression of the mutant T7 DNA ligase to be regulated with respect to the growth of the host cell. Examples of regulatory systems are those which cause gene expression to be turned on or off in response to a chemical or physical stimulus, including the presence of a regulatory compound. In prokaryotic host cells, suitable regulatory sequences include the lac, tac, and trp operon systems. In yeast host cells, suitable regulatory systems include, for example, the ADH2 system or GAL1 system. In filamentous fungi, suitable regulatory sequences include the TAKA α-amylase promoter, Aspergillus niger glucoamylase promoter, and Aspergillus oryzae glucoamylase promoter. Regulatory systems for other host cells are also well known.

[0047] Other examples of regulatory sequences are those which allow gene amplification. In eukaryotic systems, these include the dihydrofolate reductase gene which is amplified in the presence of methotrexate and the metallothionein genes which are amplified with heavy metals. In these cases, the nucleic acid sequence encoding the polypeptide of the invention will be operably linked to the regulatory sequence.

[0048] Another embodiment includes a recombinant expression vector that contains a polynucleotide encoding an engineered mutant T7 DNA ligase or a variant thereof, and one or more expression regulatory regions such as a promoter and a terminator, as well as an origin of replication, depending on the type of host into which they will be introduced. The various nucleic acids and control sequences described above can be ligated together to produce a recombinant expression vector that may include one or more convenient restriction sites to allow insertion or substitution of the nucleic acid sequence encoding the mutant T7 DNA ligase at such sites. Alternatively, the nucleic acid sequence of the mutant T7 DNA ligase can be expressed by inserting the nucleic acid sequence or a nucleic acid construct containing the sequence into an appropriate vector for expression. When producing the expression vector, the coding sequence is positioned in the vector such that the coding sequence is operably linked to the appropriate control sequences for expression.

[0049] The recombinant expression vector can be any vector (e.g., a plasmid or a virus) that can be conveniently subjected to recombinant DNA procedures and that can cause the expression of the mutant T7 DNA ligase polynucleotide sequence. The choice of the vector will generally depend on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector can be a linear plasmid or a closed circular plasmid.

[0050] The expression vector can be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity, the replication of which is independent of chromosomal replication, such as a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome. The vector can contain any means for ensuring self-replication. Alternatively, the vector can be a vector that, when introduced into the host cell, is integrated into the genome and replicated together with one or more chromosomes into which it has been integrated. In addition, a single vector or plasmid or two or more vectors or plasmids that together contain the total DNA to be introduced into the genome of the host cell can be used, or a transposon can be used.

[0051] The expression vector of the present invention preferably contains one or more selectable markers, which allow for the easy selection of transformed cells. A selectable marker is a gene whose product provides biocide resistance or virus resistance, heavy metal resistance, prototrophy for auxotrophs, etc. Examples of bacterial selectable markers are the dal gene from Bacillus subtilis or Bacillus licheniformis, or markers conferring antibiotic resistance such as ampicillin, kanamycin, chloramphenicol (Example 1) or tetracycline resistance. Suitable markers for yeast host cells are ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selectable markers for filamentous fungal host cells include, but are not limited to, amdS (acetamidase), argB (ornithine carbamoyltransferase), bar (phosphinothricin acetyltransferase), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase), sC (sulfate adenylyltransferase) and trpC (anthranilate synthase), and their equivalents. Embodiments for Aspergillus cells include the amdS and pyrG genes of Aspergillus nidulans or Aspergillus oryzae, and the bar gene of Streptomyces hygroscopicus. Selectable markers for insect, plant and mammalian cells are also well-known.

[0052] The expression vector of the present invention preferably contains one or more elements that allow the vector to integrate into the genome of the host cell or autonomous replication of the vector in the cell independent of the genome. For integration into the genome of the host cell, the vector can rely on the nucleic acid sequence encoding the polypeptide or any other element of the vector for integrating the vector into the genome by homologous or non-homologous recombination.

[0053] Alternatively, the expression vector can contain additional nucleic acid sequences for directing integration into the genome of the host cell by homologous recombination. The additional nucleic acid sequences enable the vector to integrate into one or more precise positions of one or more chromosomes in the genome of the host cell. The integration element can be any sequence homologous to the target sequence within the genome of the host cell. In addition, the integration element can be a non-coding or coding nucleic acid sequence. On the other hand, the vector can integrate into the genome of the host cell by non-homologous recombination.

[0054] For autonomous replication, the vector may also contain an origin of replication, which enables the vector to replicate autonomously in the host cell under discussion. Examples of bacterial origins of replication are the P15A ori, or the origins of replication of plasmids pBR322, pUC19, pACYC177 (which plasmids have the P15A ori) or pACYC184 that permit replication in Escherichia coli, and the origins of replication of plasmids pUB110, pE194, pTA1060 or pAM31 that permit replication in Bacillus. Examples of origins of replication for yeast host cells are the 2 micron origin of replication, ARS1, ARS4, the combination of ARS1 and CEN3, and the combination of ARS4 and CEN6. The origin of replication can be an origin of replication having a mutation that makes its function in the host cell temperature-sensitive (see, for example, Ehrlich, 1978, Proc Natl Acad Sci. USA 75:1433).

[0055] More than one copy of the nucleic acid sequence of the mutant T7 DNA ligase can be inserted into the host cell to increase the production of the gene product. An increased copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene together with the nucleic acid sequence, wherein cells containing the amplified copy of the selectable marker gene and thus the additional copy of the nucleic acid sequence can be selected by culturing the cells in the presence of an appropriate selective reagent.

[0056] Expression vectors for the mutant T7 DNA ligase polynucleotide are commercially available. Suitable commercial expression vectors include the p3xFLAGTM expression vector from Sigma-Aldrich Chemicals, St. Louis, Mo., which includes a CMV promoter and an hGH polyadenylation site for expression in mammalian host cells, and a pBR322 origin of replication and an ampicillin resistance marker for amplification in Escherichia coli. Other suitable expression vectors are pBluescriptII SK(-) and pBK-CMV commercially available from Stratagene, La Jolla, Calif., and plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen) or pPoly (Lathe et al., 1987, Gene 57:193-201).

[0057] Suitable host cells for expressing polynucleotides encoding mutant T7 DNA ligase are well known in the art and include, but are not limited to: bacterial cells such as Escherichia coli, Lactobacillus kefir, Lactobacillus brevis, Lactobacillus minutus, Streptomyces, and Salmonella typhimurium cells; fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC accession number 201178)); insect cells such as Drosophila S2 and Spodoptera Sf9 cells; animal cells such as CHO, COS, BHK, 293, and Bowes melanoma cells; and plant cells. Suitable media and growth conditions for the above host cells are well known in the art.

[0058] The polynucleotides for expressing mutant T7 DNA ligase can be introduced into cells by a variety of methods known in the art. These techniques include electroporation, biolistic particle bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion. The various methods for introducing polynucleotides into cells are known to those skilled in the art.

[0059] The polynucleotides encoding mutant T7 DNA ligase can be prepared by standard solid-phase methods according to known synthetic methods. In some embodiments, fragments of up to about 100 bases can be synthesized individually and then ligated (e.g., by enzymatic or chemical ligation methods, or polymerase-mediated methods) to form any desired continuous sequence. For example, polynucleotides can be prepared by chemical synthesis using, for example, the classical phosphoramidite method described by Beaucage et al., 1981, Tet Lett 22:1859-69, or the method described by Matthes et al., 1984, EMBO J. 3:801-05 (e.g., as it is commonly applied in automated synthesis methods). According to the phosphoramidite method, oligonucleotides are synthesized, for example, in an automated DNA synthesizer, purified, annealed, ligated, and cloned into a suitable vector. Additionally, substantially any nucleic acid can be obtained from a variety of commercial sources such as Midland Certified Reagent Company in Midland, Texas; Great American Gene Company in Ramona, California; ExpressGen in Chicago, Illinois; and Operon Technologies in Alameda, California.

[0060] The engineered mutant T7 DNA ligase expressed in a host cell can be recovered from the cells and / or the culture medium using any one or more well-known protein purification techniques, including lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography. Suitable solutions for the lysis and efficient extraction of proteins from bacteria such as Escherichia coli are commercially available from Sigma-Aldrich of St. Louis under the trade name CelLytic B.TM.

[0061] Chromatographic techniques for separating the mutant T7 DNA ligase include reverse-phase chromatography, high-performance liquid chromatography, ion-exchange chromatography, gel electrophoresis, and affinity chromatography. The purification conditions will depend in part on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, molecular shape, and will be apparent to those skilled in the art.

[0062] In some embodiments, affinity techniques can be used to separate the mutant T7 DNA ligase. For affinity chromatography purification, any antibody that specifically binds to the mutant T7 DNA ligase can be used. To generate the antibody, various host animals (including but not limited to rabbits, mice, rats, etc.) can be immunized by injection with a compound. The compound can be attached to a suitable carrier such as BSA through a side-chain functional group or a linker attached to the side-chain functional group. Various adjuvants can be used to enhance the immune response, depending on the host species, including but not limited to Freund's (complete and incomplete), mineral gels such as aluminum hydroxide, surface-active substances such as lysolecithin, pluronic polyols, polyanions, peptides, oil emulsions, keyhole limpet hemocyanin, dinitrophenol, and potentially useful human adjuvants such as BCG (Bacillus Calmette-Guérin) and Corynebacterium parvum.

[0063] Examples of Preparation of T7 DNA Ligase Mutants

[0064] T7 DNA ligase mutants are generated by conventional PCR mutagenesis, where primers are designed to contain the desired base substitutions, and during the PCR process, the mutations are incorporated into the amplicons, thereby replacing the original sequence. Preferably, the T7 DNA ligase mutants and the wild type have an added C-terminal hexameric His tag for easy purification, preceded by a hexameric series of Ser and Gly residues.

[0065] After PCR, DpnI digestion is performed, which destroys the methylated template (without the substitution), leaving only the unmethylated PCR amplicons with the substitution.

[0066] The PCR amplicons are then directly transformed into chemically competent Escherichia coli host cells, where the bacteria are pretreated with chemicals to enable them to take up and incorporate the plasmid with the amplicon. See ThermoFisher Scientific, Chemically Competent Cells web page (provides kits for generating chemically competent cells).

[0067] The mutant T7 DNA ligase polypeptides expressed by the transformed Escherichia coli host cells are characterized and selected based on a standard ligation assay with gel electrophoresis. The ligase catalyzes the formation of a phosphodiester bond between the 5' and 3' ends of complementary sticky or blunt ends of double-stranded DNA, and the degree of ligation with different T7 DNA ligase mutants can be visualized on an agarose gel using an appropriate DNA dye. In this case, the gel is stained with GelRed (Biotium, Inc., San Francisco, CA) for visualization under UV light. The performance of each T7 DNA ligase mutant is examined based on its ligation activity at reduced enzyme concentrations, and the resulting activity is compared with that of a similarly diluted wild-type ("WT") ligase, allowing determination of which mutants show increased activity compared to the wild-type under the same conditions.

[0068] Blunt-end ligation substrates were prepared for characterizing the mutant T7 DNA ligase polypeptides expressed from the transformed Escherichia coli host cells.

[0069] The DNA vector used was pUC19 (New England Biolabs, catalog number N3041S). PUC19 is a 2686 base pair long double-stranded loop. Digest PUC19.

[0070] Example: Ligation Assay and Results

[0071] To discern activity against blunt-ended substrates, each T7 DNA ligase (wild-type or variant) was tested in duplicate at a protein concentration of 100 ng / μl. For each enzyme sample, the reaction consisted of 2 μl of 5X NEBNext Quick Ligation Reaction Buffer (New England Biolabs, catalog number B6058S, consisting of 1X final concentration of 330 mM Tris-HCl, pH 7.6, 50 mM MgCl2, 5 mM ATP, 5 mM DTT); 10 μl of ΦX174-HaeIII DNA digest (New England Biolabs, catalog number M3026L); 200 mM MgCl2; and water sufficient to bring the total reaction volume to 20 μl. The reaction was incubated at 4 °C for 48 hours, after which it was treated with proteinase K at 55 °C for 30 minutes to stop any further activity and remove any ligase bound to the DNA product that might interfere with reliable gel imaging. Finally, 4 μl of stop solution (which contained 120 mM EDTA, 30% glycerol, 50 mM Tris-HCl pH 8.0, 0.0125% bromophenol blue, 0.1% SDS, and 5x GelRed nucleic acid stain (Biotium, Inc., Fremont, CA)) was added to each reaction.

[0072] Gel electrophoresis using a 0.8% agarose gel was used to visualize the ligation reaction products. Each gel had a set of wild-type T7 DNA ligase samples and variant T7 DNA ligase samples. Each gel was run at 200 V for 25 minutes.

[0073] The results compared to the wild-type are shown in Figure 1 which is a composite image of gel images of T7 DNA WT and 13 variants that meet the criteria for increased activity on sticky-ended dsDNA substrates. The identified mutants that exhibit increased ligation activity against blunt-ended substrates are as follows: E63K, D132R, E243K, D245R, E272K, D288R, E289K, and E292K.

[0074] Using the mutant T7 DNA ligase

[0075] In some embodiments, the T7 mutant ligase is used in sequencing methods, including ligating adapters to library fragments for subsequent sequencing. For example, in some embodiments, the mutant T7 DNA ligase can be used in second-generation (also known as next-generation or Next-Gen), third-generation (also known as Next-Next-Gen), or fourth-generation (also known as N3-Gen) sequencing technologies, including but not limited to pyrosequencing, ligation-based sequencing, single molecule sequencing, sequencing by synthesis (SBS), semiconductor sequencing, massively parallel cloning, massively parallel single molecule SBS, massively parallel single molecule real-time, massively parallel single molecule real-time nanopore technology, etc. Morozova and Marra provide a review of some such technologies in Genomics, 92:255 (2008), which is incorporated herein by reference in its entirety.

[0076] Many DNA sequencing techniques are suitable, including fluorescence-based sequencing methods (see, e.g., Birren et al., Genome Analysis: Analyzing DNA, 1, Cold Spring Harbor, N.Y.; which is incorporated herein by reference in its entirety). In some embodiments, mutant T7 DNA ligase can be used in automated sequencing techniques known in the art. In some embodiments, mutant T7 DNA ligase can be used for parallel sequencing of partitioned amplicons (PCT Publication No.: WO2006084132, which is incorporated herein by reference in its entirety). In some embodiments, mutant T7 DNA ligase can be used for DNA sequencing by parallel oligonucleotide extension (see, e.g., U.S. Patent Nos. 5,750,341 and 6,306,597, both of which are incorporated herein by reference). Other examples of sequencing techniques using mutant T7 DNA ligase include the Church polony technique (Mitra et al., 2003, Analytical Biochemistry 320, 55-65; Shendure et al., 2005 Science 309, 1728-1732; U.S. Patent Nos. 6,432,360, 6,485,944, 6,511,803; all of which are incorporated herein by reference in their entirety), the 454 picotiter pyrosequencing technique (Margulies et al., 2005 Nature 437, 376-380; US20050130173; which is incorporated herein by reference), the Solexa single base addition technique (Bennett et al., 2005, Pharmacogenomics, 6, 373-382; U.S. Patent Nos. 6,787,308; 6,833,246; which are incorporated herein by reference), the Lynx massively parallel signature sequencing technique (Brenner et al. (2000). Nat. Biotechnol. 18:630-634; U.S. Patent Nos. 5,695,934; 5,714,330; all of which are incorporated herein by reference in their entirety) and the Adessi PCR colony technique (Adessi et al. (2000). Nucleic Acid Res. 28, E87; WO 00018957; which is incorporated herein by reference).

[0077] Next-generation sequencing (NGS) methods share the common characteristics of large-scale parallel, high-throughput strategies, with the goal of reducing costs compared to older sequencing methods (see, e.g., Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; each incorporated herein by reference in its entirety). NGS methods can be broadly divided into methods that typically use template amplification and methods that do not use template amplification. Methods that require amplification include pyrosequencing, as the 454 technology platform (e.g., GS20 and GS FLX) by Roche, Life Technologies / Ion Torrent, the Solexa platform commercialized by Illumina and GnuBio, and the Supported Oligonucleotide Ligation and Detection (SOLiD) platform commercialized by Applied Biosystems. Non-amplification methods, also known as single molecule sequencing, are exemplified by the HeliScope platform commercialized by Helicos BioSciences and emerging platforms commercialized by VisiGen, Oxford Nanopore Technologies Ltd., and Pacific Biosciences, respectively.

[0078] In pyrosequencing (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. 6,210,891, 6,258,568; each incorporated herein by reference in its entirety), the template DNA is fragmented, end-repaired, ligated to adapters, and in situ clonally amplified by capturing individual template molecules with beads bearing oligonucleotides complementary to the adapters. Each bead bearing a single template type is partitioned into water-in-oil microbubbles, and the template is clonally amplified using a technique called emulsion PCR. After amplification, the emulsion is disrupted, and the beads are deposited into individual wells of a picotitre plate that serves as a flow cell during the sequencing reaction. In the presence of a sequencing enzyme and a luminescent reporter molecule such as luciferase, each of the four dNTP reagents is introduced into the flow cell in an ordered, iterative manner. If the appropriate dNTP is added to the 3' end of the sequencing primer, the resulting ATP production causes a sudden flash of chemiluminescence within the well, which is recorded using a CCD camera. It is possible to achieve read lengths greater than or equal to 400 bases, and 10 6Sequence reads were obtained, resulting in sequences of up to 500 million base pairs (Mb).

[0079] On the Solexa / Illumina platform (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. 6,833,246; 7,115,400; 6,969,488, each incorporated herein by reference), sequencing data are generated in the form of shorter-length reads. In this method, single-stranded fragmented DNA is end-repaired to produce 5'-phosphorylated blunt ends, and then a single A base is added to the 3' end of the fragment by Klenow-mediated addition. A-addition facilitates the addition of T-overhang adapter oligonucleotides, which are subsequently used to capture template-adapter molecules on the surface of a flow cell covered with oligonucleotide anchors. The anchors are used as PCR primers, but due to the length of the template and its proximity to other nearby anchor oligonucleotides, extension by PCR results in molecules "arching over" and hybridizing to adjacent anchor oligonucleotides to form a bridge structure on the flow cell surface. These DNA loops are denatured and cleaved. Then the forward strand is sequenced using reversible dye terminators. The sequence of the incorporated nucleotide is determined by detecting fluorescence after binding, and each fluor and blocker are removed prior to the next dNTP addition cycle. The sequence read lengths range from 36 nucleotides to more than 250 nucleotides, and the overall output per analysis run exceeds 1 billion nucleotide pairs.

[0080] Sequencing nucleic acid molecules using the SOLiD technology (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. 5,912,148; 6,130,073, each incorporated by reference) also involves fragmentation of the template, ligation to oligonucleotide adapters, ligation to beads, and clonal amplification by emulsion PCR. After this, the beads with the template are immobilized on the derivatized surface of a glass flow cell, and primers complementary to the adapter oligonucleotides are annealed. However, instead of being used for 3' extension, these primers are used to provide a 5' phosphate group for ligation to a detection probe containing two probe-specific bases, followed by six degenerate bases and one of four fluorophore tags. In the SOLiD system, the interrogation probes have 16 possible combinations of two bases at the 3' end of each probe and one of four fluorophores at the 5' end. The fluorophore color and the identity of each probe correspond to a specific color space encoding scheme. Multiple rounds (usually 7 rounds) of probe annealing, ligation, and fluorophore detection are performed, followed by denaturation, and then a second round of sequencing is performed using a primer that is offset by one base relative to the initial primer. In this way, the template sequence can be computationally reconstructed and the template bases are interrogated twice, resulting in increased accuracy. The sequence read length averages 35 nucleotides, and the overall output of each sequencing run exceeds 4 billion bases.

[0081] In certain embodiments, the techniques described herein can be used for nanopore sequencing (see, e.g., Astier et al., J. Am. Chem. Soc. 2006 Feb. 8; 128(5):1705-10, which is incorporated by reference herein). The theory behind nanopore sequencing relates to what happens when a nanopore is immersed in a conducting fluid and a potential (voltage) is applied across it. Under these conditions, a small current due to the conduction of ions through the nanopore can be observed, and the amount of current is extremely sensitive to the size of the nanopore. When each base of a nucleic acid passes through the nanopore, this causes a change in the magnitude of the current passing through the nanopore, which is different for each of the four bases, allowing the sequence of the DNA molecule to be determined.

[0082] In certain embodiments, mutant T7 DNA ligase can be used in Helicos BioSciences' HeliScope (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. 7,169,560; 7,282,337; 7,482,120; 7,501,245; 6,818,395; 6,911,345; 7,501,245, each incorporated herein by reference). The template DNA is fragmented and polyadenylated at the 3' end, and the final adenosine bears a fluorescent tag. The denatured polyadenylated template fragments are ligated to poly(dT) oligonucleotides on the flow cell surface. The initial physical location of the captured template molecules is recorded by a CCD camera, and then the tags are cleaved and washed away. Sequencing is achieved by adding polymerase and successive addition of fluorescently labeled dNTP reagents. Binding events yield a fluorophore signal corresponding to the dNTP, and the signal is captured by a CCD camera prior to each round of dNTP addition. The sequence read length ranges from 25-50 nucleotides, and the overall output per analysis run exceeds 1 billion nucleotide pairs.

[0083] Ion Torrent technology is a DNA sequencing method based on the detection of hydrogen ions released during DNA polymerization (see, e.g., Science 327(5970):1190 (2010); U.S. Patent Application Publication Nos. 20090026082, 20090127589, 20100301398, 20100197507, 20100188073, and 20100137143, which are incorporated by reference). The microwells contain template DNA strands to be sequenced. Below the microwell layer is a high-sensitivity ISFET ion sensor. All layers are contained within a CMOS semiconductor chip, similar to the chips used in the electronics industry. When a dNTP is incorporated into the growing complementary strand, a hydrogen ion is released, triggering the high-sensitivity ion sensor. If a homopolymer repeat is present in the template sequence, multiple dNTP molecules are incorporated in a single cycle. This results in a corresponding number of released hydrogens and a proportionally higher electrical signal. This technology differs from other sequencing technologies in that it does not use modified nucleotides or optics. The per-base accuracy of the Ion Torrent sequencer is approximately 99.6% for 50-base reads, and each run yields approximately 100 Mb to 100 Gb. The read length is 100-300 base pairs. The accuracy for homopolymer repeats of length 5 repeats is approximately 98%. The advantages of ion semiconductor sequencing are fast sequencing speed and low upfront and operating costs.

[0084] The mutant T7 DNA ligase can be used in another nucleic acid sequencing method developed by Stratos Genomics, Inc., and involves the use of Xpandomers. This sequencing process generally includes providing a daughter strand produced by template-directed synthesis. The daughter strand typically includes a plurality of subunits that are joined in a sequence corresponding to all or a portion of the contiguous nucleotide sequence of the target nucleic acid, wherein each subunit includes a tether, at least one probe or nucleobase residue, and at least one selectively cleavable bond. The selectively cleavable bond is cleaved, producing an Xpandomer that is longer than the plurality of subunits of the daughter strand. The Xpandomer typically includes a tether and a reporter element for resolving genetic information in the sequence corresponding to all or a portion of the contiguous nucleotide sequence of the target nucleic acid. The reporter element of the Xpandomer is then detected. Additional details related to the Xpandomer-based method are described, for example, in U.S. Patent Publication No. 20090035777, which is incorporated herein by reference.

[0085] Other single molecule sequencing methods include real-time sequencing-by-synthesis using the VisiGen platform (Voelkerding et al., Clinical Chem., 55:641-58, 2009; U.S. Patent No. 7,329,492; U.S. Patent Application Serial No. 11 / 671,956; U.S. Patent Application Serial No. 11 / 781,166; each of which is incorporated herein by reference), in which a fluorescently modified polymerase and a fluorescent acceptor molecule are used to perform strand extension on a fixed, primed DNA template, thereby obtaining detectable fluorescence resonance energy transfer (FRET) upon addition of nucleotides.

[0086] The specific methods and compositions described herein are representative of the preferred embodiments and are exemplary and not intended to limit the scope of the invention. Given this specification, those skilled in the art will envision other purposes, aspects, and embodiments and which are included within the spirit of the invention as defined by the scope of the claims. It will be apparent to those skilled in the art that various substitutions and modifications can be made to the invention disclosed herein without departing from the scope and spirit of the invention. The invention as illustratively described herein can be practiced appropriately without the presence of any one or more elements, or any one or more limitations, not expressly disclosed herein as necessary. Thus, for example, in each instance in the embodiments or examples of the invention herein, any of the terms “comprising,” “including,” “containing,” etc. should be read broadly and without limitation. The methods and processes illustratively described herein can be appropriately implemented in a different order of steps and they need not be limited to the order of steps indicated herein or in the claims. It should also be noted that, unless the context clearly indicates otherwise, as used herein and in the appended claims, the singular forms “a / an” and “the” include plural referents and the plural includes the singular. In no event should this patent application be construed as limited to the specific examples or embodiments or methods specifically disclosed herein. In no event should this patent application be construed as being limited by any statement made by any examiner or any other official or employee of the Patent and Trademark Office, unless such statement is expressly and unconditionally or specifically adopted by the applicant in responsive written material.

[0087] The invention has been described herein in broad and general terms. Each of the narrower species and subgeneric classifications that fall within the overall disclosure text also forms part of the invention. The terms and expressions that have been employed are used as terms of description and not of limitation, and are not intended to exclude any equivalents or portions thereof of the features shown and described, but it will be recognized that various modifications within the scope of the claimed invention are possible. Accordingly, it is to be understood that, although the invention has been specifically disclosed by preferred embodiments and optional features, those skilled in the art may adopt modifications and variations of the concepts disclosed herein, including but not limited to variant sequences, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.

[0088] Related sequences

[0089] The DNA sequence of wild-type T7 DNA ligase, SEQ ID NO: 1

[0090]

[0091] Amino acid sequence of wild-type T7 DNA ligase, SEQ ID NO: 2

[0092]

[0093] DNA sequence of T7 DNA ligase E63K, SEQ ID NO: 3

[0094]

[0095] Amino acid sequence of T7 DNA ligase E63K, SEQ ID NO: 4

[0096]

[0097] DNA sequence of T7 DNA ligase D132R, SEQ ID NO: 5

[0098]

[0099] Amino acid sequence of T7 DNA ligase D132R, SEQ ID NO: 6

[0100]

[0101] DNA sequence of T7 DNA ligase E243K, SEQ ID NO: 7

[0102]

[0103] Amino acid sequence of T7 DNA ligase E243K, SEQ ID NO: 8

[0104]

[0105] DNA sequence of T7 DNA ligase D245R, SEQ ID NO: 9

[0106]

[0107] Amino acid sequence of T7 DNA ligase D245R, SEQ ID NO: 10

[0108]

[0109] DNA sequence of T7 DNA ligase E272K, SEQ ID NO: 11

[0110]

[0111] Amino acid sequence of T7 DNA ligase E272K, SEQ ID NO: 12

[0112]

[0113] DNA sequence of T7 DNA ligase D288R, SEQ ID NO: 13

[0114]

[0115] Amino acid sequence of T7 DNA ligase D288R, SEQ ID NO: 14

[0116]

[0117] DNA sequence of T7 DNA ligase E289K, SEQ ID NO: 15

[0118]

[0119] Amino acid sequence of T7 DNA ligase E289K, SEQ ID NO: 16

[0120]

[0121] DNA sequence of T7 DNA ligase E292K, SEQ ID NO: 17

[0122]

[0123] Amino acid sequence of T7 DNA ligase E292K, SEQ ID NO: 18

[0124]

Claims

1. A mutant T7 DNA ligase or a bioactive fragment thereof, which comprises one or more of the following amino acid mutations, wherein the mutations are substitutions at the indicated positions in each amino acid sequence, and wherein the entire amino acid sequence is represented by adjacent even-numbered sequence identifiers: E63K (SEQ ID NO: 4), D132R (SEQ ID NO: 6), E243K (SEQ ID NO: 8), D245R (SEQ ID NO: 10), E272K (SEQ ID NO: 12), D288R (SEQ ID NO: 14), E289K (SEQ ID NO: 16), and E292K (SEQ ID NO: 18).

2. The mutant T7 DNA ligase or a bioactive fragment thereof according to claim 1, wherein each of said amino acid sequences further comprises a plurality of histidine residues at its C-terminus.

3. The mutant T7 DNA ligase or a bioactive fragment thereof according to claim 2, which has six histidine residues at its C-terminus.

4. The mutant T7 DNA ligase or a bioactive fragment thereof according to claim 2, wherein a series of alternating glycine and serine amino acid residues is adjacent to and preceding said plurality of histidine residues.

5. A polynucleotide encoding the amino acid sequence of one of the mutant T7 DNA ligases according to claim 1.

6. The polynucleotide according to claim 5, which has one of the following sequences: E63K (SEQ ID NO: 3), D132R (SEQ ID NO: 5), E243K (SEQ ID NO: 7), D245R (SEQ ID NO: 9), E272K (SEQ ID NO: 11), D288R (SEQ ID NO: 13), E289K (SEQ ID NO: 15), and E292K (SEQ ID NO: 17).

7. The mutant T7 DNA ligase or a bioactive fragment thereof according to claim 1, wherein one or more of the even-numbered amino acid sequences have conservative substitutions for some of their amino acids, but only to the extent of maintaining at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with the sequences represented by adjacent sequence identifiers.

8. A polynucleotide encoding the amino acid sequence of one of the mutant T7 DNA ligases according to claim 7.

9. A vector incorporating the polynucleotide according to claim 5.

10. A vector incorporating the polynucleotide according to claim 8.

11. A cell transformed with the polynucleotide according to claim 5 and expressing said polynucleotide.

12. A cell transformed with the vector according to claim 9 and expressing said vector.

13. A method for polynucleotide ligation between different polynucleotides or by joining the 5'-end and 3'-end of a polynucleotide to produce a circular polynucleotide, wherein the polynucleotide has blunt ends or sticky ends, the method comprising: providing a ligation mixture comprising the polynucleotides to be ligated and the mutant T7 DNA ligase or bioactive fragment according to claim 1; and providing a temperature for ligation to occur for the ligation mixture.

14. The method according to claim 13, wherein the ligation mixture comprises Tris-HCl, MgCl2, ATP, dithiothreitol and water.

Citation Information

Patent Citations

  • Methods of amplifying and sequencing nucleic acids

    US20050130173A1

  • Method and apparatus for moving stage detection of single molecular events

    US20080241951A1

  • Methods and apparatus for measuring analytes using large scale FET arrays

    US20090026082A1

  • High throughput nucleic acid sequencing by expansion

    US20090035777A1

  • Methods and apparatus for measuring analytes using large scale FET arrays

    US20090127589A1