T7 DNA ligase variants with increased ligation activity
By performing specific amino acid sequence mutations on T7 DNA ligase, a mutant T7 DNA ligase that enhances ligation activity was prepared, which solved the problem of low activity of existing T7 DNA ligase on sticky terminal substrates, and improved the ligation efficiency and efficiency during sequencing.
Patent Information
- Application Number
- CN202510069613.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-20
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-22
AI Technical Summary
The existing T7 DNA ligases have low activity on sticky terminal substrates when ligating double-stranded DNA polynucleotides, resulting in low ligation efficiency, especially in molecular biology research and sequencing.
By mutation of the specific amino acid sequence of the T7 DNA ligase, mutant T7 DNA ligases that enhance ligation activity were prepared, such as E63K, K73E, K137E, K174E, E182K, K210E, E243K, D245R, E268K, E272K, E289K, K295E and D336R, which enhance their ligation activity to viscosity-terminal substrates at low concentrations.
It enhances the ligation activity of T7 DNA ligase to sticky terminal substrates at low concentrations, improves the ligation efficiency in molecular biology research and sequencing, and reduces the amount and time requirements of enzyme use.
Smart Images

Figure CN120349980A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a T7 DNA ligase variant with increased ligation activity. Background Art Ligases are commonly used in molecular biology to form phosphodiester bonds between double-stranded nucleic acid fragments at the junction of juxtaposed 5'-phosphate and 3'-hydroxyl termini. By designing complementary overhangs between each double-stranded fragment, ligation can be directed to be position-specific and orientationally oriented. This allows for the specific integration of DNA or RNA fragments into larger vectors to meet the needs of molecular biology research. T7 DNA ligase is a thermostable ATP-dependent enzyme that efficiently catalyzes the ligation of double-stranded DNA polynucleotides that exhibit sticky ends longer than a single overhanging base. In addition, it can also repair mismatches present in nicked DNA via ligation.
[0002] Ligases are the backbone of many molecular biology protocols, enabling users to design and ligate various nucleotide molecules for uses such as cloning, protein expression, and sequencing. Depending on the user's goal, increasing the inherent ligation activity of T7 DNA ligase increases the utility of the enzyme, thereby allowing for ligation to require less enzyme or less total time to achieve complete ligation of the substrate.
[0003] Conventionally, ligases are a molecular cloning tool used to insert specific DNA fragments into vectors prior to transformation into competent cells in order to achieve the ligation of sequencing adapters as part of an NGS workflow, or as part of various other molecular biology protocols that require the ligation of two or more fragments of double-stranded DNA polynucleotides. When DNA polynucleotides do not contain unpaired bases at their 3' or 5' ends, they are considered to have blunt ends, in contrast to when overhanging bases extend from the fragment ends, in which case they are considered to have sticky ends, and these overhangs can provide sites for complementary DNA binding and ligation.
[0004] T7 DNA ligase is an ATP-dependent enzyme from a bacteriophage that can catalyze the phosphodiester bond between two complementary overhanging ends of double-stranded DNA fragments, thereby effectively joining two separate DNA fragments together. T7 DNA ligase can also repair mismatches present in nicked DNA by esterifying the 5'-phosphoryl to the 3'-hydroxyl. However, in addition, the ligation of double-stranded DNA substrates to blunt ends is possible in the presence of crowding agents such as polyethylene glycol (PEG). PEG is also used to increase the overall activity of T7 DNA ligase up to 100-fold that of the standard activity and is a common additive in ligase buffers. Summary of the Invention
[0005] The present invention relates to a mutant T7 DNA ligase which exhibits enhanced ligation activity as compared to the wild-type ligase. The following T7 DNA ligase mutants (followed by their amino acid sequence identifiers) have been identified as having such enhanced ligation activity: E63K (SEQ ID NO: 4), K73E (SEQ ID NO: 6), K137E (SEQ ID NO: 7), K174E (SEQ ID NO: 9), E182K (SEQ ID NO: 11), K210E (SEQ ID NO: 13), E243K (SEQ ID NO: 15), D245R (SEQ ID NO: 17), E268K (SEQ ID NO: 19), E272K (SEQ ID NO: 21), E289K (SEQ ID NO: 23), K295E (SEQ ID NO: 25), and D336R (SEQ ID NO: 27).
[0006] The present invention also includes mutant T7 DNA ligase amino acid sequences having at least one of the above mutations, provided that the remaining portion of the T7 DNA ligase mutant amino acid sequence has only conservative substitutions such that the molecule has at least 70%, 80%, 90%, 95%, 96%, 97%, 98% or 99% identity to the corresponding T7 DNA ligase mutant amino acid sequence in the sequence listing (hereinafter referred to as "variant sequences").
[0007] The present invention also includes the DNA sequences preceding each of the amino acid sequences of the above mutants (i.e., SEQ ID NOs: 3, 5, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, respectively), and also includes the foregoing DNA sequences and other degenerate nucleic acid sequences (collectively referred to as "degenerate nucleic acid sequences") which encode (i) each of the above T7 DNA ligase mutants, and (ii) the amino acid sequence of any one of the variant sequences.
[0008] The present invention also includes vectors which bind to any degenerate nucleic acid sequence; and cells which are transformed with any such vector or degenerate nucleic acid sequence and are capable of expressing any one of the above T7 DNA ligase mutant amino acid sequences or variant sequences.
[0009] The present invention also includes a composition or kit which contains any one of the above T7 DNA ligase mutant amino acid sequences or variant sequences, degenerate nucleic acid sequences or vectors which bind to such degenerate nucleic acid sequences. The present invention also includes a method for amplifying a target nucleic acid, wherein any one of the above T7 DNA ligase mutants or variant sequences is employed in a reaction mixture designed to amplify the target nucleic acid, and the reagent mixture is subjected to conditions for amplifying the target nucleic acid.
[0010] Compared with the wild type, the above-mentioned mutant T7 DNA ligase has greater activity in the ligation of cohesive end substrates (polynucleotides) at lower concentrations, and the variant sequences are also expected to have this greater activity. The above-mentioned mutant T7 DNA ligase can be used in sequencing methods, including ligating adapters to library fragments for subsequent sequencing; or generating oligonucleotides with regions for sequencing after replication in a plasmid.
[0011] Based on the following detailed description and the accompanying drawings, additional aspects and advantages of the present disclosure will become apparent to those skilled in the art. Only illustrative embodiments of the present disclosure are shown and described in the following detailed implementation. The present disclosure can have other and different embodiments, and several details can be modified in various obvious aspects, all of which do not depart from the present disclosure. Therefore, the description and examples in this summary of the invention are only for illustrative purposes and are not intended to show the scope of the present disclosure. Brief Description of the Drawings
[0012] Figure 1 A series of gel electrophoresis results showing the activity assays comparing wild-type T7 DNA ligase ("WT" in the upper left) with each of the T7 DNA ligase mutants. There are 12 columns in each gel, such that from left to right, each column represents the concentration of T7 DNA ligase after a 1:2 serial dilution, and in column 1, initially 100 ng of T7 DNA ligase (or mutant) is present in the solution. Each gel has an arrow at the dilution level where a significant amount of ligation product can be clearly seen. Three bands depict the substrates from the product; the supercoiled plasmid product (the upper band, labeled "1" on the right), the restriction-digested linearized plasmid substrate (the middle band, labeled "2" on the right), and the final product (the bottom band, labeled "3" on the right), which is the open circular plasmid product ligated by the wild-type T7 DNA ligase or the indicated mutant. Detailed Description
[0013] The embodiments herein and their various features and advantageous details are more fully explained with reference to the non-limiting embodiments shown in the accompanying drawings and detailed in the following description. Without departing from the present invention, those skilled in the art can conceive of various variations, changes, and substitutions. It should be understood that various alternatives for the embodiments of the present disclosure can be adopted.
[0014] Initially, for ease of reference, certain terms used in this application and their meanings as used in context are set forth. If a term used herein is not defined below, the broadest definition given by persons in the relevant art shall be given, as reflected in at least one printed publication or issued patent. Additionally, the technology is not limited by the use of the terms shown below, as all equivalents, synonyms, newly developed, and terms or technologies used for the same or similar purposes are considered to be within the scope of the claims of this invention.
[0015] When applied to any feature in the embodiments of the invention described in the specification and claims, the articles "a" and "an" as used herein mean one or more. The use of "a" and "an" does not limit their meaning to a single feature unless such a limitation is expressly stated. The article "the" before a singular or plural noun or noun phrase refers to one or more specifically identified features and may have a singular or plural meaning depending on the context in which it is used. The adjective "any" means one, some, or all, regardless of quantity.
[0016] The term "bioactive fragment" refers to any fragment, derivative, homolog, or analogue of a T7 DNA ligase mutant that has an in vivo or in vitro activity characteristic of a biomolecule; including, for example, ligase activity, or the repair of mismatches present in nicked DNA via ligation. In some embodiments, the bioactive fragment, derivative, homolog, or analogue of the mutant T7 DNA ligase has any degree of the bioactivity of the mutant T7 DNA ligase in any in vivo or in vitro assay of interest.
[0017] In some embodiments, the bioactive fragment may optionally comprise any number of contiguous amino acid residues of the mutant T7 DNA ligase. The invention also includes polynucleotides encoding any such bioactive fragment.
[0018] Bioactive fragments can result from post-transcriptional processing or translation of alternatively spliced RNA, or alternatively can be generated by engineering, bulk synthesis, or other suitable manipulations. Bioactive fragments include fragments expressed in natural or endogenous cells, as well as fragments produced in expression systems such as bacterial, yeast, plant, insect, or mammalian cells.
[0019] As used herein, the phrase "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid by another amino acid with common properties. A functional way to define the common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between the corresponding proteins of homologous organisms (Schulz (1979) Principles of Protein Structure, Springer-Verlag). Based on such analysis, amino acid groups can be defined such that the amino acids within a group preferentially exchange with each other and are thus most similar to each other in terms of their effect on the overall protein structure (Schulz (1979) ibid.). Examples of amino acid groups defined in this way can include: "charged / polar group", including Glu, Asp, Asn, Gln, Lys, Arg, and His; "aromatic or cyclic group", including Pro, Phe, Tyr, and Trp; and "aliphatic group", including Gly, Ala, Val, Leu, Ile, Met, Ser, Thr, and Cys. Within each group, subgroups can also be identified. For example, the group of charged / polar amino acids can be subdivided into multiple subgroups, including: "positively charged subgroup", including Lys, Arg, and His; "negatively charged subgroup", including Glu and Asp; and "polar subgroup", including Asn and Gln. In another example, the aromatic or cyclic group can be subdivided into multiple subgroups, including: "nitrogen-containing ring subgroup", including Pro, His, and Trp; "phenyl subgroup", including Phe and Tyr. In yet another further example, the aliphatic group can be subdivided into multiple subgroups, including: "large aliphatic non-polar subgroup", including Val, Leu, and Ile; "aliphatic weakly polar subgroup", including Met, Ser, Thr, and Cys; and "small residue subgroup", including Gly and Ala. Examples of conservative mutations include amino acid substitutions within the above-described subgroups, such as but not limited to: Lys substituting for Arg and vice versa, such that the positive charge can be maintained; Glu substituting for Asp and vice versa, such that the negative charge can be maintained; Ser substituting for Thr and vice versa, such that the free -OH can be maintained; and Gln substituting for Asn and vice versa, such that the free -NH2 can be maintained. A "conservative variant" is a polypeptide that includes one or more amino acids that have been substituted to replace one or more amino acids of a reference polypeptide (e.g., a polypeptide whose sequence is disclosed in a publication or sequence database, or whose sequence has been determined by nucleic acid sequencing) with amino acids having common properties (e.g., belonging to the same amino acid group or subgroup as described above).
[0020] When referring to a gene, "mutant" means that the gene has at least one base (nucleotide) change, deletion, or insertion relative to the native or wild-type gene. The mutation (change, deletion, and / or insertion of one or more nucleotides) can be in the coding region of the gene, or can be in an intron, 3'UTR, 5'UTR, or promoter region. As a non-limiting example, a mutant gene can be a gene with an insertion in the promoter region, which insertion can increase or decrease the expression of the gene; can be a gene with a deletion, resulting in the production of a non-functional protein, a truncated protein, a dominant negative protein, or no protein; or, can be a gene with one or more point mutations, resulting in an amino acid change in the encoded protein or abnormal splicing of the gene transcript.
[0021] The terms "mutant T7 DNA ligase of the present invention" and "mutant T7 DNA ligase" when used in this detailed description section, depending on the context, together or individually refer to mutant T7 DNA ligase polypeptides that have been tested and exhibit enhanced ligation activity, which are: E63K, K73E, K137E, K174E, E182K, K210E, E243K, D245R, E268K, E272K, E289K, K295E, and D336R; and / or variant sequences and / or degenerate nucleic acid sequences, as these terms are defined in the summary of the invention section.
[0022] "Naturally occurring" or "wild-type" refers to the form found in nature. For example, a naturally occurring or wild-type polypeptide or polynucleotide sequence is a sequence that exists in an organism and has not been intentionally modified by human manipulation.
[0023] The term "percent identity" or "homology" with respect to a nucleic acid or polypeptide sequence is defined as the percentage of nucleotide or amino acid residues in a candidate sequence that are identical to a known polypeptide after aligning the sequences to obtain the maximum percent identity and introducing gaps (if necessary) to achieve the maximum percent homology. N-terminal or C-terminal insertions or deletions should not be construed as affecting homology. Homology or identity at the nucleotide or amino acid sequence level can be determined by BLAST (Basic Local Alignment Search Tool) analysis using the algorithms employed by the programs blastp, blastn, blastx, tblastn, and tblastx (Altschul (1997), Nucleic Acids Res. 25, 3389-3402 and Karlin (1990), Proc. Natl. Acad. Sci. USA 87, 2264-2268), which are customized for sequence similarity searching. The method used by the BLAST programs is to first consider similar segments (with or without gaps) between the query sequence and the database sequences, then evaluate the statistical significance of all the identified matches, and finally summarize only those matches that meet a preselected significance threshold. For a discussion of the basic issues in sequence database similarity searching, see Altschul (1994), Nature Genetics 6, 119-129. The search parameters for histograms, descriptions, alignments, expectations (i.e., the statistical significance threshold for reporting matches to database sequences), cutoffs, matrices, and filters (low complexity) can be default settings. The default scoring matrix used by blastp, blastx, tblastn, and tblastx is the BLOSUM62 matrix (Henikoff (1992), Proc. Natl. Acad. Sci. USA 89, 10915-10919), which is recommended for query sequences longer than 85 units (nucleotide bases or amino acids).
[0024] In some embodiments, the invention relates to methods (and related kits, systems, devices, and compositions) for performing ligation reactions, the methods comprising or consisting of contacting a mutant T7 DNA ligase or a bioactive fragment thereof with a nucleic acid template in the presence of one or more nucleotides and ligating at least one of the one or more nucleotides using the mutant T7 DNA ligase or a bioactive fragment thereof.
[0025] In some embodiments, the method can include ligating a double-stranded DNA polynucleotide chain into a circular molecule. In some embodiments, the method can further include detecting a signal indicative of the ligation by using a sensor. In some embodiments, the sensor is an ISFET. In some embodiments, the sensor can include a detectable label or detectable reagent in the ligation reaction.
[0026] In some embodiments, the present invention relates to methods (and related kits, systems, devices, and compositions) for performing rolling circle amplification of nucleic acids (see U.S. Patent No. 5,714,320, incorporated by reference), which methods use a mutant T7 DNA ligase as the enzyme in the ligation step of the amplification process. Amplification includes amplifying nucleic acids in solution and clonal amplification of nucleic acids on a solid support such as nucleic acid beads, flow cells, nucleic acid arrays, or wells present on the surface of a solid support.
[0027] Preparation of mutant T7 DNA ligase
[0028] The mutant T7 DNA ligase of the present invention can be expressed in any suitable host system, including bacterial, yeast, fungal, baculovirus, plant, or mammalian host cells. For bacterial host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure include promoters obtained from the Escherichia coli lactose operon, the Streptomyces coelicolor agarase gene (dagA), the Bacillus subtilis levansucrase gene (sacB), the Bacillus licheniformis α-amylase gene (amyL), the Bacillus stearothermophilus maltogenic amylase gene (amyM), the Bacillus amyloliquefaciens α-amylase gene (amyQ), the Bacillus licheniformis penicillinase gene (penP), the Bacillus subtilis xylA and xylB genes, and the prokaryotic β-lactamase gene (Villa-Kamaroff et al., 1978, Proc. Natl. Acad. Sci. USA 75:3727-3731), as well as the tac promoter (DeBoer et al., 1983, Proc. Natl. Acad. Sci. USA 80:21-25).
[0029] For filamentous fungal host cells, suitable promoters for directing transcription of the nucleic acid constructs of the present disclosure include promoters obtained from the genes of the following enzymes: Aspergillus oryzae TAKA amylase, Rhizomucor miehei aspartic proteinase, Aspergillus niger neutral α-amylase, Aspergillus niger acid-stable α-amylase, Aspergillus niger or Aspergillus awamori glucoamylase (glaA), Rhizomucor miehei lipase, Aspergillus oryzae alkaline protease, Aspergillus oryzae triose phosphate isomerase, Aspergillus nidulans acetamidase, and Fusarium oxysporum trypsin-like protease (WO 96 / 00787), as well as the NA2-tpi promoter (a hybrid of the promoters from the Aspergillus niger neutral α-amylase and Aspergillus oryzae triose phosphate isomerase genes), and their mutant, truncated, and hybrid promoters.
[0030] In yeast hosts, useful promoters can be from the genes of the following enzymes: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP), and Saccharomyces cerevisiae 3-phosphoglycerate kinase. Other useful promoters for yeast host cells are described by Romanos et al., 1992, Yeast 8:423-488.
[0031] For baculovirus expression, insect cell lines derived from Lepidoptera (moths and butterflies), such as Spodoptera frugiperda, are used as hosts. Gene expression is controlled by strong promoters (e.g., pPolh).
[0032] Plant expression vectors are based on the Ti plasmid of Agrobacterium tumefaciens, or on tobacco mosaic virus (TMV), potato virus X, or cowpea mosaic virus. A commonly used constitutive promoter in plant expression vectors is the cauliflower mosaic virus (CaMV) 35S promoter.
[0033] For mammalian expression, cultured mammalian cell lines such as Chinese hamster ovary (CHO), COS, including human cell lines such as HEK and HeLa, can be used to produce mutant T7 DNA ligase. Examples of mammalian expression vectors include adenovirus vectors, pSV and pCMV series plasmid vectors, vaccinia virus, and retroviral vectors, as well as baculovirus. Promoters of cytomegalovirus (CMV) and SV40 are commonly used in mammalian expression vectors to drive gene expression. Non-viral promoters, such as the elongation factor (EF)-1 promoter, are also known.
[0034] A control sequence for expression may also be a suitable transcription terminator sequence, i.e., a sequence recognized by the host cell to terminate transcription. The terminator sequence is operably linked to the 3'-end of the nucleic acid sequence encoding the polypeptide. Any terminator functional in the selected host cell may be used.
[0035] For example, exemplary transcription terminators for filamentous fungal host cells may be obtained from the genes of the following enzymes: Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Aspergillus niger α-glucosidase, and Fusarium oxysporum trypsin-like protease.
[0036] Exemplary terminators for yeast host cells may be obtained from the genes of the following enzymes: Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase.
[0037] Terminators for insect, plant, and mammalian host cells are also well-known.
[0038] The control sequence may also be a suitable leader sequence, i.e., an untranslated region of the mRNA that is important for translation by the host cell. The leader sequence is operably linked to the 5'-end of the nucleic acid sequence encoding the polypeptide. Any leader sequence functional in the selected host cell may be used. Exemplary leader sequences for filamentous fungal host cells are obtained from the genes of Aspergillus oryzae TAKA amylase and Aspergillus nidulans triose phosphate isomerase. Suitable leader sequences for yeast host cells are obtained from the genes of the following: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae α-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).
[0039] The control sequence may also be a polyadenylation sequence, a sequence operably linked to the 3'-end of the nucleic acid sequence and recognized by the host cell when transcribed as a signal to add polyadenylate residues to the transcribed mRNA. Any polyadenylation sequence functional in the selected host cell may be used in the present invention. Exemplary polyadenylation sequences for filamentous fungal host cells may be from the genes of the following enzymes: Aspergillus oryzae TAKA amylase, Aspergillus niger glucoamylase, Aspergillus nidulans anthranilate synthase, Fusarium oxysporum trypsin-like protease, and Aspergillus niger α-glucosidase.
[0040] The control sequence may also be a signal peptide coding region that encodes an amino acid sequence linked to the amino terminus of a polypeptide and directs the encoded polypeptide into the secretory pathway of the cell. The 5' end of the coding sequence of the nucleic acid sequence may itself contain a signal peptide coding region that is naturally linked to a segment of the coding region encoding the secreted polypeptide in the translation reading frame. Alternatively, the 5' end of the coding sequence may contain a signal peptide coding region that is foreign to the coding sequence. In cases where the coding sequence does not naturally contain a signal peptide coding region, a foreign signal peptide coding region may be required.
[0041] Alternatively, the foreign signal peptide coding region may simply replace the native signal peptide coding region in order to enhance polypeptide secretion. However, any signal peptide coding region that directs the expressed polypeptide into the secretory pathway of the selected host cell may be used.
[0042] Effective signal peptide coding regions for bacterial host cells are signal peptide coding regions obtained from the genes of the following enzymes: Bacillus NCIB 11837 maltogenic amylase, Bacillus stearothermophilus α-amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis β-lactamase, Bacillus stearothermophilus neutral proteases (nprT, nprS, nprM), and Bacillus subtilis prsA. Further signal peptides are described by Simonen and Palva, 1993, Microbiol Rev [Microbiological Reviews] 57:109-137.
[0043] Effective signal peptide coding regions for filamentous fungal host cells may be signal peptide coding regions obtained from the genes of the following enzymes: Aspergillus oryzae TAKA amylase, Aspergillus niger neutral amylase, Aspergillus niger glucoamylase, Rhizomucor miehei aspartic proteinase, Humicola insolens cellulase, and Humicola lanuginosa lipase.
[0044] Useful signal peptides for yeast host cells may be derived from the genes of Saccharomyces cerevisiae α-factor and Saccharomyces cerevisiae invertase. Signal peptides for other host cell systems are also well known.
[0045] The control sequence may also be a propeptide encoding region encoding an amino acid sequence located at the amino terminus of the polypeptide. The resulting polypeptide is referred to as a proenzyme or pro-polypeptide (or in some cases as a zymogen). Pro-polypeptides are generally inactive and can be converted to mature, active polypeptides by catalytic or autocatalytic cleavage of the propeptide from the pro-polypeptide. The propeptide encoding region may be obtained from the genes of the following enzymes: Bacillus subtilis alkaline protease (aprE), Bacillus subtilis neutral protease (nprT), Saccharomyces cerevisiae α-factor, Rhizomucor miehei aspartic proteinase, and Myceliophthora thermophila lactase (WO 95 / 33836).
[0046] In the case where both a signal peptide and a propeptide region are present at the amino terminus of the polypeptide, the propeptide region is located immediately adjacent to the amino terminus of the polypeptide, and the signal peptide region is located immediately adjacent to the amino terminus of the propeptide region.
[0047] It may also be desirable to add regulatory sequences which allow the expression of the mutant T7 DNA ligase to be regulated relative to the growth of the host cell. Examples of regulatory systems are those which cause gene expression to be turned on or off in response to a chemical or physical stimulus, including the presence of a regulatory compound. In prokaryotic host cells, suitable regulatory sequences include the lac, tac, and trp operator systems. In yeast host cells, suitable regulatory systems include, for example, the ADH2 system or the GAL1 system. In filamentous fungi, suitable regulatory sequences include the TAKA α-amylase promoter, the Aspergillus niger glucoamylase promoter, and the Aspergillus oryzae glucoamylase promoter. Regulatory systems for other host cells are also well known.
[0048] Other examples of regulatory sequences are those which allow gene amplification. In eukaryotic systems, these include the dihydrofolate reductase gene which is amplified in the presence of methotrexate and the metallothionein genes which are amplified with heavy metals. In these cases, the nucleic acid sequence encoding the KRED polypeptide of the invention will be operably linked to the regulatory sequence.
[0049] Another embodiment includes a recombinant expression vector that comprises a polynucleotide encoding an engineered mutant T7 DNA ligase or a variant thereof, and one or more expression regulatory regions such as a promoter and a terminator, as well as an origin of replication, depending on the type of host into which they are to be introduced. The various nucleic acids and control sequences described above can be joined together to generate a recombinant expression vector, which may include one or more convenient restriction sites to allow the insertion or substitution of the nucleic acid sequence encoding the mutant T7 DNA ligase at such sites. Alternatively, the nucleic acid sequence of the mutant T7 DNA ligase can be expressed by inserting the nucleic acid sequence or a nucleic acid construct containing the sequence into an appropriate vector for expression. When generating the expression vector, the coding sequence is positioned in the vector such that the coding sequence is operably linked to an appropriate control sequence for expression.
[0050] The recombinant expression vector can be any vector (e.g., a plasmid or a virus) that can be conveniently subjected to recombinant DNA procedures and that can cause the expression of the mutant T7 DNA ligase polynucleotide sequence. The choice of the vector will typically depend on the compatibility of the vector with the host cell into which the vector is to be introduced. The vector can be a linear plasmid or a closed circular plasmid.
[0051] The expression vector can be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity, the replication of which is independent of chromosomal replication, such as a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome. The vector can contain any means for ensuring self-replication. Alternatively, the vector can be a vector that, when introduced into a host cell, is integrated into the genome and replicated together with one or more chromosomes into which it has been integrated. In addition, a single vector or plasmid or two or more vectors or plasmids that together contain the total DNA to be introduced into the genome of the host cell can be used, or a transposon can be used.
[0052] The expression vector of the present invention preferably contains one or more selectable markers, which allow for the easy selection of transformed cells. A selectable marker is a gene whose product provides biocide resistance or virus resistance, heavy metal resistance, prototrophy for auxotrophs, etc. Examples of bacterial selectable markers are the dal gene from Bacillus subtilis or Bacillus licheniformis, or markers conferring antibiotic resistance such as ampicillin, kanamycin, chloramphenicol (Example 1), or tetracycline resistance. Suitable markers for yeast host cells are ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3. Selectable markers for filamentous fungal host cells include, but are not limited to, amdS (acetamidase), argB (ornithine carbamoyltransferase), bar (phosphinothricin acetyltransferase), hph (hygromycin phosphotransferase), niaD (nitrate reductase), pyrG (orotidine-5'-phosphate decarboxylase), sC (sulfate adenylyltransferase), and trpC (anthranilate synthase), and their equivalents. Embodiments for Aspergillus cells include the amdS and pyrG genes of Aspergillus nidulans or Aspergillus oryzae, and the bar gene of Streptomyces hygroscopicus. Selectable markers for insect, plant, and mammalian cells are also well known.
[0053] The expression vector of the present invention preferably contains one or more elements that allow the vector to integrate into the genome of the host cell or to replicate autonomously in the cell independently of the genome. For integration into the genome of the host cell, the vector can rely on the nucleic acid sequence encoding the polypeptide or any other element of the vector used for integrating the vector into the genome by homologous or non-homologous recombination.
[0054] Alternatively, the expression vector can contain additional nucleic acid sequences for directing integration into the genome of the host cell by homologous recombination. The additional nucleic acid sequences enable the vector to integrate into one or more precise positions of one or more chromosomes in the genome of the host cell. The integration element can be any sequence homologous to the target sequence within the genome of the host cell. In addition, the integration element can be a non-coding or coding nucleic acid sequence. On the other hand, the vector can be integrated into the genome of the host cell by non-homologous recombination.
[0055] For autonomous replication, the vector may also contain an origin of replication, which enables the vector to replicate autonomously in the host cell under discussion. Examples of bacterial origins of replication are P15A ori, or the origins of replication of plasmids pBR322, pUC19, pACYC177 (which plasmids have P15A ori) or pACYC184 that permit replication in Escherichia coli, and the origins of replication of plasmids pUB110, pE194, pTA1060 or pAM31 that permit replication in Bacillus. Examples of origins of replication for yeast host cells are the 2 micron origin of replication, ARS1, ARS4, the combination of ARS1 and CEN3, and the combination of ARS4 and CEN6. The origin of replication can be an origin of replication having a mutation that renders its function in the host cell temperature-sensitive (see, for example, Ehrlich, 1978, Proc Natl Acad Sci. USA 75:1433).
[0056] More than one copy of the nucleic acid sequence of mutant T7 DNA ligase can be inserted into the host cell to increase the production of the gene product. An increased copy number of the nucleic acid sequence can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selectable marker gene together with the nucleic acid sequence, wherein cells containing the amplified copy of the selectable marker gene and thus the additional copy of the nucleic acid sequence can be selected by culturing the cells in the presence of an appropriate selective reagent.
[0057] Expression vectors for mutant T7 DNA ligase polynucleotides are commercially available. Suitable commercial expression vectors include the p3xFLAGTM expression vector from Sigma-Aldrich Chemicals, St. Louis, Mo., which includes a CMV promoter and an hGH polyadenylation site for expression in mammalian host cells, as well as a pBR322 origin of replication and an ampicillin resistance marker for amplification in Escherichia coli. Other suitable expression vectors are pBluescriptII SK(-) and pBK-CMV, commercially available from Stratagene, La Jolla, Calif., and plasmids derived from pBR322 (Gibco BRL), pUC (Gibco BRL), pREP4, pCEP4 (Invitrogen), or pPoly (Lathe et al., 1987, Gene 57:193-201).
[0058] Suitable host cells for expressing polynucleotides encoding mutant T7 DNA ligase are well known in the art and include, but are not limited to: bacterial cells such as Escherichia coli, Lactobacillus kefir, Lactobacillus brevis, Lactobacillus minutus, Streptomyces, and Salmonella typhimurium cells; fungal cells such as yeast cells (e.g., Saccharomyces cerevisiae or Pichia pastoris (ATCC accession number 201178)); insect cells such as Drosophila S2 and Spodoptera Sf9 cells; animal cells such as CHO, COS, BHK, 293, and Bowes melanoma cells; and plant cells. Suitable media and growth conditions for the above host cells are well known in the art.
[0059] The polynucleotides for expressing mutant T7 DNA ligase can be introduced into cells by a variety of methods known in the art. These techniques include electroporation, biolistic particle bombardment, liposome-mediated transfection, calcium chloride transfection, and protoplast fusion. The various methods for introducing polynucleotides into cells are known to those skilled in the art.
[0060] The polynucleotides encoding mutant T7 DNA ligase can be prepared by standard solid-phase methods according to known synthetic methods. In some embodiments, fragments of up to about 100 bases can be synthesized individually and then ligated (e.g., by enzymatic or chemical ligation methods, or polymerase-mediated methods) to form any desired continuous sequence. For example, polynucleotides can be prepared by chemical synthesis using, for example, the classical phosphoramidite method described by Beaucage et al., 1981, Tet Lett 22:1859-69, or the method described by Matthes et al., 1984, EMBO J. 3:801-05 (e.g., as it is commonly applied to automated synthesis methods). According to the phosphoramidite method, oligonucleotides are synthesized, for example, in an automated DNA synthesizer, purified, annealed, ligated, and cloned into a suitable vector. Additionally, substantially any nucleic acid can be obtained from a variety of commercial sources such as Midland Certified Reagent Company in Midland, Texas; Great American Gene Company in Ramona, California; ExpressGen in Chicago, Illinois; and Operon Technologies in Alameda, California.
[0061] The engineered mutant T7 DNA ligase expressed in a host cell can be recovered from the cells and / or the culture medium using any one or more well-known protein purification techniques, including lysozyme treatment, sonication, filtration, salting out, ultracentrifugation, and chromatography. Suitable solutions for lysing bacteria such as Escherichia coli and efficiently extracting proteins are commercially available from Sigma-Aldrich of St. Louis under the trade name CelLytic B.TM.
[0062] Chromatographic techniques for separating the mutant T7 DNA ligase include reverse-phase chromatography, high-performance liquid chromatography, ion-exchange chromatography, gel electrophoresis, and affinity chromatography. The purification conditions will depend in part on factors such as net charge, hydrophobicity, hydrophilicity, molecular weight, and molecular shape, and will be apparent to those skilled in the art.
[0063] In some embodiments, affinity techniques can be used to separate the mutant T7 DNA ligase. For affinity chromatography purification, any antibody that specifically binds the mutant T7 DNA ligase can be used. To generate the antibody, various host animals (including but not limited to rabbits, mice, rats, etc.) can be immunized by injection of a compound. The compound can be attached to a suitable carrier such as BSA through a side-chain functional group or a linker attached to the side-chain functional group. Various adjuvants can be used to enhance the immune response, depending on the host species, including but not limited to Freund's (complete and incomplete), mineral gels such as aluminum hydroxide, surface-active substances such as lysolecithin, pluronic polyols, polyanions, peptides, oil emulsions, keyhole limpet hemocyanin, dinitrophenol, and potentially useful human adjuvants such as BCG (Bacillus Calmette-Guérin) and Corynebacterium parvum.
[0064] Examples of preparing mutants of T7 DNA ligase
[0065] T7 DNA ligase mutants are generated by conventional PCR mutagenesis, where primers are designed to contain the desired base substitutions, and during the PCR process, the mutations are incorporated into the amplicons, replacing the original sequence. The T7 DNA ligase mutants and the wild type preferably have an added C-terminal hexameric His tag for easy purification, preceded by a hexamer series of Ser and Gly residues (Gly Ser Gly Ser Ser Gly His His His His His His).
[0066] After PCR, DpnI digestion is performed, which destroys the methylated template (without the substitution), leaving only the unmethylated PCR amplicons with the substitution.
[0067] The PCR amplicons are then directly transformed into chemically competent Escherichia coli host cells, where the bacteria are pretreated with chemicals to enable them to take up and incorporate the plasmid of the amplicon. See ThermoFisher Scientific, Chemically Competent Cells web page (providing kits for generating chemically competent cells).
[0068] The mutant T7 DNA ligase polypeptides expressed by the transformed Escherichia coli host cells are characterized and selected based on a standard ligation assay with gel electrophoresis. The ligase catalyzes the formation of a phosphodiester bond between the 5' and 3' ends of complementary sticky or blunt ends of double-stranded DNA, and the degree of ligation with different T7 DNA ligase mutants can be visualized on an agarose gel using an appropriate DNA dye. In this case, the gel is stained with GelRed (Biotium, Inc., San Francisco, CA) for visualization under UV light. The performance of each T7 DNA ligase mutant is examined based on its ligation activity at reduced enzyme concentrations, and the resulting activity is compared with that of a similarly diluted wild-type ("WT") ligase, allowing determination of which mutants show increased activity compared to the wild type under the same conditions.
[0069] The ligation substrate for characterizing the mutant T7 DNA ligase polypeptides expressed by the transformed Escherichia coli host cells is prepared as follows.
[0070] The DNA vector used is pUC19 (New England Biolabs, catalog number N3041S). PUC19 is a 2,686 base pair long double-stranded loop. pUC19 is digested with (New England Biolabs, catalog number R3733S), which cuts the 5' strand after adding a random (N1) nucleotide to the recognition sequence and cuts the 3' strand after adding five additional random (N5) nucleotides to the complementary recognition sequence. The 5' recognition sequence is GGTCTC. The cleavage site of pUC19 is designated as 5'-GGTCTC(N1) / (N5)-3'.
[0071] 5 μl of pUC19 at a concentration of 1 mg / ml is combined with 2.5 μl of 20,000 units / ml of 5 μl of 10X rCutSmart TMBuffer (New England Biolabs, catalog number B6004S) (50 mM potassium acetate, 20 mM Tris-acetate, 10 mM magnesium acetate, 100 μg / ml recombinant albumin) was combined with 35 μl of water. After combining all components, the composition was incubated at 37 °C for digestion. After 1 hour at 37 °C, the reaction was incubated at 80 °C for 20 minutes to heat inactivate. The mixture was then diluted with water to a concentration of 10 ng / μl.
[0072] Example: Ligation Assay and Results
[0073] The ligation procedure was carried out as follows. Each T7 DNA ligase (whether wild-type or variant) was diluted with enzyme diluent (50% glycerol, 10 mM tris-HCl) in serial dilutions such that 10 different concentrations were obtained for each sample. The starting concentration for each serial dilution was 100 ng / ul T7 DNA ligase, diluted two-fold in the next dilution such that the final concentration was 50% of the sample concentration before dilution, and so on, for a total of 10 1:2 dilutions. Then 2 μl of each enzyme or serial dilution was added to a PCR plate such that the approximate amounts of enzyme in lanes 1 to 12 were 100, 50, 25, 12.5, 6.25, 3.1, 1.5, 0.7, 0.39, 0.2, 0.1, and 0.05 ng, respectively.
[0074] For 2 μl of enzyme, 13 μl of a master mixture consisting of 2 μl of 5X NEBNext Quick Ligation Reaction Buffer (New England Biolabs, catalog number B6058S; 1X components yield 66 mM Tris-HCl, 10 mM MgCl2, 1 mM ATP, 10 mM DTT [dithiothreitol], 7.5% polyethylene glycol [PEG 6000], pH 7.6), 0.5 μl of 10 ng / μl digested pUC19, and 13.5 μl of water was added to each reaction to achieve a total reaction volume of 20 μl. The reactions were incubated at 16 °C for 20 minutes. Then 4 μl of termination solution (120 mM EDTA, 30% glycerol, 50 mM Tris-HCl pH 8.0, 0.0125% bromophenol blue, 0.1% SDS, and 5x GelRed nucleic acid stain (Biotium, Inc., Fremont, CA)) was added to each reaction.
[0075] Gel electrophoresis using a 0.8% agarose gel was used to visualize the ligation reaction products. Each gel had a series of samples of wild-type T7 DNA ligase and 12 series of samples of variant T7 DNA ligase. Each gel was run at 200 V for 25 minutes.
[0076] The results compared to the wild type are shown in Figure 1 which is a composite image of gel images of T7 DNA WT and 13 variants that meet the criteria for increased activity on sticky-end dsDNA substrates. The identified mutants that exhibit increased ligation activity are as follows: E63K, K73E, K137E, K174E, E182K, K210E, E243K, D245R, E268K, E272K, E289K, K295E, and D336R.
[0077] The comparative activity results are listed in Table 1 below, showing the increase in variant activity compared to WT as calculated from the lane differences for each T7 DNA ligase mutant, where each lane difference greater than WT is assigned a 2-fold increase in activity. For example, a 1-lane activity improvement relative to WT is assigned a value of 2 (as indicated by little or no distinct upper band representing supercoiled plasmid product and little or no distinct middle band representing restriction-digested linear bottom plasmid in the lane), while a 2-lane activity improvement relative to WT is assigned a value of 4 and so on.
[0078] Table 1: T7 DNA Ligase Variants with Activity Greater than WT
[0079] Mutant Activity relative to WT (fold) E63K 2 K73E 2 K146E 8 E174K 2 E182K 2 K210E 2 E243K 8 D245R 4 E268K 2 E272K 16 E289K 2 K295E 4 D336R 2
[0080] Using the mutant T7 DNA ligase
[0081] In certain embodiments, the T7 mutant ligase is used in sequencing methods, including ligating adapters to library fragments for subsequent sequencing. For example, in some embodiments, the mutant T7 DNA ligase can be used in second-generation (also known as next-generation or Next-Gen), third-generation (also known as Next-Next-Gen), or fourth-generation (also known as N3-Gen) sequencing technologies, including but not limited to pyrosequencing, sequencing-by-ligation, single molecule sequencing, sequencing-by-synthesis (SBS), semiconductor sequencing, massively parallel cloning, massively parallel single molecule SBS, massively parallel single molecule real-time, massively parallel single molecule real-time nanopore technology, etc. A review of some such technologies is provided by Morozova and Marra in Genomics, 92:255 (2008), which is incorporated herein by reference in its entirety.
[0082] Many DNA sequencing techniques are suitable, including fluorescence-based sequencing methods (see, e.g., Birren et al., Genome Analysis: Analyzing DNA, 1, Cold Spring Harbor, N.Y.; which is incorporated herein by reference in its entirety). In some embodiments, mutant T7 DNA ligase can be used in automated sequencing techniques known in the art. In some embodiments, mutant T7 DNA ligase can be used for parallel sequencing of partitioned amplicons (PCT Publication No.: WO2006084132, which is incorporated herein by reference in its entirety). In some embodiments, mutant T7 DNA ligase can be used for DNA sequencing by parallel oligonucleotide extension (see, e.g., U.S. Patent Nos. 5,750,341 and 6,306,597, both of which are incorporated herein by reference). Other examples of sequencing techniques using mutant T7 DNA ligase include the Church polony technique (Mitra et al., 2003, Analytical Biochemistry 320, 55-65; Shendure et al., 2005 Science 309, 1728-1732; U.S. Patent Nos. 6,432,360, 6,485,944, 6,511,803; all of which are incorporated herein by reference in their entirety), the 454 picotiter pyrosequencing technique (Margulies et al., 2005 Nature 437, 376-380; US20050130173; which is incorporated herein by reference), the Solexa single base addition technique (Bennett et al., 2005, Pharmacogenomics, 6, 373-382; U.S. Patent Nos. 6,787,308; 6,833,246; which are incorporated herein by reference), the Lynx massively parallel signature sequencing technique (Brenner et al. (2000). Nat. Biotechnol. 18:630-634; U.S. Patent Nos. 5,695,934; 5,714,330; all of which are incorporated herein by reference in their entirety) and the Adessi PCR colony technique (Adessi et al. (2000). Nucleic Acid Res. 28, E87; WO 00018957; which is incorporated herein by reference).
[0083] Next-generation sequencing (NGS) methods share the common characteristics of large-scale parallel, high-throughput strategies, with the goal of reducing costs compared to older sequencing methods (see, e.g., Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; each incorporated herein by reference in its entirety). NGS methods can be broadly divided into methods that typically use template amplification and methods that do not use template amplification. Methods that require amplification include pyrosequencing by Roche as the 454 technology platform (e.g., GS20 and GS FLX), the Ion Torrent by Life Technologies / Ion Torrent, the Solexa platform commercialized by Illumina and GnuBio, and the Supported Oligonucleotide Ligation and Detection (SOLiD) platform commercialized by Applied Biosystems. Non-amplification methods, also known as single-molecule sequencing, are exemplified by the HeliScope platform commercialized by Helicos BioSciences and emerging platforms commercialized by VisiGen, Oxford Nanopore Technologies Ltd., and Pacific Biosciences, respectively.
[0084] In pyrosequencing (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. 6,210,891, 6,258,568; each incorporated herein by reference in its entirety), template DNA is fragmented, end-repaired, ligated to adapters, and in situ clonally amplified by capturing individual template molecules with beads carrying oligonucleotides complementary to the adapters. Each bead carrying a single template type is partitioned into water-in-oil microbubbles, and the template is clonally amplified using a technique called emulsion PCR. After amplification, the emulsion is disrupted, and the beads are deposited into individual wells of a picotitre plate that serves as a flow cell during the sequencing reaction. In the presence of a sequencing enzyme and a luminescent reporter molecule such as luciferase, each of the four dNTP reagents is introduced into the flow cell in an ordered, iterative manner. If the appropriate dNTP is added to the 3' end of the sequencing primer, the resulting ATP production causes a sudden flash of chemiluminescence within the well, which is recorded using a CCD camera. Read lengths of greater than or equal to 400 bases are possible, and 10 6Individual sequence reads, resulting in sequences of up to 500 million base pairs (Mb).
[0085] In the Solexa / Illumina platform (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. 6,833,246; 7,115,400; 6,969,488, each incorporated herein by reference), sequencing data is generated in the form of shorter length reads. In this method, single-stranded fragmented DNA is end-repaired to produce 5'-phosphorylated blunt ends, and then a single A base is added to the 3' end of the fragment by Klenow-mediated addition. A-addition facilitates the addition of T-overhang adapter oligonucleotides, which are subsequently used to capture template-adapter molecules on the surface of a flow cell covered with oligonucleotide anchors. The anchors are used as PCR primers, but due to the length of the template and its proximity to other nearby anchor oligonucleotides, extension by PCR results in molecules "arching over" and hybridizing to adjacent anchor oligonucleotides to form a bridge structure on the surface of the flow cell. These DNA loops are denatured and cleaved. Then the forward strand is sequenced with reversible dye terminators. The sequence of the incorporated nucleotide is determined by detecting fluorescence after binding, and each fluor and blocker is removed prior to the next dNTP addition cycle. The sequence read lengths range from 36 nucleotides to more than 250 nucleotides, and the overall output per analysis run exceeds 1 billion nucleotide pairs.
[0086] Sequencing nucleic acid molecules using the SOLiD technology (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. 5,912,148; 6,130,073, each incorporated by reference) also involves fragmentation of the template, ligation to oligonucleotide adapters, ligation to beads, and clonal amplification by emulsion PCR. After this, the beads with the templates are immobilized on a derivatized surface of a glass flow cell, and primers complementary to the adapter oligonucleotides are annealed. However, this primer is not for 3' extension but for providing a 5' phosphate group for ligation to a detection probe containing two probe-specific bases, followed by six degenerate bases and one of four fluorophore tags. In the SOLiD system, the interrogation probes have 16 possible combinations of two bases at the 3' end of each probe and one of four fluorophores at the 5' end. The fluorophore color, as well as the identity of each probe, corresponds to a specific color space encoding scheme. Multiple rounds (usually 7 rounds) of probe annealing, ligation, and fluorophore detection are performed, followed by denaturation, and then a second round of sequencing is performed using a primer that is offset by one base relative to the initial primer. In this way, the template sequence can be computationally reconstructed, and the template bases are interrogated twice, resulting in increased accuracy. The sequence read length averages 35 nucleotides, and the overall output per sequencing run exceeds 4 billion bases.
[0087] In certain embodiments, the techniques described herein can be used for nanopore sequencing (see, e.g., Astier et al., J. Am. Chem. Soc. 2006 Feb. 8; 128(5):1705-10, which is incorporated by reference herein). The theory behind nanopore sequencing relates to what happens when a nanopore is immersed in a conductive fluid and a potential (voltage) is applied across it. Under these conditions, a small current due to the conduction of ions through the nanopore can be observed, and the amount of current is extremely sensitive to the size of the nanopore. When each base of a nucleic acid passes through the nanopore, this causes a change in the magnitude of the current passing through the nanopore, which is different for each of the four bases, allowing the sequence of the DNA molecule to be determined.
[0088] In certain embodiments, mutant T7 DNA ligase can be used in Helicos BioSciences' HeliScope (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. 7,169,560; 7,282,337; 7,482,120; 7,501,245; 6,818,395; 6,911,345; 7,501,245, each of which is incorporated herein by reference). The template DNA is fragmented and polyadenylated at the 3' end, and the final adenosine bears a fluorescent label. The denatured polyadenylated template fragments are ligated to poly(dT) oligonucleotides on the flow cell surface. The initial physical location of the captured template molecules is recorded by a CCD camera, and then the label is cleaved and washed away. Sequencing is achieved by adding polymerase and successive addition of fluorescently labeled dNTP reagents. Binding events yield a fluorophore signal corresponding to the dNTP, and the signal is captured by the CCD camera prior to each round of dNTP addition. The sequence read length ranges from 25 to 50 nucleotides, and the overall output per analysis run exceeds 1 billion nucleotide pairs.
[0089] Ion Torrent technology is a DNA sequencing method based on the detection of hydrogen ions released during DNA polymerization (see, e.g., Science 327(5970):1190 (2010); U.S. Patent Application Publication Nos. 20090026082, 20090127589, 20100301398, 20100197507, 20100188073, and 20100137143, which are incorporated by reference). The microwells contain template DNA strands to be sequenced. Below the microwell layer is a high-sensitivity ISFET ion sensor. All layers are contained within a CMOS semiconductor chip, similar to those used in the electronics industry. When a dNTP is incorporated into the growing complementary strand, a hydrogen ion is released, triggering the high-sensitivity ion sensor. If a homopolymer repeat is present in the template sequence, multiple dNTP molecules are incorporated in a single cycle. This results in a corresponding number of released hydrogens and a proportionally higher electronic signal. This technology differs from other sequencing technologies in that it does not use modified nucleotides or optics. The per-base accuracy of the Ion Torrent sequencer is approximately 99.6% for 50-base reads, and each run generates approximately 100 Mb to 100 Gb. The read length is 100 - 300 base pairs. The accuracy for homopolymer repeats of length 5 repeats is approximately 98%. The advantages of ion semiconductor sequencing are fast sequencing speed and low upfront and operating costs.
[0090] The mutant T7 DNA ligase can be used in another nucleic acid sequencing method developed by Stratos Genomics, Inc., and involves the use of Xpandomers. This sequencing process generally includes providing daughter strands produced by template-directed synthesis. The daughter strands typically comprise a plurality of subunits that are joined in a sequence corresponding to all or a portion of the contiguous nucleotide sequence of the target nucleic acid, wherein each subunit comprises a tether, at least one probe or nucleobase residue, and at least one selectively cleavable bond. The selectively cleavable bond is cleaved to produce an Xpandomer that is longer than the plurality of subunits of the daughter strand. The Xpandomer typically comprises a tether and a reporter element for resolving genetic information in the sequence corresponding to all or a portion of the contiguous nucleotide sequence of the target nucleic acid. The reporter element of the Xpandomer is then detected. Additional details regarding Xpandomer-based methods are described, for example, in U.S. Patent Publication No. 20090035777, which is incorporated herein by reference.
[0091] Other single molecule sequencing methods include real-time sequencing-by-synthesis using the VisiGen platform (Voelkerding et al., Clinical Chem., 55:641-58, 2009; U.S. Patent No. 7,329,492; U.S. Patent Application Serial No. 11 / 671,956; U.S. Patent Application Serial No. 11 / 781,166; each of which is incorporated herein by reference), in which strand extension of a fixed, primed DNA template is carried out using a fluorescence-modified polymerase and a fluorescence acceptor molecule to obtain detectable fluorescence resonance energy transfer (FRET) upon addition of nucleotides.
[0092] The specific methods and compositions described herein are representative of preferred embodiments and are exemplary and not intended to limit the scope of the invention. Given the present specification, other objects, aspects, and embodiments will occur to those skilled in the art and are included within the spirit of the invention as defined by the scope of the claims. It will be apparent to those skilled in the art that various substitutions and modifications can be made to the invention disclosed herein without departing from the scope and spirit of the invention. The invention as illustratively described herein can be practiced appropriately without any one or more elements, or any one or more limitations, not expressly disclosed herein as being necessary. Thus, for example, in each instance in the embodiments or examples of the invention herein, any of the terms “comprising,” “including,” “containing,” etc. should be read broadly and without limitation. The methods and processes illustratively described herein can be appropriately implemented in different orders of steps, and they need not be limited to the order of steps indicated herein or in the claims. It should also be noted that, unless the context clearly indicates otherwise, as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents, and the plural includes the singular form. In any case, this patent application should not be construed as limited to the specific examples or embodiments or methods specifically disclosed herein.
[0093] The invention has been described herein in broad and general terms. Each of the narrower species and subgeneric classifications falling within the overall disclosure text also forms part of the invention. The terms and expressions that have been employed are used as terms of description and not of limitation, and are not intended to exclude any equivalents or portions of the features shown and described, but it will be recognized that various modifications within the scope of the claimed invention are possible. Accordingly, it is to be understood that, although the invention has been specifically disclosed by way of preferred embodiments and optional features, modifications and variations of the concepts disclosed herein may be employed by those skilled in the art, including but not limited to variant sequences, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.
[0094] Related sequences
[0095] DNA sequence of wild-type T7 DNA ligase, SEQ ID NO: 1
[0096] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 1020 1021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0097] Amino acid sequence of wild-type T7 DNA ligase, SEQ ID NO: 2
[0098] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEADGIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMDEFTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0099] DNA sequence of T7 DNA ligase E63K, SEQ ID NO: 3
[0100] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTC AAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841 AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 10201021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0101] Amino acid sequence of T7 DNA ligase E63K, SEQ ID NO: 4
[0102] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALKHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEADGIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMDEFTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0103] DNA sequence of T7 DNA ligase K73E, SEQ ID NO: 5
[0104] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGG GAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841 AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 1020 1021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0105] Amino acid sequence of T7 DNA ligase K73E, SEQ ID NO: 6
[0106] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRW E RLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEADGIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMDEFTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0107] Amino acid sequence of T7 DNA ligase K137E, SEQ ID NO: 7
[0108] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP060061 ALKHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120121 EELFVEPIRKKDKVPFELHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240241 ENEADGIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMDEFTETVKEATLS 300301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0109] DNA sequence of T7 DNA ligase K174E, SEQ ID NO: 8
[0110] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGA GAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841 AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 1020 1021 CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0111] Amino acid sequence of T7 DNA ligase K174E, SEQ ID NO: 9
[0112] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHV ENMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEADGIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMDEFTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0113] DNA sequence of T7 DNA ligase E182K, SEQ ID NO: 10
[0114] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAA AAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841 AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 1020 1021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0115] Amino acid sequence of T7 DNA ligase E182K, SEQ ID NO: 11
[0116] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 Q KYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEADGIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMDEFTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0117] DNA sequence of T7 DNA ligase K210E, SEQ ID NO: 12
[0118] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACA GAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841 AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 1020 1021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0119] Amino acid sequence of T7 DNA ligase K210E, SEQ ID NO: 13
[0120] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQ ERAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEADGIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMDEFTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0121] DNA sequence of T7 DNA ligase E243K, SEQ ID NO: 14
[0122] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAAC AAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841 AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 1020 1021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0123] Amino acid sequence of T7 DNA ligase E243K, SEQ ID NO: 15
[0124] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 EN K ADGIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMDEFTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0125] DNA sequence of T7 DNA ligase D245R, SEQ ID NO: 16
[0126] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCG CGTGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841 AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 1020 1021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0127] Amino acid sequence of T7 DNA ligase D245R, SEQ ID NO: 17
[0128] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEA R GIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMDEFTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0129] DNA sequence of T7 DNA ligase E268K, SEQ ID NO: 18
[0130] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTT AAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841 AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 1020 1021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0131] Amino acid sequence of T7 DNA ligase E268K, SEQ ID NO: 19
[0132] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEADGIIQGLVWGTKGLANEGKVIGF K VLLESGRLVNATNISRALMDEFTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0133] DNA sequence of T7 DNA ligase E272K, SEQ ID NO: 20
[0134] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTG AAAAGCGGTCGCCTCGTCAATGCGACA 840 841 AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 1020 1021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0135] Amino acid sequence of T7 DNA ligase E272K, SEQ ID NO: 21
[0136] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEADGIIQGLVWGTKGLANEGKVIGFEVLL K SGRLVNATNISRALMDEFTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0137] DNA sequence of T7 DNA ligase E289K, SEQ ID NO: 22
[0138] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841 AACATCAGTCGGGCACTGATGGATAAG TTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 1020 1021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0139] Amino acid sequence of T7 DNA ligase E289K, SEQ ID NO: 23
[0140] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEADGIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMD K FTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0141] DNA sequence of T7 DNA ligase K295E, SEQ ID NO: 24
[0142] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661 AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTG GAA GAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCTGACGGTTCACTGCGC 1020 1021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0143] Amino acid sequence of T7 DNA ligase K295E, SEQ ID NO: 25
[0144] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEADGIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMDEFTETV E KATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETPDGSLRHPSFVMFRGTEDNPQEKM 358
[0145] DNA sequence of T7 DNA ligase D336R, SEQ ID NO: 26
[0146] 001 ATGAATATCAAGACTAATCCGTTTAAAGCAGTATCGTTCGTGGAAAGCGCGATCAAAAAA 060061 GCCTTGGACAACGCTGGTTATTTAATCGCAGAAATTAAATATGATGGTGTCAGAGGGAAC 120 121 ATCTGCGTCGATAATACGGCCAATTCGTATTGGCTGAGCCGTGTGTCTAAAACTATTCCG 180 181 GCACTCGAACACCTGAATGGTTTTGATGTTAGATGGAAACGCCTTTTAAATGACGATCGG 240 241 TGTTTTTACAAAGATGGCTTTATGCTGGATGGGGAACTGATGGTTAAAGGCGTCGATTTC 300 301 AATACCGGATCTGGGTTATTACGTACGAAATGGACTGACACAAAAAATCAAGAATTTCAC 360 361 GAAGAATTATTTGTAGAACCAATTCGAAAAAAGGATAAAGTGCCTTTTAAGTTACATACA 420 421 GGCCATCTGCATATAAAGCTCTATGCGATACTGCCCCTTCACATTGTGGAAAGCGGTGAG 480 481 GATTGTGACGTCATGACGCTGCTGATGCAGGAACATGTGAAAAACATGTTACCTCTGTTA 540 541 CAAGAATATTTTCCAGAGATTGAGTGGCAGGCCGCGGAATCCTATGAAGTTTATGACATG 600 601 GTAGAACTCCAGCAGTTGTATGAACAAAAACGCGCCGAAGGGCACGAAGGATTGATCGTC660 661AAAGATCCCATGTGTATCTATAAACGGGGTAAAAAATCCGGTTGGTGGAAAATGAAACCG 720 721GAAAACGAAGCGGATGGTATAATTCAGGGACTGGTGTGGGGTACGAAAGGATTAGCGAAT 780 781 GAAGGCAAAGTCATCGGTTTTGAAGTTTTGCTGGAAAGCGGTCGCCTCGTCAATGCGACA 840 841AACATCAGTCGGGCACTGATGGATGAGTTCACAGAGACCGTGAAAGAAGCGACCTTGTCT 900 901 CAGTGGGGCTTCTTTTCTCCTTACGGTATAGGAGATAATGATGCTTGTACTATTAATCCG 960 961 TATGACGGATGGGCATGTCAGATCAGTTACATGGAAGAAACTCCT CGT GGTTCACTGCGC 1020 1021CATCCCAGCTTCGTGATGTTCCGGGGTACTGAAGATAATCCCCAAGAGAAAATGTAA 1077
[0147] Amino acid sequence of T7 DNA ligase D336R, SEQ ID NO: 27
[0148] 001 MNIKTNPFKAVSFVESAIKKALDNAGYLIAEIKYDGVRGNICVDNTANSYWLSRVSKTIP 060061 ALEHLNGFDVRWKRLLNDDRCFYKDGFMLDGELMVKGVDFNTGSGLLRTKWTDTKNQEFH 120 121 EELFVEPIRKKDKVPFKLHTGHLHIKLYAILPLHIVESGEDCDVMTLLMQEHVKNMLPLL 180 181 QEYFPEIEWQAAESYEVYDMVELQQLYEQKRAEGHEGLIVKDPMCIYKRGKKSGWWKMKP 240 241 ENEADGIIQGLVWGTKGLANEGKVIGFEVLLESGRLVNATNISRALMDEFTETVKEATLS 300 301 QWGFFSPYGIGDNDACTINPYDGWACQISYMEETP R GSLRHPSFVMFRGTEDNPQEKM 358。
Claims
1. A mutant T7 DNA ligase or a bioactive fragment thereof, which comprises one or more of the following amino acid mutations, wherein said mutations are substitutions at the indicated positions in each amino acid sequence, and wherein the entire amino acid sequence is represented by adjacent sequence identification numbers: E63K (SEQ ID NO: 4), K73E (SEQ ID NO: 6), K137E (SEQ ID NO: 7), K174E (SEQ ID NO: 9), E182K (SEQ ID NO: 11), K210E (SEQ ID NO: 13), E243K (SEQ ID NO: 15), D245R (SEQ ID NO: 17), E268K (SEQ ID NO: 19), E272K (SEQ ID NO: 21), E289K (SEQ ID NO: 23), K295E (SEQ ID NO: 25), and D336R (SEQ ID NO: 27).
2. A polynucleotide that encodes the amino acid sequence of one of the mutant T7 DNA ligases according to claim 1.
3. A mutant T7 DNA ligase or a bioactive fragment thereof, which comprises one or more of the following amino acid mutations, wherein said mutations are substitutions at the indicated positions in each amino acid sequence, and wherein the entire amino acid sequence is represented by adjacent sequence identification numbers: E63K (SEQ ID NO: 4), K73E (SEQ ID NO: 6), K137E (SEQ ID NO: 7), K174E (SEQ ID NO: 9), E182K (SEQ ID NO: 11), K210E (SEQ ID NO: 13), E243K (SEQ ID NO: 15), D245R (SEQ ID NO: 17), E268K (SEQ ID NO: 19), E272K (SEQ ID NO: 21), E289K (SEQ ID NO: 23), K295E (SEQ ID NO: 25), and D336R (SEQ ID NO: 27), wherein each of said amino acid sequences has conservative substitutions for some of its amino acids, but only to the extent of maintaining at least 70% sequence identity with the sequence represented by said sequence identification number.
4. The mutant T7 DNA ligase according to claim 3, wherein each of said amino acid sequences has conservative substitutions for some of its amino acids, but only to the extent of maintaining at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with the sequence represented by said sequence identification number.
5. A polynucleotide that encodes the amino acid sequence of one of the mutant T7 DNA ligases according to claim 3.
6. A polynucleotide that encodes the amino acid sequence of one of the mutant T7 DNA ligases according to claim 4.
7. The polynucleotide according to claim 2, which has one of the following DNA sequences: E63K (SEQ ID NO: 3), K73E (SEQ ID NO: 5), K174E (SEQ ID NO: 8), E182K (SEQ ID NO: 10), K210E (SEQ ID NO: 12), E243K (SEQ ID NO: 14), D245R (SEQ ID NO: 16), E268K (SEQ ID NO: 18), E272K (SEQ ID NO: 20), E289K (SEQ ID NO: 22), K295E (SEQ ID NO: 24), and D336R (SEQ ID NO: 26).
8. A vector incorporating the polynucleotide according to claim 2.
9. A vector incorporating the polynucleotide according to claim 7.
10. A cell transformed with the polynucleotide according to claim 8 and expressing the polynucleotide.
11. A cell transformed with the polynucleotide according to claim 9 and expressing the polynucleotide.
12. A method for polynucleotide ligation between different polynucleotides or by ligating the 5'-end and 3'-end of a polynucleotide to produce a circular polynucleotide, wherein the polynucleotide has blunt ends or sticky ends, the method comprising: providing a ligation mixture comprising the polynucleotides to be ligated and the mutant T7 DNA ligase or bioactive fragment according to claim 1; and placing the ligation mixture at a temperature at which ligation occurs.
13. A method for polynucleotide ligation between different polynucleotides or by ligating the 5'-end and 3'-end of a polynucleotide to produce a circular polynucleotide, wherein the polynucleotide has blunt ends or sticky ends, the method comprising: providing a ligation reaction mixture comprising a buffer, the polynucleotides to be ligated, and the mutant T7 DNA ligase or bioactive fragment according to claim 1; and placing the ligation reaction mixture under temperature conditions suitable for ligation.
14. The method according to claim 13, wherein the ligation reaction mixture comprises Tris-HCl, MgCl2, ATP, dithiothreitol, and water.
15. A method for polynucleotide ligation between different polynucleotides or by ligating the 5'-end and 3'-end of a polynucleotide to produce a circular polynucleotide, wherein the polynucleotide has blunt ends or sticky ends, the method comprising: providing a ligation reaction mixture comprising a buffer, the polynucleotides to be ligated, and the mutant T7 DNA ligase or bioactive fragment according to claim 3; and placing the ligation reaction mixture under temperature conditions suitable for ligation.
16. The method according to claim 15, wherein the ligation reaction mixture comprises Tris-HCl, MgCl2, ATP, dithiothreitol, and water.
Citation Information
Patent Citations
Methods of amplifying and sequencing nucleic acids
US20050130173A1
Method and apparatus for moving stage detection of single molecular events
US20080241951A1
Methods and apparatus for measuring analytes using large scale FET arrays
US20090026082A1
High throughput nucleic acid sequencing by expansion
US20090035777A1
Methods and apparatus for measuring analytes using large scale FET arrays
US20090127589A1