Exonuclease activity-weakened nucleic acid polymerase mutant for single-molecule sequencing

By modifying the amino acid sequence of nucleic acid polymerase to weaken its 3'→5' exonuclease activity and optimize its extension progression, the problems of insufficient sequencing accuracy and read length in single-molecule sequencing technology have been solved, achieving higher sequencing performance and clinical application potential.

WO2025237410A1PCT designated stage Publication Date: 2025-11-20GENEUS TECH CHENGDU CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/095468
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-17
Filing Date
2025-05-16
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

In existing single-molecule sequencing technologies, the 3'→5' exonuclease activity of nucleic acid polymerases leads to base insertion errors in sequencing results, reducing sequencing accuracy and read length, making it difficult to meet the high requirements of clinical applications.

Method used

By modifying the amino acid sequence of wild-type polymerase to reduce its 3'→5' exonuclease activity and coupling it to protein nanopore molecules via the SpyTag-SpyCatcher method, the extension progression and thermal stability of the polymerase are optimized, thereby improving sequencing accuracy and read length.

Benefits of technology

It achieved higher sequencing accuracy and read length, meeting the high requirements of clinical applications, improving the performance of single-molecule sequencing technology, and promoting its application in the field of precision medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025095468_20112025_PF_FP_ABST
    Figure CN2025095468_20112025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a polymerase mutant applicable to single-molecule sequencing, which has weakened 3'→5' exonuclease activity relative to a wild-type polymerase, while substantially maintaining or even improving the extension processivity, sequencing accuracy, and read length. The present invention also provides use of the polymerase mutant in single-molecule sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

Exo minus nucleic acid polymerase mutants for single molecule sequencing

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese Patent Application No. CN202410619809.6, filed May 17, 2024, which is incorporated by reference herein in its entirety. TECHNICAL FIELD

[0003] The present invention belongs to the field of single molecule sequencing and the necessary nucleic acid polymerase modification of the “sequencing by synthesis” scheme. Specifically, the present invention relates to a polymerase mutant suitable for single molecule sequencing technology, and a polymerase mutant with overall improved performance is obtained. BACKGROUND

[0004] In recent years, single molecule sequencing technology has begun to emerge and show great application prospects. Single molecule sequencing refers to obtaining a nucleic acid base sequencing signal that can be identified based on only one nucleic acid molecule without amplification, achieving the purpose of sequencing. Compared with traditional sequencing technology based on multiple copies of nucleic acid molecules, 1. Single molecule sequencing does not require amplification, avoiding the “amplification bias” problem of traditional sequencing technology, providing more uniform and complete coverage for complex sequencing samples, and improving the detection sensitivity of low-abundance samples; 2. Single molecule sequencing avoids the problem of asynchronous multiple molecule signals, thereby greatly extending the sequence length that can be obtained each time. The leap in read length has a significant advantage for subsequent sequence assembly and information acquisition of genetic structure variation, allowing for nearly perfect interpretation of complex genomes and metagenomes, making it possible for sequencing technology to be widely used in clinical applications; 3. The principle of single molecule sequencing technology supports simultaneous biochemical reactions and signal acquisition during sequencing, without the need to pause. This “real-time” feature makes the sequencing process more convenient and efficient. In addition, single molecule sequencing can be combined with nanopore detection without the need for fluorescent labeling, which also lays a good foundation for a significant reduction in sequencing costs.

[0005] Single-molecule sequencing represents the future direction of sequencing technology due to its great advantages / potential in data quality, detection cost, convenience, timeliness, etc. It is of great significance for sequencing technology to enter the clinical application on a large scale, promote the development of precision medicine, and promote the boom of related markets. Current single-molecule sequencing technologies in the sequencing field are mainly based on two schemes: one is based on directly reading nucleic acid chain bases using nanopores, represented by Oxford Nanopore; the other is based on "sequencing by synthesis", which monitors the type of nucleotide used in each step of the nucleic acid polymerase replication process to achieve sequencing purposes, represented by Pacific Biosciences (based on fluorescence signal) and Roche sequencing department (based on nanopore current signal). Compared with the scheme of directly reading nucleic acid bases, the single-molecule sequencing technology of "sequencing by synthesis" is based on traditional and mature biochemical methods, and the generated sequencing signal is simpler and clearer. The number of bases and the number of signals have a direct correspondence (i.e. each base corresponds to a piece of signal, and the starting boundary of the signal is clear), which is more conducive to the detection of homopolymer in complex genomes, so the final sequencing accuracy is high, and it can better meet the high requirements of clinical applications. With the maturation of the technical scheme, the industry generally expects that the single-molecule "sequencing by synthesis" technology will soon realize the boom of the clinical application market in the next few years, which will not only realize huge market benefits, but also promote the entire medical industry into a new era of precision medicine.

[0006] For the "sequencing by synthesis" scheme, the nucleic acid polymerase drives the entire biochemical process of sequencing, and its importance is self-evident. The comprehensive performance of the polymerase is required to be higher for the implementation of this scheme on a single-molecule basis. These requirements include: 1. can efficiently extend the modified nucleotide; 2. the catalytic activity of the polymerase exhibits good kinetic characteristics matched with the signal detection method; 3. good extension progressiveness, i.e. the number of bases that a polymerase molecule can continuously extend before releasing the template molecule; 4. good thermal stability; 5. good strand displacement function; 6. other properties of the polymerase, including compatible working temperature, salt concentration with the sequencing platform environment, good fidelity, and low level of exonuclease activity. Among them, the kinetic characteristics and extension progressiveness of the polymerase are the key factors for determining the accuracy and read length of single-molecule sequencing technology, and are important contents of polymerase optimization and modification.

[0007] Based on good extension modification nucleotide ability and strand substitution function, the wild type polymerase with sequence of SEQ ID NO: 1 is identified as the preferred enzyme for our single molecule sequencing platform. Its gene fragment is derived from a bacteriophage Actinomyces naeslundii Phage Av-1, which is identified as a bacteriophage DNA polymerase in the literature. The SEQ ID NO: 1 polymerase has very strong 3'→5' exonuclease activity, and has 3'→5' exonuclease proofreading function for newly synthesized mismatched bases during the process of sequencing by synthesis, causing the error type of real-time generated sequencing data to change from mismatch to base insertion. Base insertion is also a type of mismatch, but due to the insertion of error bases, the alignment position of the following base pairs changes, and in sequence alignment, insertion errors have a higher penalty than mismatch errors, which will reduce the correction accuracy of the consensus sequence. Therefore, weakening the exonuclease activity of the polymerase is an important way to improve the accuracy of the sequencing by synthesis technology. In addition, many articles report that mutants with weakened exonuclease activity exhibit a characteristic of weakened extension progression, and therefore, optimizing and modifying the SEQ ID NO: 1 wild type polymerase to obtain a mutant with weakened exonuclease activity and good extension progression is an important content for improving the accuracy and read length of single molecule sequencing technology and promoting the application of single molecule sequencing technology. SUMMARY

[0008] In one aspect, provided herein are polymerase mutants comprising any one or any combination of the following mutations relative to a wild type polymerase:

[0009] S24G, S24C, S24V, C26A, C26V, D28A, T31I, T33G, T33S, T33Y, F81V, F81I, L87M, N78D, D82A, D88Y, S119P, K137A, K137H, K137P, K137S, S141M, I146F, I146L, N151C, N151L, N151T, P164G, P164M, P164T, Y178F, Q179T, D182L, D182Q, Q193W, Q193A, T243A, T243S, Y244T, K401M, K401P, T404M, T404E, I502L, D516K, D516Y, and L545N,

[0010] wherein the wild type polymerase comprises the amino acid sequence set forth in SEQ ID NO: 1.

[0011] In some embodiments, the polymerase mutant has weakened 3'→5' exonuclease activity relative to the wild type polymerase.

[0012] In some embodiments, the 3' 5' exonuclease activity of the polymerase mutant is 5% - 79% or 5% - 93% of the wild-type polymerase.

[0013] In some embodiments, the polymerase mutant comprises any one or any combination of the following mutations relative to the wild-type polymerase:

[0014] S24G, S24C, S24V, C26A, C26V, T33S, T33Y, L87M, D88Y, S119P, K137A, K137S, S141M, I146F, I146L, N151C, N151L, N151T, P164G, P164M, P164T, Y178F, Q179T, D182Q, Q193W, Q193A, T243A, T243S, Y244T, K401M, K401P, T404M, T404E, and L545N.

[0015] In some embodiments, the polymerase mutant has attenuated 3' 5' exonuclease activity relative to the wild-type polymerase, and processivity of extension with modified nucleotides is 110% - 207% or 104% - 244% of the wild-type polymerase.

[0016] In some embodiments, the modified nucleotide is a dN4P or dN6P nucleotide.

[0017] In some embodiments, the polymerase mutant comprises any one of the following combinations of mutations relative to the wild-type polymerase:

[0018] K116R + Y372F + L318N + H394I + Q313N + S24G;

[0019] K116R + Y372F + L318N + H394I + Q313N + S24C;

[0020] K116R + Y372F + L318N + H394I + Q313N + C26V;

[0021] K116R + Y372F + L318N + H394I + Q313N + T33S;

[0022] K116R + Y372F + L318N + H394I + Q313N + T33Y;

[0023] K116R + Y372F + L318N + H394I + Q313N + D88Y;

[0024] K116R+Y372F+L318N+H394I+Q313N+K401M;

[0025] K116R+Y372F+L318N+H394I+Q313N+Q179T;

[0026] K116R+Y372F+L318N+H394I+Q313N+K401M;

[0027] K116R+Y372F+L318N+H394I+Q313N+K137A;

[0028] K116R+Y372F+L318N+H394I+Q313N+S141M;

[0029] K116R+Y372F+L318N+H394I+Q313N+I146L;

[0030] K116R+Y372F+L318N+H394I+Q313N+Y178F;

[0031] K116R+Y372F+L318N+H394I+Q313N+Q193A;

[0032] K116R+Y372F+L318N+H394I+Q313N+T243S;

[0033] K116R+Y372F+L318N+H394I+Q313N+T404M;

[0034] K116R+Y372F+L318N+H394I+Q313N+N151L;

[0035] K116R+Y372F+L318N+H394I+Q313N+L545N;

[0036] L265A+E378Q+T384K+K491R+E493K+T33G;

[0037] L265A+E378Q+T384K+K491R+E493K+T33Y;

[0038] L265A+E378Q+T384K+K491R+E493K+D88Y;

[0039] L265A+E378Q+T384K+K491R+E493K+K137S;

[0040] L265A + E378Q + T384K + K491R + E493K + S24C;

[0041] L265A + E378Q + T384K + K491R + E493K + P164M;

[0042] L265A + E378Q + T384K + K491R + E493K + Y178F;

[0043] L265A + E378Q + T384K + K491R + E493K + Y244T; and

[0044] L265A + E378Q + T384K + K491R + E493K + L545N.

[0045] In some embodiments, the polymerase mutant has higher sequencing accuracy and read length relative to the wild-type polymerase.

[0046] In some embodiments, the polymerase mutant has a median sequencing accuracy higher than 67.5% and / or a median sequencing read length higher than 6350 bp.

[0047] In some embodiments, the polymerase mutant is coupled to a protein nanopore molecule.

[0048] In some embodiments, the polymerase mutant is coupled to the protein nanopore molecule via SpyTag-SpyCatcher approach.

[0049] In some embodiments, the protein nanopore molecule is a-hemolysin heptamer.

[0050] In another aspect, provided herein is use of the above polymerase mutant in nucleic acid sequencing.

[0051] In some embodiments, the nucleic acid sequencing is single molecule sequencing.

[0052] In some embodiments, the nucleic acid sequencing is sequencing by synthesis approach. BRIEF DESCRIPTION OF DRAWINGS

[0053] FIG. 1 is a schematic diagram of dN6P and dA4P-FG structure.

[0054] FIG. 2 shows the LL35 template structure diagram and sequence information.

[0055] FIG. 3 shows representative results of 3'→5' exonuclease activity of wild-type polymerase (L001) and some mutants.

[0056] FIG. 4 shows the RCA88 template structure diagram and sequence information.

[0057] Figure 5 shows representative results of the progressive evaluation of some mutants by fluorescence signal intensity at 3 hours.

[0058] Figure 6 shows representative results of the 3'→5' exonuclease activity of some mutants.

[0059] Figures 7A-7E show representative results of the progressive evaluation of some mutants by fluorescence signal intensity over time.

[0060] Figure 8 shows representative results of the progressive evaluation of some mutants by fluorescence signal intensity at 3 hours. DETAILED DESCRIPTION

[0061] Unless otherwise indicated, all technical and scientific terms used herein have the meanings that are commonly understood by one of ordinary skill in the art.

[0062] The term "or" refers to a single element of the list of alternatives, unless otherwise indicated by context. The term "and / or" means any one, any two, any three, any more, or all of the listed alternatives.

[0063] The terms "comprising," "including," "having," and the like, are meant to be interpreted as "including but not limited to," unless otherwise indicated by context.

[0064] The term "about" means up to plus or minus 1%, more specifically 5%, more specifically 10%, more specifically 15%, and in some cases up to or down to 20% of the stated value, with the range of deviation including integer values and, where applicable, non-integer values, constituting a continuous range.

[0065] The term "wild-type polymerase" or "polymerase wild-type" refers herein to a DNA polymerase from the bacteriophage Actinomyces naeslundii Phage Av-1 having the amino acid sequence set forth in SEQ ID NO: 1. The modifier "wild-type" as used herein does not mean that it is naturally occurring in nature, but rather that the wild-type polymerase is used herein as a reference for determining the mutations and mutation positions included in the polymerase mutants provided herein, or can be considered as the parent protein of the polymerase mutants provided herein. In addition, in some cases, the wild-type polymerase is also used as a reference for studying various activities or properties of the polymerase mutants provided herein.

[0066] The term "polymerase mutant" or "polymerase mutant strain" refers herein to a protein that differs in amino acid sequence from a wild-type polymerase. The difference can be in the presence of one or more amino acid substitutions, deletions, or insertions at one or more positions in the amino acid sequence.

[0067] The term "amino acid substitution", also referred to as "amino acid replacement", refers herein to the substitution of one amino acid (referred to as the original amino acid) at a particular position in an amino acid sequence with another amino acid (the substitution amino acid). For example, if the amino acid at position 2 of a wild-type protein is a glutamic acid residue (E), and the corresponding position in a mutant protein is an alanine residue (A), then the mutant protein is said to have an amino acid substitution: alanine for glutamic acid at position 2. For amino acid substitutions, the following nomenclature is used herein: original amino acid, position, and substitution amino acid, and the IUPAC single letter abbreviations for amino acid names are used. For the example described above, this can be denoted as E2A. When two or more amino acid substitutions are included in a mutant protein, then the amino acid substitution combination is said to be present. The term "amino acid substitution combination" refers herein to the presence of amino acid substitutions at two or more positions in a mutant protein. For example, if the amino acid at position 2 of a wild-type protein is a glutamic acid residue (E), and the corresponding position in a mutant protein is an asparagine residue (N); and, simultaneously, the amino acid at position 23 of a wild-type protein is a glutamic acid residue (E), and the corresponding position in a mutant protein is an alanine residue (A), then the mutant protein is said to have an amino acid substitution combination: E2N and E23A. When multiple amino acid substitutions are present in a mutant, the multiple mutations can be separated by a "+" sign, e.g., "E2N + E23A + Y65R" represents the substitution of glutamic acid with asparagine at position 2, glutamic acid with alanine at position 23, and tyrosine with arginine at position 65.

[0068] The term "amino acid sequence" is synonymous with the terms "polypeptide", "protein", and "peptide" and is used interchangeably herein. Amino acid residues in an amino acid sequence can be represented using conventional one-letter or three-letter amino acid designations, wherein the amino acid sequence is presented in the standard amino to carboxy terminal orientation (i.e., N→C).

[0069] The term "3' to 5' exonuclease activity" refers to the ability of a DNA polymerase to cleave deoxyribonucleotide monomers from the 3' end of a nucleotide chain in the 5' direction. This activity of a DNA polymerase can be used to cleave off mispaired deoxyribonucleotides at the 3' end of a newly synthesized strand, followed by continued addition of correctly paired deoxyribonucleotides using polymerase activity, i.e., to provide a proofreading function to the DNA polymerase in synthesizing a new nucleotide chain. This function is important for DNA replication in an organism, but for single molecule sequencing applications, the proofreading function causes base insertion errors in the sequencing results, and thus the 3' to 5' exonuclease activity of a DNA polymerase used for single molecule sequencing should be reduced or attenuated. As used herein, "attenuated 3' to 5' exonuclease activity" refers to a 3' to 5' exonuclease activity that is reduced by about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or more, or even eliminated, relative to the 3' to 5' exonuclease activity of a wild-type DNA polymerase.

[0070] The term "elongation processivity" as used in reference to a DNA polymerase refers to the number of nucleotides that a DNA polymerase molecule can continuously extend before releasing or falling off of a template molecule. In some embodiments, the polymerase mutants provided herein have an elongation rate that is not less than about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or even better than the elongation processivity of a wild-type polymerase, e.g., more than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%.

[0071] The term "SpyTag-SpyCatcher" refers to a method of coupling two proteins of interest by means of a polypeptide spytag and a polypeptide spycachter that can spontaneously form a covalent linkage under mild conditions. For example, a polypeptide spytag can be fused to the C-terminus of a nanopore protein, while a polypeptide spycathcer is expressed in fusion with a polymerase. Because spytag and spycachter can spontaneously form a covalent linkage, the above system can help to achieve coupling of an independently expressed and purified nanopore protein with a nucleic acid polymerase.

[0072] The term "modified nucleotide" refers to a nucleotide that is modified to be suitable for single molecule sequencing. Examples of modified nucleotides include the 4-phosphate nucleotides (dN4P nucleotides) or 6-phosphate nucleotides (dN6P nucleotides) used in the examples herein. In addition, the modified nucleotides can further be linked to a fluorescent group, such as FG.

[0073] The present application provides a series of DNA polymerase mutants with high amplification performance and weakened 3'-5' exonuclease activity. Specifically, the present application constructs a single-point saturated mutation library for the exonuclease activity domain of the wild-type polymerase of SEQ ID NO: 1, and finally obtains a series of amino acid sites and mutations with weakened 3'-5' exonuclease activity by designing an efficient and accurate exonuclease activity screening scheme. Through characterization and identification, the mutant exhibits good comprehensive performance in amplification performance, extension performance, electrolyte tolerance, and catalytic kinetics of modified nucleotide substrates, and shows higher sequencing accuracy and sequencing read length in single molecule sequencing technology based on the "sequencing by synthesis" scheme.

[0074] The present application is further illustrated by specific examples below.

[0075] Example 1: Cloning, expression and purification of DNA polymerase and mutants

[0076] Vector construction: The SEQ ID NO: 1 DNA sequence was synthesized by Shenguo Bioengineering (Shanghai) Co., Ltd. and cloned into the pET26b vector (Novagen). In order to test the performance of the DNA polymerase on the chip, we fused a SpyCatcher protein to the DNA polymerase. In order to improve the performance of the DNA polymerase, the present application studies the mutation of multiple amino acid sites. The construction of the mutation is by designing primers containing the corresponding mutation sites (Shenguo Bioengineering) and introducing the mutation site into the gene by PCR (phusion polymerase kit, New England Biolabs). The PCR product is purified by the Tiagene (QIAquick Gel Extraction Kit) gel recovery kit, and recombined into the vector according to the same sequence on the primer and the vector. The plasmid is transformed into DH5a E. coli competent cells by chemical transformation. After sequencing (Shenguo Bioengineering) verification, the plasmid is extracted and stored at -20℃, and the strain is stored at -80℃.

[0077] Protein expression: The plasmid is transformed into BL21(DE3) E. coli competent cells by chemical transformation. A single colony on the plate is picked into 4 mL of LB medium containing kanamycin, and cultured at 37℃; 220 RPM overnight. The seed liquid is inoculated into new resistant medium in a 1:100 manner (250 mL, 1000 mL), and the inducer (IPTG = 250 μΜ) is added when the OD value is (1.5), and expressed at 16℃; 220 RPM overnight. 10-50 μL of bacterial liquid is centrifuged to remove the supernatant, and 1 x bugbuster (EMD Millipore) is added to the bacterial pellet to extract the protein. The protein is purified by Ni-NTA affinity chromatography (EMD Millipore) and stored at -20℃. 10X Protein Extraction Reagent, Milipore) to break the cells, and a precast gel (Precast-GL gel Hepes SDS-PAGE 10% 10 Well) was used to detect the expression of the protein.

[0078] Protein purification: The cells were collected by centrifugation, and the bacterial pellet was resuspended in lysis buffer (50 mM Tris-HCl (pH 8.0), 1 M NaCl, 10% glycerol, 0.1% Tween, 2 mM TCEP) and sonicated on ice for 20 min (power 60%, ultrasound 3 s, interval 6 s). The lysis product of the bacterial solution was collected by centrifugation, and the target protein was purified by NI column affinity chromatography and further purified by size exclusion chromatography (Superdex 200 Increase 10 / 300 GL). The final target protein was stored in buffer (50 mM Tris 7.5, 100 mM KCl, 0.1 mM EDTA, 1 mM TECP, 0.5% Tween-20, 0.5% NP-40, 50% glycerol) and stored at -80°C.

[0079] Example 2: Detection of 3'→5' exonuclease activity of polymerase mutants

[0080] Table 1. Reagent composition table for extension rate test experiment

[0081] Template preparation: The structure of the LL35 template is shown in FIG. 2. The template strand was heat treated in a metal bath at 95°C for 5 min and naturally cooled to 25°C. The complex obtained by annealing was named LL35.

[0082] 3'→5' exonuclease activity detection: Referring to the formula in Table 1, the reaction was carried out in a 96-well enzyme plate (Black 96 well cell culture plate, Flat Bottom, WHB): first mix LL35 and wild-type polymerase or its mutants, and incubate on ice for 10 min. Add Tris, KCl, ethylene glycol, dA4P-FG / dG6P / dC6P / dT6P gradually, and add Mn 2+The detection was carried out under the condition of 480 nm excitation light and 516 nm emission light. The 3' end base of the template LL35 is C, which forms a wrong pairing with the corresponding base T of the template. The 3' mismatched base C can be removed by the 3'→5' exonuclease function of the polymerase, and then the polymerase switches from the exonuclease activity to the extension activity, forms base complementary pairing with T by using DA4P-FG, and extends the synthesized chain. In this process, the by-product P-P-P-FG of the extension reaction is digested by alkaline phosphatase to release the fluorescent substance FG. Since only the base A among the four modified nucleotides (A / T / C / G) added in the reaction system has a fluorescent modification, the extension of LL35 without exonuclease activity only adds three modified bases T / C / G, and none of the three bases has a fluorescent modification, so the by-product does not produce FG fluorescence. Therefore, in this embodiment, the strength of the 3'→5' exonuclease activity of the polymerase can be quantified by the increase of the fluorescence value (516 nm) within the reaction time.

[0083] The test results show that the DNA polymerase SEQ ID NO: 1 and its mutants have 3'→5' exonuclease activity in the presence of modified nucleotides dA4P-FG / dG6P / dC6P / dT6P. After structural modification, the DNA polymerase mutants including at least one of the following mutations weaken the 3'→5' exonuclease activity to about 5% to 93% of the 3'→5' exonuclease activity of the wild type (SEQ ID NO: 1): S24G, S24C, S24V, C26A, C26V, D28A, T31I, T33G, T33S, T33Y, F81V, F81I, L87M, N78D, D82A, D88Y, S119P, K137A, K137H, K137P, K137S, S141M, I146F, I146L, N151C, N151L, N151T, P164G, P164M, P164T, Y178F, Q179T, D182L, D182Q, Q193W, Q193A, T243A, T243S, Y244T, K401M, K401P, T404M, T404E, I502L, D516K, D516Y, L545N (see FIGS. 3 and 6).

[0084] Example 3: Progressive evaluation of polymerase mutants with weakened exonuclease activity

[0085] Table 2. Reagent composition table of the progressive evaluation experiment

[0086] Template preparation: RCA88 structure see Figure 4. Template chain T2 (5'-P) mixed with primer chain P2 at a molar ratio of 1:1.2, metal bath 95℃ heat treatment for 5min, natural cooling to 25℃; add T4 ligase and buffer, 10℃ overnight ligation; 65℃ heat treatment for 30min, centrifugation at 12000g for 5min to collect the supernatant; use molecular sieve chromatography to collect the target template, named RCA88, and quantitate the concentration of RCA88T according to A260 (Nanodorp).

[0087] According to the formula in Table 2, the reaction was carried out in a 96-well enzyme plate (Black 96 well cell culture plate, Flat Bottom, WHB): first mix RCA88 and polymerase, incubate on ice for 30min. Add Tris, KCl, ethylene glycol, dA4P-FG / dG6P / dC6P / dT6P, interfering DNA, alkaline phosphatase and alkaline phosphatase buffer step by step, extend at 30℃ for 3h, add Mn 2+ , and detect at 480nm excitation light and 516nm emission light.

[0088] Good extension progress, i.e. the number of bases that can be continuously extended by one polymerase molecule before the template molecule is released. In a single molecule sequencing scheme, one polymerase corresponds to one nucleic acid template molecule to obtain sequence signals. Once the polymerase releases the template molecule, it will not be able to start sequencing the same template molecule again in a single molecule setting, thereby seriously affecting read length and data output.

[0089] In this embodiment, the polymerase synthesizes a new chain using dA4P-FG / dG6P / dC6P / dT6P, and the number of bases that can be continuously extended by the polymerase is equal to the sum of the number of dA4P-FG / dG6P / dC6P / dT6P consumed by the four modified bases, wherein in the consumption process, the byproduct P-P-P-FG is digested by alkaline phosphatase to release the fluorescent substance FG, so the consumption of dA4P-FG can be quantitated by the increase in fluorescence value (516nm) within the reaction time. If the polymerase dissociates from the template DNA (RCA88) in the reaction system, 100-fold excess of interfering DNA in the system will quickly occupy the enzyme binding site, preventing the recombination of template RCA88 and DNA polymerase, and at the same time, the interfering DNA is single-stranded DNA and cannot be extended because there is no complementary primer strand. Therefore, in this embodiment, the length of the progress of the polymerase mutant is judged by detecting the high and low of the fluorescence value at 516nm in the reaction system.

[0090] The test results show that the exonuclease activity-reduced mutants, including at least one mutation selected from the group consisting of S24G, S24C, S24V, C26A, C26V, T33S, T33Y, L87M, D88Y, S119P, K137A, K137S, S141M, I146F, I146L, N151C, N151L, N151T, P164G, P164M, P164T, Y178F, Q179T, D182Q, Q193W, Q193A, T243A, T243S, Y244T, K401M, K401P, T404M, T404E, L545N (see FIG. 5 and FIGS. 7A-7E and FIG. 8) have a progressive activity of about 102% to 244% of the wild-type SEQ ID NO: 1 compared to SEQ ID NO: 1.

[0091] Example 4: Performance evaluation of exonuclease activity-reduced polymerase mutants in sequencing applications

[0092] The DNA polymerase and its mutants are coupled to the protein nanopore by means of fusion proteins. The alpha-hemolysin pore heptamer fused with SpyTag is mixed with the DNA polymerase or its mutants fused with SpyCatcher at a molar ratio of 1:2, and placed at room temperature for 30 min. The coupling efficiency is detected by electrophoresis.

[0093] After coupling, the complex (template-polymerase-alpha hemolysin) is embedded in the lipid bilayer above the etched sequencing unit pores on the semiconductor sensor chip. The dynamic capture of modified nucleotides by the enzyme-template-nanopore complex is recorded at 120 mV / -180 mV. According to the current block caused by the modified nucleotides when captured by the DNA polymerase, the information such as the type of modified nucleotides, block frequency, residence time, etc. is determined, compared with the template sequence, and thus the kinetic characteristics of the DAN polymerase are evaluated.

[0094] The test results show that the DNA polymerase comprising at least one mutation selected from the group consisting of S24G, S24C, S24V, C26A, C26V, T33S, T33G, T33Y, L87M, D88Y, S119P, K137A, K137S, S141M, I146F, I146L, N151C, N151L, N151T, P164G, P164M, P164T, Y178F, Q179T, D182Q, Q193W, Q193A, T243A, T243S, Y244T, K401M, K401P, T404M, T404E, L545N improves the accuracy and read length. Specifically, the polymerase L014 (K116R+Y372F+L318N+H394I+Q313N+Y178F) obtains a sequence result of read length >8000 bp base signal in sequencing, and the accuracy of the sequencing result is 87.8% compared with the DNA sequence to be tested. Table 3 shows the characteristic data of some representative mutants in sequencing applications.

[0095] Table 3 Characteristic data of representative mutants in sequencing applications

[0096] The present application obtains a polymerase mutant with weakened exonuclease activity by deep optimization and modification of nucleic acid polymerase, further improves the sequencing accuracy and read length for the single molecule sequencing scheme based on the combination of nanopore and sequencing by synthesis, obtains high-quality sequencing data, and has industrialization prospects.

[0097] The polymerase mutant sequence mentioned herein is as follows:

Claims

1. A polymerase mutant comprising any one or any combination of the following mutations relative to a wild-type polymerase: S24G, S24C, S24V, C26A, C26V, D28A, T31I, T33G, T33S, T33Y, F81V, F81I, L87M, N78D, D82A, D88Y, S119P, K137A, K137H, K137P, K137S, S141M, I146F, I146L, N151C, N151L, N151T, P164G, P164M, P164T, Y178F, Q179T, D182L, D182Q, Q193W, Q193A, T243A, T243S, Y244T, K401M, K401P, T404M, T404E, I502L, D516K, D516Y, and L545N, wherein the wild-type polymerase comprises the amino acid sequence set forth in SEQ ID NO:

1.

2. The polymerase mutant of claim 1, having attenuated 3'→5' exonuclease activity relative to the wild-type polymerase.

3. The polymerase mutant of claim 1 or 2, having 5%-79% or 5%-93% of the 3'→5' exonuclease activity of the wild-type polymerase.

4. The polymerase mutant of any one of claims 1-3, comprising any one or any combination of the following mutations relative to the wild-type polymerase: S24G, S24C, S24V, C26A, C26V, T33S, T33Y, L87M, D88Y, S119P, K137A, K137S, S141M, I146F, I146L, N151C, N151L, N151T, P164G, P164M, P164T, Y178F, Q179T, D182Q, Q193W, Q193A, T243A, T243S, Y244T, K401M, K401P, T404M, T404E, and L545N.

5. The polymerase mutant of any one of claims 1-4, having attenuated 3'→5' exonuclease activity relative to the wild-type polymerase, and a processivity of extension with modified nucleotides that is 110%-207% or 104%-244% of the wild-type polymerase.

6. The polymerase mutant of any one of claims 1-5, wherein the modified nucleotide is a dN4P or dN6P nucleotide.

7. The polymerase mutant of any one of claims 1-6, comprising any one of the following combinations of mutations relative to the wild-type polymerase: K116R + Y372F + L318N + H394I + Q313N + S24G; K116R + Y372F + L318N + H394I + Q313N + S24C; K116R + Y372F + L318N + H394I + Q313N + C26V; K116R+Y372F+L318N+H394I+Q313N+T33S; K116R+Y372F+L318N+H394I+Q313N+T33Y; K116R+Y372F+L318N+H394I+Q313N+D88Y; K116R+Y372F+L318N+H394I+Q313N+S119P; K116R+Y372F+L318N+H394I+Q313N+Q179T; K116R+Y372F+L318N+H394I+Q313N+K401M; K116R+Y372F+L318N+H394I+Q313N+K137A; K116R+Y372F+L318N+H394I+Q313N+S141M; K116R+Y372F+L318N+H394I+Q313N+I146L; K116R+Y372F+L318N+H394I+Q313N+Y178F; K116R+Y372F+L318N+H394I+Q313N+Q193A; K116R+Y372F+L318N+H394I+Q313N+T243S; K116R+Y372F+L318N+H394I+Q313N+T404M; K116R+Y372F+L318N+H394I+Q313N+N151L; K116R+Y372F+L318N+H394I+Q313N+L545N; L265A+E378Q+T384K+K491R+E493K+T33G; L265A+E378Q+T384K+K491R+E493K+T33Y; L265A+E378Q+T384K+K491R+E493K+D88Y; L265A+E378Q+T384K+K491R+E493K+K137S; L265A+E378Q+T384K+K491R+E493K+S24C; L265A+E378Q+T384K+K491R+E493K+P164M; L265A+E378Q+T384K+K491R+E493K+Y178F; L265A+E378Q+T384K+K491R+E493K+Y244T; and L265A+E378Q+T384K+K491R+E493K+L545N.

8. The polymerase mutant of any one of claims 1-7, which has higher sequencing accuracy and read length relative to the wild-type polymerase.

9. The polymerase mutant of any one of claims 1-8, which has a median sequencing accuracy higher than 67.5% and / or a median sequencing read length higher than 6350 bp.

10. The polymerase mutant of any one of claims 1-9, coupled to a protein nanopore molecule.

11. The polymerase mutant of any one of claims 1-10, coupled to the protein nanopore molecule by SpyTag-SpyCatcher approach.

12. The polymerase mutant of any one of claims 1-11, wherein the protein nanopore molecule is an a-hemolysin heptamer.

13. Use of the polymerase mutant of any one of claims 1-12 in nucleic acid sequencing.

14. The use of claim 13, wherein the nucleic acid sequencing is single molecule sequencing.

15. The use of claim 13 or 14, wherein the nucleic acid sequencing is by sequencing by synthesis approach.

Citation Information

Patent Citations

  • Polymerase enzyme

    CN107922929A

  • Labeled nucleotide analogs, reaction mixtures, and methods and systems for sequencing

    CN108697499A

  • Phi29 DNA polymerase and encoding gene and application thereof

    CN110573616A

  • Novel DNA polymerase variant

    JP2021023199A

  • Recombinant Polymerases With Increased Phototolerance

    US20130217007A1