Polymerases for mixed aqueous-organic media and uses thereof - Patents.com

Modified Taq DNA polymerases with specific amino acid changes and organic solvents enhance PCR amplification of high-GC content DNA by improving thermal stability and fidelity, addressing denaturation and degradation issues in existing PCR methods.

JP2024538743A5Inactive Publication Date: 2026-04-085PRIME BIOSCIENCES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-01-04
Publication Date
2026-04-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing PCR methods face challenges in amplifying high-GC content DNA regions due to incomplete denaturation, formation of secondary structures, and rapid polymerase deactivation, leading to low yield and poor fidelity, especially in the presence of low-molecular-weight organic solvents.

Method used

Development of modified Taq DNA polymerases with specific amino acid changes and the use of PCR buffers containing low molecular weight organic solvents like amides, sulfoxides, and diols, enhancing polymerase stability and compatibility in organic aqueous media.

Benefits of technology

The modified polymerases demonstrate increased thermal stability and fidelity, enabling effective amplification of high-GC content DNA regions with improved yield and reduced degradation, even at higher temperatures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to the field of molecular biology known as protein engineering, which is concerned with designing enzymes with properties that exceed those of previously reported enzymes. More particularly, the present invention relates to compositions comprising engineered polymerase enzymes with various properties that exceed those of previously reported polymerase enzymes, and compositions using such enzymes for performing polynucleotide amplification reactions in organic-aqueous media.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is generally related to molecular biology and methods of molecular biology for selecting nucleic acids encoding gene products. More particularly, the present invention relates to compositions and methods for enhancing polynucleotide amplification reactions in organic aqueous media.

Background Art

[0002] Polymerase chain reaction (PCR), an in vitro method for amplifying DNA sequences, is a core technique in modern biology. This technique was first discovered in 1985 by the group of Kary Mullis (Saiki et al., 1985, 1986). In summary, the process consists of two steps: selecting a region of target DNA to be amplified, and aligning it with two oligonucleotide primers (each extended from its 3' end by the DNA polymerase enzyme). A typical PCR reaction includes target DNA, two oligonucleotide primers, DNA polymerase, deoxynucleotide triphosphate (dNTP), reaction buffer, and a magnesium salt. The PCR reaction consists of three basic steps: denaturing double-stranded DNA (dsDNA) to single-stranded DNA, annealing primers to single-stranded DNA (ssDNA), and extending the primers using DNA polymerase. In a typical process, the denaturation step involves heating the reaction mixture in a reaction buffer to a temperature typically between 92°C and 97°C, annealing the primers to a single DNA strand by cooling the mixture to approximately 50°C and 60°C, and extending the primers with DNA polymerase at approximately 72°C. Repeating the three-step cycle doubles the amount of the target sequence. If the process is repeated many times, the theoretical yield after 20 to 35 repeated cycles can reach amplifications of over one billion times for a selected region. The Klenow fragment of DNA polymerase I, the polymerase used by Mullis's team in their early studies, is unstable at DNA denaturation temperatures, and therefore a new enzyme had to be added in each cycle. Thus, although the concept was very interesting, the process was inefficient and initially not promising as a routine testing technique.In 1988, the introduction of thermostable polymerases, starting with Taq DNA polymerase (derived from Thermus aquaticus—a thermophilic bacterium found in the hot springs of Yellowstone National Park), proved useful in making PCR an acceptable testing technique (Saiki et al., 1988).

[0003] While the basic PCR process may seem surprisingly simple, its practical application in research and industry has been fraught with countless obstacles and difficulties. Of course, progress has been made in many areas, and the usefulness of the technology has improved, but as the cited literature suggests, further progress is still ongoing. Generally, the progress made can be divided into three main categories: a) improving or engineering and creating better reaction media (reaction buffers); b) discovering and / or developing better polymerases; and c) developing improved protocols and novel and improved instruments. This invention relates to both (a) and (b), but places particular emphasis on (b).

[0004] One major problem with the PCR process is that when the target being amplified has a high GC content, the yield of the product is low or nonexistent, and / or its fidelity is poor (Henke et al., 1997). Within the complementary strand of DNA, there are only two hydrogen bonds between A and T nucleotides, while there are three hydrogen bonds connecting G and C nucleotides, making high GC-content regions of DNA resistant to thermal denaturation. When the GC content of a region exceeds 50%, 95 cEven heating to op does not necessarily cause complete denaturation, and heating to higher temperatures often leads to other problems, including nucleic acid chain degradation, due to depurine and deamination, as well as slow hydrolysis of phosphodiester bonds (Lindahl et al., 1972, 1974). Even more troublesome is the fact that DNA polymerase begins to rapidly deactivate at temperatures above 95°C. For example, the half-life of Taq polymerase is 40 minutes at 95°C, 9 minutes at 97.5°C, and 0.3 minutes at 100°C (Innis et al., 1995). Another problem associated with high GC targets is the formation of secondary structures (hairpins, dumbbells, etc.) within denatured single-strand DNA (ssDNA). Such structures can interfere with polymerase progression during the extension reaction and may be involved in the generation of nonspecific products (Fry et al., 1992).

[0005] Early researchers found that adding certain organic compounds such as formamide (HCONH2), dimethyl sulfoxide (DMSO), and betaine could help amplify some DNA targets that were difficult to amplify (Sarkar et al., 1990; Pomp et al., 1991; Henke et al., 1997).

[0006] Despite the above advancements, many high-GC targets could not be amplified even with the help of such adjuvants (Chakrabarti, 2002). Chakrabarti and Schutt found that certain low-molecular-weight organic solvents dramatically improved the PCR amplification of DNA targets that were highly GC-containing and could not be amplified separately, and that the fidelity of the amplified products was also significantly improved in the presence of such solvents; they found that four groups of low-molecular-weight solvents (amide, sulfoxide, sulfone, and polyol (especially diol)) were particularly effective and significantly more potent than others previously described (Chakrabarti, 2002, 2004; Chakrabarti et al., 2001 Nucleic Acids Research; Chakrabarti et al., 2001 Gene; Chakrabarti et al., U.S. Patent No. 6,949,368; Ibid. No. 7,276,357 B2; and Ibid. No. 7,772,383 B2). These inventors are licensed by major biotechnology companies and are currently the preferred method for amplifying high-GC DNA targets in both industry and university laboratories.

[0007] While the aforementioned low molecular weight organic solvents have proven highly effective in amplifying many high-GC DNA targets, they suffer from the limitation of their application range due to the reduction in the half-life of DNA polymerase in their presence. This is especially true at higher temperatures.

[0008] What is required is a modified DNA polymerase that is heat-resistant and has better overall compatibility for implementation in the presence of the solvents outlined above and in aqueous organic media in general. [Overview of the project]

[0009] This invention generally relates to molecular biology and molecular biological methods for selecting nucleic acids encoding gene products. More specifically, this invention relates to compositions and methods for enhancing polynucleotide amplification reactions in organic aqueous media.

[0010] The compositions and methods described herein provide improved variant DNA polymerases for use in specific applications.

[0011] For example, in some embodiments, a) G3D, M4I, L5Q, F8L, E9V, P10S, V14A, L16P, H21R, A23P, L22M, F27S, A29T, G32D, G38D, K53N, A54V, L55P, A61V, D67G, P71L, R74L,R74H,R74C, K82N, G84D, A86V, P87Q, P89S, E90D, A97T, V103A, D104G, A109V, R110Q, P114S, G115D, E117D, A118V, A118T, K128R, V136A, L149P, L162P, K171T, A180V, R183H, T186I, G187S, D191N, L193R, G195S, G200S, E201K, K202R, R205H, K206Q, G212D, S213N, S213G, N220D, L224Q, I228V, H235Y, D237G, W243R, D244E, D244V, L254P, K260N, F258S, R261H, P264S, E267K, E277G, L287Q, S290G, K292N, P302L, P302S, V310L, L311M, D320N, A326V, R328H, H333R, K346R, L351M, E363D, L365Q, P382T, N384D, E388D, T399A, A414S, A454E, A454L, A454V, A458V, L461Q, F482I, L461R, V474I, G499D A502T, I503T, E507K, S515N, S515G, A516G, E520G, A521V, I528T, K531R, Q534R, T539A, S543G, D551N, D551G, V586A, V586M, Q592R, L606M, A608T, S612R, I665V, F667Y, H676L, H676R, H676Y, Q680R, E681K, K702R, A705V, V720L, V730I, D732G, D732N, E734G, V737D, V737A, S739G, V740A, V740I, E742K, F749V, F749IF749L, K762R, K767R, L768M, E773K, L781P, E797G, E797Q, V799A, P812Q, Q782H, A814V, L813M, E825Q, and E832K (for example, L5Q, F8L, P10S, L16P, A23P, A29T, K31R, G38D, A61V, P89S, A97T, A118V, L162P, K171T, T186I, E201K, R205K, K206Q, G208S, K219E, N220D, I228V, M236T, D244E, D244V, R261H, D273G, L287Q, S290G, V310L, H333R, K346R, L351M, P382T, E388D, E434D, A454E, L461Q, L461R, V474I, F482I, I503T, E507K, S515N, A521V, Q534R, S543G, D551G, D551N, Q592R, L606M, A608V, S612R, H676L, Q680R, K702R, D732N, E734G, S739G, E742K, F749I, F749V, F749L, K762R, K767R, L768M, Q782H, or E832K; for example, L5Q, F8L, P10S, L16P, A23P, A29T, T186I, K31R, G38D, A97T, A118V, L162P, R205K, G208S, K219E, N220D, I228V, D273G, S290G, K346R, P382T, E388D, E434D, A454E, L461Q, L461R, V474I, F482I, I503T, E507K, S515N, A521V, Q534R, D551G, L606M, A608V, S612R, Q680R, K702R, E734G, S739G, E742K, F749V, F749I, F749L, K762R, K767R, L768M,A modified Taq DNA polymerase having the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) having one or more amino acid changes selected from the group consisting of Q782H or E832K (for example, L5Q, P10S, A23P, A29T, T186I, L461R, E507K, A608V, S612R, E742K, F749L, F749I, K762R, K767R, or E832K), b) A PCR buffer containing, for example, one or more low molecular weight organic solvents selected from amides, sulfoxides, sulfones, or diols, wherein one or more low molecular weight organic solvents are present in an amount of about 0.05 to about 3.0 moles (e.g., 0.05 to 2.5 moles, 0.05 to 2 moles, 0.05 to 1.5 moles, 0.05 to 1.0 moles, 0.1 to 3.0 moles, 0.1 to 2.5 moles, 0.1 to 2). The PCR buffer contains the following components at concentrations in the range of 1.5 mol, 0.1-1.0 mol, 0.5-3.0 mol, 0.5-2.5 mol, 0.5-2.0 mol, 0.5-1.5 mol, 0.5-1.0 mol, 1.0-3.0 mol, 1.0-2.5 mol, 1.0-2.0 mol, 1.0-1.5 mol, 1.5-3.0 mol, 1.5-2.5 mol, 1.5-2.0 mol, or 2.0-3.0 mol): Compositions containing the above are presented herein.

[0012] In some embodiments, at least one amino acid change is selected from, for example, P10S, L16P, A29T, K31R, G38D, A61V, A118V, L162P, T186I, G208S, N220D, I228V, D244V, D273G, S290G, K346R, L351M, E388D, A454E, L461Q, L461R, F482I, I503T, S515N, A521V, Q534R, D551G, L606M, A608V, S612R, Q680R, E734G, S739G, F749V, F749I, L768M, or E832K. In some embodiments, at least one amino acid change is selected from the group consisting of, for example, F8L, P10S, L16P, A29T, K31R, G38D, A61V, A97T, or L162P. In some embodiments, at least one mutation is selected from, for example, A186I, D244V, R205K, G208S, K219E, N220D, I228V, D273G, S290G, K346R, P382T, E388D, E434D, A454E, L461Q, L461R, V474I, F482I, I503T, E507K, S515N, A521V, Q534R, D551G, or L606M. In some embodiments, at least one amino acid change is A608V. In some embodiments, at least one of the amino acid substitutions is selected from, for example, S612R, Q680R, K702R, S739G, E742K, L768M, F749I, F749V, K762R, K767R, or Q782H. In some embodiments, at least one of the amino acid changes is E832K. In some embodiments, up to 12 amino acid substitutions may be present in the Taq polymerase.

[0013] Further embodiments include compositions comprising a modified Taq DNA polymerase suitable for PCR reactions in an organic aqueous medium, wherein the organic aqueous medium comprises one or more low molecular weight organic solvents selected from the group consisting of, for example, amides, sulfoxides, sulfones, and diols, and the amino acid sequence of the modified Taq DNA polymerase is, for example, L30P, A54V, E434D, K206Q, S612R, V730I, and F749V; P10S, A61V, T186I, D244V, K314R, E520G, V586A, S612R, V730I, and F749V; G12T, A54V, T186I, D244V, F667Y, and F749V; P10S, A61V, F73S, T186I, R205K, K219E, M236T, A608V, S612R, and 2494ΔG; P10S, L30P, A61V, L365P, V586A, S612R, and E832K; P10S, A61V, D244V, S612R, and E832K; L30P and 2494ΔG; A29T, G200S, D237G, and F749I; L16P, F73S, E388D, Q680R, and F749I; F73S, K346R, A454E, and F749V; F73S, A118V, and F749I; A23P, L162P, I228V, L461R, A521V, E734G, F749I, and L768M; K31R, F482I, Q534R, A608V, and F749I; A23P and F749I; G38D, F73S, A454V, and F749V; N220D, I503T, S515N, and F749V; A29T, F73S, S290G, L461R, D551G, L606M, S739G, and F749I; E434D, A608V, and K762R; E434D, E507K, and K762R; E434D, E507K, E742K, and F749I; P10S, P382T, E434D, and E507K; R205K, K219E, E434D, V474I, A608V, inS661R, E742K, and F749I; A97T, A608V, K702R, and K762R; F8L, P10S, E434D, E507K, K762R, and K767R; P10S, E507K, Q680R, and K762R;E507K, A608V, Q782H, and F749I; E434D, A608V, E742K, and F749I; E520G, V586A, S612R, and 2493ΔA; P10S, V730I, and 2493ΔA; V586A, S612R, S674S, and 2494ΔGA; E434D and 2494ΔGA; Y116Stop2494ΔG; A54V; A61V; F749V; E832K; T186I, V586A, S612R, and 2494ΔG; A64V and 2493ΔA; D244V, K314R, V586A, and S612R; A61V, T161I, V586A, S612R, and 2494ΔG; G12T, A61V, and 2494ΔG; A29T, K53R, R205K, K219E, D320N, A326V, N415D, L461R, E602D, and A608V; A29T, K53R, R205K, K219E, D244E, D320N, A326V, N415D, L461R, and A608V; A29T, K53R, R223P, D320N, A326V, N415D, L461R, E602D, and A608V; A29T, D238E, R328H, L461R, A608V, E745K, and F749I; A29T, F73S, D238E, R328H, D551N, A608V, E745K, and F749I; A29T, D238E, R328H, D551N, A608V, and F749V; A109V, L224Q, T399A, A502T, A608V, and F749I; A109V, L224Q, T399A, A502T, A608V, S739G, and F749I; A29T, L224Q, T399A, A454E, A608V, S739G, and F749I; K53R, F73S, A141P, P382S, A472G, R556G, and F749I; R110L, K219E, M236T, E274K, R492L, A608V, E626D, K767R, and E825K; R110L, K219E, M236T, N415Y, R492L, A608V, K767R, and E832N; K82I, K219E, M236T, N415Y, R492L, A608V, E626V, and K793R;P10S, F73S, K219E, M236T, E337D, E507K, A608V, and K767R; P10S, F73S, K219E, E337D, E434D, V474I, A608V, and K767R; P10S, F73S, K219E, E337D, E434D, A608V, and K767R; P10S, V14A, R205K, K219E, M236T, N384D, V474I, A608V, S612R, and K762R; A composition is presented that is 90% identical to the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) having amino acid changes selected from the group consisting of P10S, V14A, K219E, N384D, E434D, V474I, A608V, S612R, and K767R; P10S, V14A, R205K, K219E, N384D, V474I, A608V, S612R, and F749I; and R110L, R205K, K219E, N415Y, S543I, A608V, E626D, K767R, and E825K.

[0014] An additional embodiment is a composition comprising one or more DNA polymerases having increased thermal stability compared to wild-type Taq DNA polymerase in a PCR buffer containing 0-10% by weight of one or more organic cosolvents, wherein the one or more DNA polymerases are, for example, P10S, G12T, L16P, A23P, A29T, L30P, K31R, G38D, A61V, A64V, F73S, Y116Stop, A118V, T161I, L162P, T186I, G200S, N220D, I228V, D237G, D244V, S290G, K314R, K346R, E388D, E434D, A454E, A454V, L461R, F482I, I503T, S515N, E520G, A521V, Q534R, D551G, V586A, L606M, A608V, S612R, Q680R, V730T, E734G, S739G, F749I, F749V, L768M, The present invention provides a composition comprising a modified Taq DNA polymerase having an amino acid sequence consisting of the amino acid sequence of a wild-type Taq DNA polymerase (SEQ ID NO: 41) having one or more amino acid changes selected from the group consisting of 2493ΔA or 2494ΔG (for example, P10S, A29T, L30P, K31R, F73S, A118V, G200S, G237G, K346R, S434D, A454E, F482I, E520G, Q534R, V586A, A608V, S612R, V730I, F749I, F749V, 2493ΔA, or 2494ΔG).In some embodiments, one or more DNA polymerases are, for example, F749V; F30L and 2494ΔG; E520G, V586A, S612R, and 2493ΔA; E434D and 2494Δ; P10S, V730I, and 2493ΔA; V116Stop and 2494ΔG; A64V and 2493ΔA; T186I, V586A, S612R, and 2494ΔG; V586A, S612R, and 2494ΔG; D244V, K314R, V586A, and S612R; A61V, T161I, V586A, S612R, and 2494ΔG; G12T, A61V, and 2494ΔG; A29T, G200S, D237G, and F749I; L16P, F73S, E388D, Q680R, and F749I; F73S, K346R, A454E, and F749V; F73S, A118V, and F749I; A23P, L162P, I228V, L461R, A521V, E734G, F749I, and L768M; K31R, F482I, Q534R, A608V, and F749I; A23P and F749I; G38D, F73S, A454V, and F749V; It has an amino acid sequence that is at least 90% identical to the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) having amino acid changes selected from the group consisting of N220D, I503T, S515N, and F749V; or A29T, F73S, S290G, L461R, D551G, L606M, S739G, and F749I.

[0015] Other embodiments include a composition comprising one or more DNA polymerases having increased fidelity compared to wild-type Taq DNA polymerase in a PCR buffer containing 0-10% by weight of one or more organic cosolvents, wherein one or more DNA polymerases are, for example, P10S, G12T, A23P, K31R, A54V, A61V, F73S, Y116Stop, A118V, L162P, T186I, K206Q, I228V, D244V, K314R, L461R, F482I, A521V, Q534R, V586A, A608V, S 612R, E734G, F749I, L768M, E832K, 2494ΔG, A23P, K31R, L162P, I228V, L461R, F482I, A521V, E734G, F749I, or L768M (for example, K31R, A54V, F73S, A118V, T186I, K206Q, D244V, K314R, F482I, Q534R, V586A, A608V, S612R, F749I, E832K, or 2494ΔG; The present invention provides a composition comprising a modified Taq DNA polymerase having an amino acid sequence consisting of the amino acid sequence of a wild-type Taq DNA polymerase (SEQ ID NO: 41) having one or more amino acid changes selected from the group consisting of A54V, T186I, or E832K. In some embodiments, one or more DNA polymerases have an amino acid sequence that is at least 90% identical to the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 1) having amino acid changes selected from the group consisting of, for example, A54V; T186I; E832K; D244V, K314R, V586A, and S612R; K206Q and 2494ΔG; G12T, A61V, and 2494ΔG; P10S; K31R, F482I, Q534R, A608V, and F749I; F73S, A118V, and F749I; or A23P, L162P, I228V, L461R, A521V, E734G, F749I, and L768M.

[0016] One particular embodiment provides a composition comprising one or more DNA polymerases, wherein, in a PCR buffer containing 0-10% by weight of one or more organic cosolvents, the nucleotide integration rate is increased and the processing capacity is increased compared to wild-type Taq DNA polymerase, and the composition comprises one or more DNA polymerases having an amino acid sequence consisting of the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) having one or more amino acid changes selected from the group consisting of, for example, A29T, V310L, A454L, H676R, E687K, D732G, V737D, V740A, F749V, or 2494ΔG (e.g., V310L, F749Y, or 2494ΔG). In some embodiments, one or more DNA polymerases have an amino acid sequence that is at least 90% identical to the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) having amino acid changes selected from the group consisting of, for example, F749V; F310L; 2494ΔG; A454L, F749V, and 2494ΔG; H676R and D732G; E687K and 2494ΔG; A29T and V737D; or V740A and F749V.

[0017] The present invention is not limited to specific organic cosolvents. Examples include, but are not limited to, low molecular weight amides, low molecular weight sulfoxides, low molecular weight sulfones, or low molecular weight diols. In some embodiments, amides include, for example, formamide, N-methylformamide, N,N-dimethylformamide (DMF), acetamide, N-methylacetamide, N,N-dimethylacetamide, propionamide, isobutylamide, 2-pyrrolidone, N-methylpyrrolidone (NMP), N-hydroxyethylpyrrolidone (HEP), N-formylpyrrolidine, and N-formylmorpholine; The sulfoxide is selected from δ-valerolactam, ε-caprolactam, or 2-azacyclooctanone; the sulfoxide is selected from, for example, dimethyl sulfoxide (DMSO), n-propyl sulfoxide, n-butyl sulfoxide, methyl sec-butyl sulfoxide, or tetramethylene sulfoxide; the sulfone is, for example, dimethyl sulfone, diethyl sulfone, di(n-isopropyl) sulfone, tetramethylene sulfone (sulfolane), or 2,4-dimethyl sulfone The diol is selected from, for example, 1,2-propanediol, 1,3-propanediol, 1,2-butanediol, 1,3-butanediol, 1,4-butanediol, 1,2-pentanediol, 2,4-pentanediol, 1,5-pentanediol, 1,2-cyclopentanediol, 1,2-hexanediol, 1,6-hexanediol, or 2-methyl-2,4-pentanediol. In some embodiments, the amide solvent is N,N-dimethylformamide (DMF) at a concentration of about 0.5 to about 1.5 molars; isobutylamide at a concentration of about 0.1 to about 1.0 molars; 2-pyrrolidone at a concentration of about 0.1 to about 1.0 molars; or N-methylpyrrolidone at a concentration of about 0.1 to about 1.0 molars.In some embodiments, the sulfoxide is dimethyl sulfoxide (DMSO) at a concentration of about 0.5 to about 3.0 molars, or tetramethylene sulfoxide at a concentration of about 0.1 to about 1.0 molars. In some embodiments, the sulfone is tetramethylene sulfone (sulfolane) at a concentration of about 0.1 to about 1.0 molars. In some embodiments, the diol is 1,3-propanediol at a concentration of about 0.5 to about 3.0 molars; 1,4-butanediol at a concentration of about 0.5 to about 2.0% molars; or 1,5-pentanediol at a concentration of about 0.5 to about 1.0% molars.

[0018] Although the Taq polymerase variants of this application are described above in conjunction with the consideration of solvents and / or reaction media, the Taq polymerase variants are considered herein as compositions in themselves, independently of either the consideration of solvents or reaction media.

[0019] In further embodiments, kits or systems comprising the modified DNA polymerase and organic cosolvent described herein are presented herein. In some embodiments, the modified DNA polymerase has an amino acid sequence consisting of the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) having one or more amino acid changes, in which case the one or more amino acid changes are, for example, L30P, A54V, E434D, K206Q, S612R, V730I, F749V; P10S, A61V, T186I, D244V, K314R, E520G, V586A, S612R, V730I, F749V; G12T, A54V, T186I, D244V, F667Y, F749V; P10S, A61V, F73S, T186I, R205K, K219E, M236T, A608V, S612R, 2494ΔG ; P10S, L30P, A61V, L365P, V586A, S612R, E832K ; P10S, A61V, D244V, S612R, E832K ; L30P, 2494ΔG ; E520G, V586A, S612R, 2493ΔA ; P10S, V730I, 2493ΔA; V586A, S612R, S674S, 2494ΔGA; E434D, 2494ΔGA; Y116Stop2494ΔG; A54V; A61V; F749V; E832K; T186I, V586A, S612R, 2494ΔG; A64V, 2493ΔA; D244V, K314R, V586A, S612R; A61V, T161I, V586A, S612R, 2494ΔG; G12T, A61V, 2494ΔG; T186I; K206Q and 2494ΔG; P10S; F310L; 2494ΔG; A454L, F749V, 2494ΔG; H676R and D732G; E687K and 2494ΔG; A29T and V737D; V740A and F749V; A29T, K53R, R205K, K219E, D320N, A326V, N415D, L461R, E602D, A608V; A29T、K53R、R205K、K219E、D244E、D320N、A326V、N415D、L461R、A608V; A29T、K53R、R223P、D320N、A326V、N415D、L461R、E602D、A608V; A29T、D238E、R328H、L461R、A608V、E745K、F749I; A29T, F73S, D238E, R328H, D551N, A608V, E745K, F749I; A29T, D238E, R328H, D551N, A608V, F749V; A109V, L224Q, T399A, A502T, A608V, F749I; A109V, L224Q, T399A, A502T, A608V, S739G, F749I; A29T, L224Q, T399A, A454E, A608V, S739G, F749I; K53R, F73S, A141P, P382S, A472G, R556G, F749I; R110L、K219E、M236T、E274K、R492L、A608V、E626D、K767R、E825K; R110L、K219E、M236T、N415Y、R492L、A608V、K767R、E832N; K82I, K219E, M236T, N415Y, R492L, A608V, E626V, K793R; P10S, F73S, K219E, M236T, E337D, E507K, A608V, K767R; P10S, F73S, K219E, E337D, E434D, V474I, A608V, K767R; P10S, F73S, K219E, E337D, E434D, A608V, K767R; P10S、V14A、R205K、K219E、M236T、N384D、V474I、A608V、S612R、K762R; P10S、V14A、K219E、N384D、E434D、V474I、A608V、S612R、K762R、K767R; P10S, V14A, R205K, K219E, N384D, V474I, A608V, S612R, F749I; R110L, R205K, K219E, N415Y, S543I, A608V, E626D, K767R, E825K; K31R, F482I, Q534R, A608V, F749I; F73S, A118V, F749I; A23P, L162P, I228V, L461R, A521V, E734G, F749I, L768M; A29T, G200S, D237G, F749I; L16P, F73S, E388D, Q680R, F749I; F73S, K346R, A454E, F749V; A23P, F749I; G38D, F73S, A454V, F749V; N220D, I503T, S515N, F749V; The model is selected from the group consisting of A29T, F73S, S290G, L461R, D551G, L606M, S739G, and F749I.

[0020] Additional embodiments are described herein.

[0021] The patent or application file shall contain at least one colored drawing. A copy of this patent or patent application publication, including the color drawing, shall be provided by the office upon request and payment of the necessary fees.

[0022] While the subject matter of this disclosure has been described in general terms, please refer here to the attached drawings (which are not necessarily represented to an accurate scale). [Brief explanation of the drawing]

[0023] [Figure 1]This figure shows the crystal structure of Taq DNA polymerase. The figure describes the crystal structure of Taq DNA polymerase, which consists of 834 amino acids. The depiction can be likened to a partially closed right hand, with each domain identified as the "palm," "thumb," and "fingers." The palm is the polymerase active site. The amino acid domains from the 1st to the 288th amino acid in the longitudinal direction are the 5' to 3' (exonuclease) activity site (a sheet of catalytic amino acids) located at the base of the palm. The thumb and fingers hold the extended DNA conformation in place. [Figure 2] This figure shows the 3D structure of Taq polymerase. This figure illustrates the location of a specific important mutant in the 3D structure of Taq polymerase. [Figure 3] This figure shows a partial list of the structures of organic cosolvents. This figure lists the chemical structures of exemplary organic cosolvents useful in embodiments of the present invention. [Figure 4] This figure shows Winsor's R theory for emulsion formation. R is expressed as the ratio between the tendency of a surfactant monolayer to be convex toward oil and the tendency of the same layer to be convex toward water. When R=1, crystalline layered micelles (open-end type) are formed; when R>1 or <1, ​​closed-type spherical micelles—oil-external or water-external—are formed. For further details, see "Detailed Description of Preferred Embodiments". [Figure 5] This figure shows a list of emulsifiers that can be used in the CSR of the present invention. This figure presents a list of emulsifiers that can be used to create the W / O emulsion of the present invention. They belong to a class of surfactants called nonionic surfactants. Preferred emulsifiers may consist of one or more molecules belonging to the chemical group shown in Figure 5. [Figure 6]This figure shows a list of fluorinated surfactants that can be used as emulsifiers in the CSR of the present invention. This figure presents examples of fluorinated surfactants that can be used to create the W / O emulsion of the present invention, particularly when the oil used is a fluorinated synthetic oil. Fluorinated surfactants are characterized by having conventional hydrophilic tails, such as polyethylene oxy chains or highly hydrophobic fluorocarbon chains. When present in a water-oil system even at very low concentrations, the low surface energy of the fluorocarbon tails results in very low interfacial tension between the oil phase and the water phase. Numerous companies worldwide produce fluorocarbon surfactants today, but the first and most widely known producer of fluorinated surfactants is 3M. ​​3M products are marketed under the Novac® brand. Representative examples include 3M® fluorinated surfactants 4430, 4432, and 4434. These products have been found to be used in a variety of applications in industry, medicine, and biotechnology. [Figure 7]Stable water-in-oil reversed-phase emulsion containing single cells: This figure shows a polydisperse emulsion. The oil phase is light oil and the emulsifier is a mixture of nonionic surfactants. These figures show the structure of a water-in-oil reversed-phase emulsion consisting of polar internal droplets, an organic aqueous medium (1,4-butanediol being the organic component and constituting 5% of the composition), a "compound 1×Taq buffer" containing 20 mM Tris-HCl, 50 mM KCl, 50 μM tetramethylammonium chloride, 250 μM dNTPs, and a pair of 1 μM adjacent PCR primers, and an expression cell. The oil phase was light oil. The emulsifier used was a mixture of Span 80, Tween 80, and Triton X100. The average droplet size of the internal phase was 25 μM, and the size of individual droplets ranged from 15 μM to 50 μM. Images A and B show microscopic fluorescence images of GFP-expressing Escherichia coli (E. coli) cells in solution and emulsion. Images C and D show bright-field images of the emulsion obtained under a light microscope before and after the PCR cycle. As can be seen, the high-temperature denaturation step dissolves the cell walls, and therefore no original cells are observed after CSR. [Figure 8]Emulsion integrity - This figure shows that no crossover is observed between droplets during the PCR period. This figure demonstrates the integrity of the emulsion droplets in Figure 7 when used as standalone containers for performing PCR reactions, meaning that there is no crossover of reactants from one droplet to another during the PCR reaction period. Lane 1: DNA marker. Lanes 2 and 3: Emulsions from PCR performed in the absence of an organic co-solvent. Lanes 4 and 5: Emulsions from PCR performed in the presence of an organic co-solvent (1,4-butanediol). Here, the same experiment as in lane 2 qn3 was repeated, except that the taq buffer contained 5% 1,4-butanediol. Lanes 6, 7, and 8: Solutions from PCR performed in the absence of an organic co-solvent. These were control experiments for the experiments in lanes 2 and 3. Lane 6 contained T1, T2, and their respective primers, as well as polymerase. The gel shows that both amplicons were amplified as expected. Lane 7 contained only T1, its primer, and polymerase. The gel showed only one amplification band, the T1 amplification band, as expected. Lane 8 contained only T2, its primer, but not polymerase. The gel showed no amplification band, as expected. Lanes 9, 10, and 11: Solutions of PCR performed in the presence of an organic cosolvent (1,4-butanediol). Lanes 9, 10, and 11 were repeats of lanes 6, 7, and 8, except that in each case 5% 1,4-butanediol was present in the reaction mixture. The results were similar to those of lanes 6, 7, and 8. [Figure 9A]Top panel: Stable water-in-oil reversed-phase emulsion containing a single cell (before and after PCR): This figure shows a monodisperse emulsion prepared using a μEncapsulator from Dolomite Microfluidics (UK). The oil phase is a low-viscosity fluorinated synthetic oil, and the emulsifier is a nonionic fluorinated surfactant. This figure shows the structure of a monodisperse water-in-oil reversed-phase emulsion prepared using a mechanical device (μEncapsulator from Dolomite Microfluidics (UK)) according to the manufacturer's instructions. The first two plates in the figure show that the monodisperse droplets contain a single GFP-expressing bacterium, with or without 5% organic cosolvent (1,4-butanediol), and the number of bacteria per droplet does not exceed one. The second two plates show the same droplets after simulated PCR [25× (30 seconds at 94°C, 30 seconds at 55°C, and 3 minutes at 72°C)] for 5 minutes at 95°C, followed by holding at 4°C. In the plate after PCR, no cells are observed inside the droplets because the cell walls were thermally dissolved during the PCR period. The second portion of Figure 9 shows the amplified polymerase DNA of the expression cells after isolation by emulsion extraction using perfluoro-1-octanol (Sigma Cat #370533) and subsequent centrifugation. The uppermost aqueous layer containing the amplified DNA was analyzed on a 1% agarose gel. NC = negative control; PC = positive control; M = 1kb DNA marker. [Figure 9B]Bottom panel: Preparation of a stable aqueous-oil-aqueous emulsion (biemulsion) using Dolomite Microfluidics (UK). A primary emulsion (PE) was prepared (A), followed by PCR. A portion of the primary emulsion was used to isolate DNA and run on a 1% agarose gel (B). Panel B: Lane 1 is the DNA marker, Lane 2 is the negative control, and Lane 3 is the positive control. After PCR, the primary emulsion was collected and the biemulsion was prepared as described in the "Materials and Methods" section (C). The biemulsion is shown in (D). After PCR, the positive control was stained with SYBR Green I and visualized under a fluorescence microscope (E). Images before and after sorting are shown in Panels F and G, respectively. FACS sorting was performed on the bilayer emulsions, and a total of 1.6 million events were randomly captured. A threshold of 5000 was applied to gate the parent DE (H and I), followed by sorting of SYBR-positive bilayer emulsions (J). SSC: Side scatter; FSC: Forward scatter; A: Area; H: Height. [Figure 10] This figure shows a schematic diagram of CSR. The Taq polymerase gene was diversified by epPCR, followed by digestion of the PCR product using XbaI and SalI restriction enzymes, and then cloning into the pASK-IBA5C plasmid. [Figure 11] This figure shows the development of selection pressure for CSR selection within 5% 1,4-butanediol. The activity of WT Taq DNA polymerase was determined in the absence and presence of 5% 1,4-butanediol to develop appropriate selection pressure for CSR selection experiments. [Figure 12] This figure shows the amount of DNA and the peak area of ​​its melting curve. A linear correlation exists between the amount of DNA and the peak area of ​​its melting curve. [Figure 13]This figure shows the effect of aqueous organic media on DNA melting and enzyme efficiency. A) Computational prediction of GC content of c-Jun fragments used in this study. B) c-Jun DNA Tm was determined in 0-10% BD. C) From the linear correspondence between BD concentration and the decrease in DNA melting temperature (Tm), the slope of the plot (dTm / [BD]) was determined to be 5.9 K / M. D) Polymerization efficiency of selected mutants was evaluated in 0-1.2 M (0-10%) organic solvents, and it was revealed that the melting temperature of wild-type Taq polymerase decreased compared to the mutants, suggesting that mutant polymerase is resistant to the denaturing effect of BD. E) Cq change rate of mutants across the entire BD concentration range. [Figure 14] This figure shows the amplification efficiency of engineered polymerases in the presence of a co-solvent on Taq and c-jun templates. Selected clones were used to evaluate their amplification efficiency for wild-type and its variants, using two different templates and varying co-solvent concentrations. The figure shows representative qPCR results for the clones used in real-time PCR assays. To evaluate efficiency, each polymerase with equal activity was tested under identical conditions. The following PCR cycles were used for (A) Taq template in 5% BD, (B) c-jun template in 0–8% BD using selected clones from the initial screening round and associated synthetic clones, (C) Taq template, and (D) c-jun template in 0–10% BD using selected clones from the later screening round: 1 minute at 98.3°C, 6 minutes at 95°C, followed by 17 cycles of 30 seconds at 94°C, 30 seconds at 57.8°C, and 30 seconds at 72°C. [Figure 15] This figure (6) shows a segment of the Taq variant gene created for NGS analysis. It shows the fragment (amplicon for NGS) corresponding to the sequence in the parental wild-type Taq polymerase. [Figure 16]This figure shows the evaluation of WT Taq polymerase and Taq polymerase variant L-5-2-F01 in amplification of GC-enriched targets derived from human genomic DNA, using the following PCR cycling protocol (1 minute at 98.3°C + 6 minutes at 95°C, followed by 25 cycles of 30 seconds at 94°C, 30 seconds at 57°C, and 50 seconds at 72°C), employing high denaturation temperatures and up to 5% and 7% BD, respectively. Final extension was performed at 72°C for 2 minutes, followed by holding at 4°C. In a 50 μL reaction volume, the PCR mixture contained 1× PCR buffer (Invitrogen), 1.5 mM MgCl2, 0.25 mM dNTPs, 25 ng of human gDNA (Promega #G1471), 0.5 μM forward and reverse primers, and 2.5 U of polymerase. PCR products were analyzed on a 1% agarose gel. The expected amplicon size is indicated in the figure (as base pairs). WT Taq: (A) 0% BD (B) 5% BD; L-5-2-F01: (C) 0% BD (D) 7% BD. M = 1 kb DNA ladder, the numbers 0.5 and 1 are in kbp. Amplification of the target was not possible in 1% and 2% BD using WT Taq (data not shown). The characteristics of the target are shown in Figure 19. [Figure 17]This figure shows the evaluation of WT Taq polymerase and Taq polymerase variant L-5-2-F01 in amplification of GC-enriched targets derived from human genomic DNA, using the following PCR cycling protocol (1 minute at 98.3°C + 6 minutes at 95°C, followed by 25 cycles at 94°C for 30 seconds, 57°C for 30 seconds, and 72°C for 50 seconds), employing high denaturation temperatures and up to 7% and 10% BD, respectively. The final extension was performed at 72°C for 2 minutes, followed by holding at 4°C. In a 50 μL reaction volume, the PCR mixture contained 1× PCR buffer (Invitrogen), 1.5 mM MgCl2, 0.25 mM dNTPs, 25 ng of human gDNA (Promega #G1471), 0.5 μM forward and reverse primers, and 2.5 U of polymerase. The PCR products were analyzed on a 1% agarose gel. The expected amplicon size is indicated in the figure (as base pairs). WT Taq: (A) 0% BD (B) 7% BD; L-5-2-F01: (C) 0% BD (D) 10% BD (D). M = 1 kb DNA ladder, the numbers 0.5 and 1 are in kbp units. Target characteristics are described in Figure 19. [Figure 18]This figure shows the evaluation of WT Taq polymerase and Taq polymerase variant L-5-2-F01 in amplification of GC-enriched targets derived from human genomic DNA, using a moderate denaturation temperature and a maximum of 7% BD, with the following PCR cycling protocol (2 minutes at 94°C, followed by 30 cycles at 95°C for 30 seconds, 57°C for 30 seconds, and 72°C for 50 seconds). Final extension was performed at 72°C for 2 minutes, followed by holding at 4°C. In a 50 μL reaction volume, the PCR mixture contained 1× PCR buffer (Invitrogen), 1.5 mM MgCl2, 0.25 mM dNTPs, 25 ng of human gDNA (Promega #G1471), 0.5 μM forward and reverse primers, and 2.5 U of polymerase. PCR products were analyzed on a 1% agarose gel. Expected amplicon sizes are indicated in the figure (as base pairs). WT Taq: (A) 0% BD (B) 7% BD; L-5-2-F01: (C) 0% BD (D) 7% BD. M = 1 kb DNA ladder, the numbers 0.5 and 1 are in kbp units. Target characteristics are described in Figure 19. [Figure 19] This figure shows the GC content for each template region in a GC-enriched template. [Modes for carrying out the invention]

[0024] The subject matter of this disclosure is described more fully below with reference to the accompanying drawings illustrating some (but not all) embodiments of the invention. Throughout, similar figures refer to similar elements. The subject matter of this disclosure may be embodied in many different forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided so as to satisfy the applicable legal requirements of this disclosure. In fact, many modifications and other embodiments of the subject matter of this disclosure described herein (which relate to the subject matter of this disclosure and have the benefit of the teachings presented in the above description and the related drawings) will be conceivable to those skilled in the art. Therefore, it is understood that the subject matter of this disclosure is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the accompanying claims.

[0025] 1.Definition Amino acids: As used herein, this term refers to both the 20 natural and non-natural amino acids that make up the entire protein kingdom.

[0026] Cassette mutagenesis: As used herein, the term cassette mutagenesis refers to the process of cutting a cassette from a double-stranded plasmid and replacing it with another (synthetic) cassette containing a mutant.

[0027] Codon: As used herein, the term codon refers to a set of three nucleotide bases within a DNA sequence that codes for a single amino acid. Protein gene sequences typically begin with the ATG codon (coding methionine M) and end with the TAA, TAG, and TGA codons (these codons do not code for any amino acid and merely signal the end of a coding gene).

[0028] Codon Optimization: As used herein, the term codon optimization refers to the process of optimizing the selection of codons that code for a particular amino acid. There are 61 codons that code for the 20 amino acids in a protein. There are more types of codons than amino acids, meaning that one amino acid can be coded by two or more codons. Different organisms exhibit a bias towards which codons they prefer to use to code for a particular amino acid. This bias can affect protein expression in an organism. In molecular biology, when a gene is inserted into a new organism, optimization is often performed to improve the expression of the new gene in that organism for codons that the organism has a positive bias towards for the same amino acid.

[0029] Contig: As used herein, the term contig refers to a set of duplicate DNA segments that together represent a consensus region of DNA.

[0030] Co-solvent: As used herein, the term co-solvent refers to a low molecular weight organic compound that, when added to a PCR reaction buffer, can enhance the amplification reaction in various ways in some embodiments.

[0031] CSR is an abbreviation for Compartmentalized Self-Replication.

[0032] Deep sequencing, also known as high-throughput sequencing or next-generation sequencing (NGS), involves sequencing genomic regions multiple times (often thousands of times). This process allows researchers to detect rare clonal types, even those present in minute quantities (as small as 0.1% of the original sample).

[0033] DNA Shuffling: As used herein, the term DNA shuffling refers to the digestion of a gene into random fragments by DNase 1, and the recombination of these fragments into a full-length gene, typically by primerless and modified PCR. The fragments act as primers for each other based on sequence homology, and recombination occurs when a fragment from one copy of a gene anneals to a fragment from another copy. Modifications to PCR involve an alternating extension process (StEP), in which the annealing and extension steps are significantly shortened to generate alternating DNA fragments and facilitate crossover events (shuffling or fragment switching) along the full length of the template sequence. DNA shuffling can also be generated using restriction enzymes, in which case the fragments can be recombined using DNA ligase. DNA shuffling is an important technique for creating diversification in directed evolution experiments. Diversification results from combining useful mutations from two or more genes into a single gene.

[0034] The effective range of a cosolvent, as used herein, refers to the optimal concentration of a particular cosolvent in an amplification reaction. In some embodiments, the optimal concentration varies depending on the selected cosolvent. The effective concentration can be determined, for example, using the methods described herein.

[0035] Enzyme activity (polymerase activity): One unit of polymerase activity is defined as the amount of polymerase required to synthesize 10 mmol of product within 30 minutes. Therefore, the term refers to the efficiency and selectivity of DNA polymerase.

[0036] Enzyme induction and expression: Enzyme induction is the process by which a molecule (e.g., a drug) induces (initiates or enhances) the expression of an enzyme. Expression is related to production efficiency—high levels of expression of the relevant gene are required to create overproduction.

[0037] Expression cells: For the purposes of this document, these are E. coli cells containing a pool of diversified mutant Taq DNA polymerase genes.

[0038] Fidelity: This term refers to the accuracy of DNA polymerization by template-dependent DNA polymerases. Fidelity is maintained by both 3'→5' exonuclease activity and DNA polymerase activity. Fidelity is measured as the error rate. High fidelity is defined as a mutation rate of 4.45 × 10⁻⁶. -6 This refers to less than one doubling per nt. Low fidelity enzymes are used in error-prone PCR (for example, for mutagenesis).

[0039] Frameshift mutations are a type of mutation that involves the addition (insertion) or deletion of a DNA sequence with a number of base pairs that is not divisible by 3 (e.g., the addition or deletion of nucleotides in numbers such as 1, 2, 4, 5, 7, etc.). Since cells read genes in groups of three base pairs, "divisibility by 3" is of great significance. Each group of three base pairs corresponds to one of 20 different amino acids used to construct proteins. If a mutation disrupts this reading frame, the entire sequence after the mutation is read incorrectly. Therefore, frameshift mutations can dramatically alter proteins by incorporating new nonsense codons or linkage termination codons (TAA, TAG, TGA) and causing premature termination of translation. Polypeptides resulting from such mutations are highly likely to be nonfunctional. The earlier the deletion or insertion occurs in the sequence, the greater the degree of protein alteration. Frameshift mutations can be both dangerous and beneficial. Frameshift mutations are thought to be the underlying cause of dangerous genetic disorders such as Tay-Sachs disease, as well as a predisposition to several types of cancer and familial hypercholesterolemia. Positive effects have been found in some hemophilia cases. These individuals exhibited resistance to the HIV virus and possessed a rare frameshift mutation, CCR5 Δ32, which means a 32-base pair deletion in the CCR5 gene. The CCR5 protein is a cell surface protein that acts as an anchor, through which the AIDS virus (HIV) gains access to cells. A 32-base pair deletion in the CCR5 gene disables the production of the CCR5 protein and therefore disrupts the HIV docking point. [Collins, FS, The Language of Life, Harper Perennial, New York, pp. 169-173, 2010].

[0040] In this case, the inventors observed that many of their variants possess Del A @2493, Del G @2494, and Del GA @2494-2425. Such deletions at the end of the Taq gene mean that, as shown below, the stop codon is moved further away, making the mutant gene longer, and as a result the variant enzyme has 13 more amino acids than the parent protein.

[0041] TIFF2023059361000007.tif69167

[0042] Gene tiling: In this case, the entire genome is broken down into fragments (tiles). This is a whole-genome microarray.

[0043] High GC targets: The average GC content of genomic DNA is approximately 40%. Any polynucleotide with a GC content exceeding 40%, especially those with a GC content exceeding 50%, are called high GC targets. Examples of high GC genes include the 996-base pair c-jun with a 64% GC content and the 660-base pair GTP with a 58% GC content. An example of an extremely high GC gene is the fragile X locus expansion (containing long CGG repeats) in autism patients, which has a GC content exceeding 90%.

[0044] His-tagged polymerase: This is an abbreviation for polymerase tagged with polyhistidine. This tag helps the polymerase molecule bind better to the metal, and therefore makes it easier to purify by column chromatography.

[0045] Ligation: The process of inserting a DNA segment into a plasmid.

[0046] Microarray: A grid of DNA segments consisting of known sequences, used to test and map DNA fragments.

[0047] Next-generation sequencing (NGS): A high-throughput method for deciphering DNA sequence changes.

[0048] The effect of cosolvents is defined in the text titled "Organic Aqueous Media" within the chapter "Detailed Description of Preferred Embodiments".

[0049] Polymerase Processing Capacity: Processing capacity refers to the ability of a DNA polymerase to perform the polymerization sequence without dissociating from the growing DNA strand. This capacity is measured by the length of the polymerized nucleotide chain (e.g., 20 nt, 30 nt, etc.) without interruption by DNA polymerase dissociation. High processing capacity refers to a length greater than 20 nt. Enzymes with higher processing capacity may operate efficiently even at low concentrations.

[0050] Saturated mutagenesis: Also known as single-site saturated mutagenesis, this is a process in which a library is generated by substituting a single amino acid within a specific site with any conceivable amino acid.

[0051] Synthetic Sequencing: This is Illumina Corporation's proprietary high-throughput next-generation sequencing method. The process uses reversible, individually isolated, fluorescently tagged dNTPs to synthesize genes via a modified PCR process, and a 4-pass / band-filter camera / sensor records every nucleotide event for all four types of nucleotides on thousands of templates simultaneously in a massively parallel manner.

[0052] Silent Mutation: A silent mutation is a type of point mutation in which a single base is changed in the protein-coding portion of a gene, but it does not affect the amino acid sequence in the encoded protein. Such a mutation has no effect on the protein it encodes or on the phenotype of the organism.

[0053] Site-directed mutation (SMT), also known as site-specific or oligonucleotide-directed mutation, is an in vitro process that uses specially designed primers to introduce a desired mutation at a specific site within a double-stranded DNA plasmid. Commercially available kits, including instructions, are available to perform this process. Further details are provided in the "Detailed Description of Preferred Embodiments."

[0054] StEP: This is an abbreviation for the alternating extension process, a modified form of PCR in which the annealing and extension steps are significantly shortened to generate alternating DNA fragments and to facilitate a crossover event along the full length of the template sequence. See the section on shuffling for more details.

[0055] Transformation: The process of introducing ligated DNA into cells.

[0056] Non-natural amino acids: These are amino acids that do not occur in natural proteins, but can be introduced into protein structures to create non-natural (synthetic) proteins. explanation

[0057] This invention generally relates to molecular biology and molecular biological methods for selecting nucleic acids encoding gene products. More specifically, this invention relates to compositions and methods for enhancing polynucleotide amplification reactions in organic aqueous media.

[0058] This specification presents an artificially designed DNA polymerase particularly suitable for use in mixed organic aqueous media. Herein, the properties of the aqueous organic media of interest to the inventors are described, and in detail, the various techniques (directed evolution, enrichment of evolved species, next-generation sequencing (NGS), gene synthesis, and theoretical calculations) that the inventors have used in unique ways, either alone or in combination, to arrive at the preferred polymerase compositions of the present invention.

[0059] In some embodiments, methods for engineering DNA polymerase variants particularly suitable for in vitro use in PCR reactions in the presence of specific organic solvents are presented herein. In directed evolution experiments, solvent and temperature were used as “selective pressures.” The primary objective was to develop a list of variants and / or mutations that provided superior temperature and solvent tolerance, but the very fact that the surviving sequences passed compatibility tests to new media also meant that some or many of them possessed other phenotypic properties that conferred additional compatibility to them. In addition to thermal stability, such phenotypic properties included, among others, enzyme activity, DNA binding affinity, processing capacity, ability to amplify long templates, and extension rate (V). max Examples include nucleotide count / second and fidelity. Other properties such as salt tolerance, tolerance to inhibitors, and amplification yield are also included in other properties that may be attributable to or associated with better suitability to harsh in vitro conditions.

[0060] Parent polymerase: The inventors used Taq polymerase as a prototype parent polymerase to develop the desired variant. It is a type A834 amino acid polymerase isolated from the thermophilic Eubacterium Thermus aquaticus (Taq) line YT1 (Lawyer et al., 1989). Some important properties of Taq polymerase include a half-life of 9 minutes at 97.5°C; an optimal activity temperature of 75°C to 80°C; a processing capacity of 50 to 60 nucleotides; an elongation rate of 75 nucleotides / second; and 5'-to-3' nic translation exonuclease activity, but lacking 3'-to-5' proofreading exonuclease activity (see Chakrabarti, 2002).

[0061] The DNA polymerases that can be used in the development of variants according to the present invention are not limited to Taq DNA polymerase alone. Such DNA polymerases can be selected from any type of DNA polymerase, including naturally occurring (wild-type) polymerases, and artificially created polymerases, including truncated fragments derived from natural polymerases; also included in the list are chimeric DNA polymerases, fusion polymerases, and other modified polymerases.

[0062] The naturally occurring polymerases (wild-type) commonly used in PCR reactions are heat-stable polymerases belonging to either the A or B family, meaning they are homologous to E. coli Pol I and II, respectively. The most common B family polymerases are of archaeal (extremophile) origin. Bacterial Taq polymerase belongs to the A family. Common archaeal polymerases—Pfu (Stratagene), Vent / Deep Vent (New England Biolab), KOD (Toyobo), Tgo (Roche), and Pwo (Roche)—belong to the B family.

[0063] Truncated Pol is a polymerase derived from a natural polymerase by removing a specific segment. Examples include the Klenow fragment from *E. coli* Pol I, and the 544-amino acid Stoffel fragment, created by removing a segment from the 834-amino acid Taq DNA polymerase (which helps improve thermal stability).

[0064] A chimeric polymerase is a polymerase that contains sequences derived from two or more natural polymerases. An example is Kofu, which has one segment derived from KOD and one segment derived from Pfu.

[0065] Fusion polymerases are polymerases created by attaching a specific segment of a non-polymerase protein to a native or chimeric polymerase to confer certain desirable properties to it. Examples include Phusion (New England Biolab), PfuUltra® II Fusion (Stratagene), and Herculase II Fusion (Stratagene), which were created by fusing a small basic chromatin-like Sso7d protein to a chimera derived from Deep Vent and Pfu.

[0066] As an example of a modified polymerase, a) Taq polymerase T8 variant derived through directed evolution, containing six mutations - F73S, R205K, K219E, M236T, E434D, and A608V (Ghadessy et al., 2001; Hollinger et al., U.S. Patent No. 7,514,210B2); b) Variants of Kofu and Taq pol described by Bourn et al. (U.S. Patents 8,481,685B2 and 10,457,968B2 granted to KAPA Biosystems); c) The high-speed cycling Taq variant described by Arezi et al. (2014); and d) Variants of archaeal family B polymerases, such as Pfu and ShIB, which have evolved into knockout uracil-binding pockets (Connolly et al., 2009; Tubeleviciute et al., 2010). These are some examples.

[0067] The present invention is not limited to the polymerases listed above, but can be used in any way with Deep Vent, Herculase II Fusion, Klenow fragments, KOD, Kofu, Pfu, PfuUltra II Fusion, Phusion, Pwo, Stoffel fragments, Taq, Tgo, T8, Vent, and various engineered DNA polymerase variants. However, Taq DNA polymerase is the most common and widely used "workhorse" polymerase for PCR reactions, and therefore the inventors have chosen it for their experiments.

[0068] Organic solvents and organic aqueous media: Enzymes evolved in nature to catalyze reactions in water. The use of media that are aquatic and can penetrate organic solvent domains is a human invention to meet the specific needs of in vitro applications of enzymes, both for research and industrial purposes. From this perspective, enzymatic reactions in organic aqueous media constitute an independent field of study.

[0069] Enzymatic reactions occurring in the presence of organic solvents can span the entire spectrum, from almost entirely organic to almost entirely aqueous. Reactions in two-phase mixtures consisting of water and an organic solvent immiscible to water (where the former may be suspended within the latter, or vice versa) can also be included (Koskinen and Klibanov, 1996).

[0070] In the case of organic media, a small amount of water, even a trace concentration, is always necessary for the enzyme to function. Therefore, the field of enzymatic reactions in organic media begins not with 100% organic media, but with nearly organic media. As Kuntz and Kauzmann state, water is an "enzyme molecular lubricant" (Kuntz et al., 1974). An excellent review of early studies on the usefulness of organic solvents for enzymatic reactions is provided by Klibanov (Klibanov, 2001).

[0071] Most of the enzymatic reactions tested to date are hydrolases. Even in those cases, the organic molecules used as substrates in these reactions were limited to small ones, compared to some enzymes that carry out other types of reactions on biomolecules. PCR reactions belong to a different type of reaction, which means that while prior art enzymatic reaction media involving organic solvents had some significance for the inventors, they could not provide any specific guidance.

[0072] In our early research, we identified certain organic cosolvents that, when miscible with water, demonstrated superiority for PCR amplification of many substrates, particularly those with high GC content. Such organic cosolvents belong to four chemical classifications defined by us: low molecular weight amides, sulfoxides, sulfones, and polyols (especially diols) (Chakrabarti, 2002, 2004; Chakrabarti et al., 2001 Nucleic Acid Res; Gene, 2001 Gene; Biotechniques, 2002; U.S. Patent No. 6,949,368; U.S. Patent No. 7,276,357B2; and U.S. Patent No. 7,772,358B2). Early DMF, DMSO, and glycerol have also been reported to have some beneficial effects in PCR amplification of high GC targets (Sarker et al., 1990; Pomp et al., 1991; Henkel et al., 1997). A comprehensive list of some of the more useful low-molecular-weight organic cosolvents is presented below, and some of their chemical structures are shown in Figure 3. a) When selected from low molecular weight amides, the members are formamide, N-methylformamide, N,N-dimethylformamide (DMF), acetamide, N-methylacetamide, N,N-dimethylacetamide, propionamide, isobutylamide, 2-pyrrolidone, N-methylpyrrolidone (NMP), N-hydroxyethylpyrrolidone (HEP), N-formylpyrrolidine, N-formylmorpholine; δ-valerolactam, ε-caprolactam, 2-azacyclooctanone (16 compounds). b) When selected from low molecular weight sulfoxides, the members are dimethyl sulfoxide (DMSO), n-propyl sulfoxide, n-butyl sulfoxide, methyl sec-butyl sulfoxide, and tetramethylene sulfoxide (5 compounds: Figure 3b). c) When selected from low molecular weight sulfones, the members are dimethyl sulfone, diethyl sulfone, di(n-propyl) sulfone, tetramethylene sulfone (sulfolane), and 2,4-dimethylsulfolane, and butadiene sulfone (sulfolene) (6 compounds: Figure 3c). d) When selected from low molecular weight diols, the members are 1,2-propanediol, 1,3-propanediol, 1,2-butanediol, 1,3-butanediol, 1,4-butanediol, 1,2-pentanediol, 2,4-pentanediol, 1,5-pentanediol, 1,2-cyclopentanediol, 1,2-hexanediol, 1,6-hexanediol, and 2-methyl-2,4-pentanediol (13 compounds; Figure 3d). e) Besides diols, triols, namely glycerol (as already mentioned), have also been found to help enhance the amplification of certain high-GC targets. f) Other organic compounds that may belong to the preferred organic component include, in particular, betaine.

[0073] When used as part of a PCR buffer, these cosolvents provide an organic aqueous reaction medium that is essentially primarily aqueous (opposite to the spectrum of the primarily organic reaction media described above). These cosolvents have proven particularly effective in amplifying high GC-containing polynucleotide targets by providing the following benefits: ●Decreased melting temperature of double-stranded DNA: This meant that even very high-temperature melting DNA targets could be denatured more effectively and completely at temperatures below 95°C, i.e., temperatures that did not cause DNA damage. Targets that were not amplified in standard aqueous buffers could also be amplified in these modified buffers. ●In particular, given that such mixed organic aqueous buffers readily release the secondary structures within the ssDNA strand, which are the main cause of the cessation of the extension reaction and thus the formation of nonspecific products, the specificity of the product is improved. ●Polymerase molecules in such mixed solvents exhibit improved fidelity.

[0074] However, such systems have significant limitations that hinder their wider application. The primary limitation is the thermal stability of DNA polymerase in such systems. This limitation manifests itself in different dimensions for each different member of the list. These dimensions are expressed in terms of the effective range, potency, and specificity of each cosolvent, which differ from compound to compound (Chakrabarti R., 2004).

[0075] The effective range of a co-solvent is defined as the concentration at which amplification of a given target reaches its peak, and beyond this concentration, amplification begins to be inhibited. In other words, the effective range of a co-solvent had a certain concentration range (outside which it showed no beneficial effect). This range differed depending on the compound, and also differed even for the same compound if the target was different.

[0076] The effectiveness of a co-solvent is defined as the maximum concentration measurement volume of target band amplification obtainable for any given target amplification within the effective range of the co-solvent. This represents the maximum effectiveness of the co-solvent at the most effective concentration within its effective range.

[0077] The specificity of a co-solvent at a particular concentration is defined (expressed as a percentage) as the ratio of the volume of target band amplification to the total volume of all bands, including undesirable nonspecific bands. For example, false positives and false negatives in PCR-based disease diagnosis are results of poor reaction specificity. The use of co-solvent-based PCR is extremely useful in this area, and some co-solvent patent licensors have licensed their patents specifically for this purpose.

[0078] Most PCR cosolvents as defined above provided excellent amplification in terms of amplification range and specificity of the amplified product. However, for DNA targets, especially those with high GC content and resistance to amplification under standard conditions, the performance was often extremely limited due to the narrow concentration range in which the cosolvent was effective.

[0079] Further investigation revealed that these defects were due to a decrease in enzyme stability (shortening of half-life) in the presence of such co-solvents (Chakrabarti, 2002, 2004). While the half-life of polymerase decreases with increasing temperature, this means that the enzyme's thermal stability decreased with increasing temperature in the presence of co-solvents. Indeed, the thermal stability of DNA polymerase at 92°C–95°C (the range in which the denaturation step of the PCR reaction is typically performed) was significantly reduced by the addition of the most potent and specific co-solvent. The inventors argued that the reason for these defects lies in the fact that most DNA polymerases used in PCR are of natural origin or created from simple manipulations of natural products. Nature has evolved such enzymes for performance in aqueous environments. The inventors further argued that if DNA polymerases could be artificially and optimally designed to function in organic aqueous media as outlined above, the opportunities and scope of application for such novel polymerases would be significantly broadened. The design may include additional factors such as amplification speed (high-speed PCR) in addition to efficacy and specificity.

[0080] When the properties of cosolvents are meticulously tested in terms of their most important characteristics (DNA melting, thermal stability, potency (activity), effective range, and processing capacity) in relation to their influence on the behavior of enzymes in PCR reactions, it is suggested that each group of cosolvents possesses one or two members that best represent that group (Chakrabarti, R., 2002, 2004). Such members include N-methylpyrrolidone and 2-pyrrolidone in the case of amides; dimethyl sulfoxide (DMSO) and tetramethylene sulfoxide in the case of sulfoxides; sulfolane (tetramethylene sulfone) in the case of sulfones; and 1,3-propanediol and 1,4-butanediol in the case of diols. These compounds are widely used today in commercial buffers for applications such as the diagnosis of autism spectrum disorders and the amplification of other targets that are difficult to amplify, under license from Chakrabarti Advanced Technology, which is also the authorizing body of this application.

[0081] Another notable observation was that representative members of the cosolvent exhibited different behaviors with respect to different types of substrates, but their effects on polymerases were remarkably similar. This suggests that the mechanisms by which these solvents destabilize dsDNA and enzymes in general, particularly polymerases, are similar. In the case of dsDNA, this involves the relaxation of hydrogen bonds that maintain the double helix, and in the case of enzymes, the relaxation of hydrogen bonds that maintain their folded structure (which is involved in their activity). Therefore, the ability of a solvent to destabilize the secondary structure of DNA or enzymes is a transferable property, transferable from one solvent to another by multiplying by a solvent-specific coefficient. This coefficient can depend on various factors, particularly the geometric fit of the molecule within the complex three-dimensional structure (Chakrabarti, 2002).

[0082] By using the solvent intrinsic coefficients, in order to experimentally confirm conclusions regarding the transferability when transferring DNA destabilization data from one solvent to another, while using four different oligonucleotide dsDNAs (having 0%, 50%, 70%, and 100% GC content), for some of the representative members of the solvents in the inventors' list, melting point (T m ) depression experiments were conducted. The results (Example 14) clearly show that the melting point depression by any one of these solvents is directly proportional to its concentration, and thus is represented by a generalized equation containing a coefficient specific to each solvent, and is more or less independent of the target GC content. In the inventors' previous research (Chakrabarti in 2002), it was demonstrated that there is a direct correlation between the molar concentration of the different organic co-solvents herein and the decrease in the half-life of wild-type Taq polymerase at 72 °C and 95 °C (the decrease in t 1 / 2 ). In Example 15, the inventors, while using 1,4-butanediol as an example, demonstrated that just as the T m depression of DNA by an organic co-solvent is independent of the GC content of the DNA, the decrease in t 1 / 2 is independent of whether the polymerase is wild-type Taq or its variant with a specific mutation introduced. It is a self-evident truth that if X is directly proportional to Y and also directly proportional to Z, then Y should be directly proportional to Z. In this case, since both the T m depression of DNA and the T 1 / 2 decrease of DNA polymerase are proportional to the molar concentration of the solvent, the former functions (T m and T 1 / 2 ) should be in a directly proportional relationship with each other. This means that the findings of the T m depression of dsDNA should be generally transferable to the t 1 / 2 characteristics of the polymerase. Therefore, if a DNA polymerase resistant to one of the solvents on the inventors' list is developed, it will necessarily show a similar resistance, although not exactly the same, to the other solvents within the inventors' list.

[0083] Such experiments make the challenge of designing polymerases for mixed aqueous organic media easier. Since 1,4-butanediol was one of the solvents that exhibited DNA melting point depression that was approximately in the middle of the range of solvents that the inventors found to be effective PCR enhancers, the inventors selected 1,4-butanediol as a cosolvent for this purpose.

[0084] As previously described, the inventors used a combination of various techniques to design DNA polymerases that are particularly effective in the presence of organic cosolvents. Such techniques include directed evolution, CSR enrichment, NGS, gene synthesis and site-directed mutagenesis, and theoretical calculations (ΔΔG and ΔΔS). vib This included ( ), each of which is described in detail below.

[0085] Directed Evolution: Protein engineering involves manipulating amino acids at different locations within a protein to improve the stability and function of enzymes for in vitro applications. Directed evolution is the most widely used method to achieve this goal. The manipulation is carried out at the gene level of the protein, i.e., in the coding DNA. Directed evolution techniques most commonly rely on building a large library of variant genes through random mutagenesis (see below), followed by high-throughput screening, and selection to identify such members of a library encoding proteins with desirable properties. The process may be repeated several times until the desired level of performance is achieved.

[0086] Although the theoretical framework for directed evolution was established very early on (Eigen 1984), the power of the technique was first demonstrated in practice by Frances H. Arnold's group, who reported that directed evolution of the hydrolytic enzyme subtilisin E resulted in an active variant enzyme even in the highly unnatural (denaturing) environment of an organic aqueous medium (the organic solvent used was DMF). To evolve subtilisin E, they created and re-diversified a DNA sequence library through three rounds of random mutagenesis and screening, using error-prone PCR. The selection criterion used was the hydrolysis of the milk protein casein. The enzyme secreted by bacterial colonies was transferred to agar plates containing both DMF and casein. Active variants were detected on the agar plate in the presence of DMF by the halo formed by the variant together with casein. Plasmid DNA was isolated from clones secreting enzyme variants that produced a larger halo than the halo around the parent enzyme, and these were subjected to further rounds of mutagenesis. The final variant enzyme exhibited 256-fold higher activity than the wild type in 60% (v / v) DMF (Artnold 1993). This experiment (for which Arnold was later awarded the Nobel Prize in Chemistry in 2018) led to further exploration and development of the technology, which has since increased exponentially.

[0087] One of the essential components of directed evolution is the generation of diversity at the genetic (DNA) level. Various methods are available for this purpose. Two of the most commonly used methods for generating diversity are error-prone PCR (epPCR) and DNA shuffling (StEP PCR).

[0088] Error-prone PCR (epPCR) is the most common and efficient method used to introduce random mutations into genes. This involves performing a PCR reaction using a low-fidelity polymerase, i.e., a polymerase lacking proofreading or 3'→5' exonuclease function, in order to intentionally introduce copy errors. The degree of error can be further increased by adding modified nucleotides. The lack of proofreading function allows for the random misincorporation of nucleotides during the extension process. One drawback of this method is that, as is obvious, when used to develop diversity libraries for directed evolution, the numerous random mutations introduced (often as many as 95%) have proven to be less desirable than the start gene for the intended specific traits, and therefore the method is used in conjunction with a method that allows for easy isolation and / or removal of undesirable genes from desirable genes. For the same reason, if the desired pool is further diversified by epPCR after the desirable has been isolated from the undesirable, it can often be more detrimental than beneficial, as some of the better clones are inactivated.

[0089] DNA shuffling involves the digestion of a gene into random fragments by DNase I, and the reassembly of these fragments into the full-length gene, typically via primerless and modified PCR (Stemmer, 1994). The fragments act as primers for each other based on sequence homology, and recombination occurs when a fragment from one copy of the gene anneals to a fragment from another copy. Modifications to PCR involve alternating extension processes (StEP), in which the annealing and extension steps are significantly shortened to generate alternating DNA fragments and facilitate crossover events (shuffling or fragment switching) along the full length of the template sequence (Zhao et al., 2006).

[0090] By using numerous very short St-PCR cycles, adjacent segments within the fragment acting as primers are only extended by a few nucleotides until the final full-length gene sequence is generated. Unlike epPCR, which uses low-fidelity polymerases, StEP PCR prefers to use high-fidelity polymerases to avoid the excessive number of novel mutations that would result from the numerous repetitions of the StEP PCR cycle (approximately 150 cycles). DNA shuffling can also be achieved using restriction enzymes, in which case the fragments can be rejoined using DNA ligases.

[0091] DNA shuffling is a crucial technique for creating diversification for directed evolution experiments. Diversification results from combining useful mutations from two or more genes into a single gene. While the primary purpose of shuffling is to reconstruct existing mutations, avoiding the introduction of novel mutations is nearly impossible. In our case, when we performed StEP PCR, we found that mutations were introduced at almost every amino acid position, albeit at very low intensity. We considered this a beneficial phenomenon because the novel CSRs of the shuffled products encountered more diversity than planned.

[0092] StEP shuffling is a convenient method for generating chimeric libraries from two or more target sequences. In this invention, the inventors used DNA shuffling between two or more mutant polymerases (each containing two or more preferred mutations) to generate novel variant polymerase structures containing multiple preferred mutations.

[0093] epPCR and DNA shuffling are, among other things, the two most widely used methods for diversity generation, but other methods are also available to accomplish the same thing. Two such methods are as follows: a) Random priming in vitro recombination. This involves priming a template polynucleotide(s) using random sequence primers, and an extension step to generate a pool of short DNA fragments containing controlled levels of point mutations. The fragments are reassembled over a cycle of denaturation, annealing, and further enzymatic DNA polymerization to produce a library of full-length sequences (Shao et al., 1998). b) Saturated Mutagenesis: Saturated mutagenesis, also known as site-saturated mutagenesis, is a method used to prepare diversified libraries where a deeper examination of amino acid changes at any specific site or a predetermined number of sites is desired in directed evolution projects (Reetz et al., 2007). Here, a library is provided in which a single codon (or set of codons) is substituted with all possible amino acids, and all 20 native amino acids are contained at one or more predetermined sites. Saturation can be achieved by site-directed PCR using randomized codons in primers, or by artificial gene synthesis.

[0094] Selective Pressure: After creating a diversified library, the next major challenge in directed evolution is selecting selection criteria. These are the criteria that newly evolved enzymes are expected to satisfy. In the case of Francis Arnold's semen study on the evolution of subtilisin E, the selection criterion chosen by his group was the hydrolysis of casein in the presence of the organic solvent DMF, which is normally toxic to the wild-type enzyme. The selection criteria may differ depending on the desirability of the enzyme's performance on substrates(s) that may differ from its natural substrate(s) or in media different from its natural medium. In this case, selective pressure (high temperature and solvent) was applied during the PCR reaction of CSR (see below) and during the screening of products by RT qPCR. Selective pressure allowed only variants that developed "compatibility" to the new criteria through mutation to survive; other variants with less compatibility, including the wild type, did not survive under selective pressure(s) and disappeared from the colony. Selection pressure can be applied gradually over several steps, increasing in intensity step by step, with the goal of eventually reaching the final selection criteria. Diversification can be performed only once, either at the start or between selection rounds. When random mutation (epPCR) is used for diversification, as previously described, additional intermediate diversification steps may be detrimental rather than beneficial by inactivating some of the good variants through additional mutations. This may not apply to some other diversification methods (such as DNA shuffling of preferred sequences).

[0095] Two major challenges in directed evolution, therefore: a) The ability to generate robust diversification libraries that may contain variant genes with desirable mutations; and b) Creative methods to eliminate genes that are not very compatible and to recover genes that are better compatible to survive selective pressure. This includes the following: For the directed evolution of many enzymes, particularly DNA polymerases, compartmentalized self-replication (CSR) has been found to be especially suitable.

[0096] Selection using compartmentalized self-renewal (CSR): This selection technique, which uses a water-in-oil reversed-phase emulsion system, appears to be unparalleled suitable for the directed evolution of enzymes such as DNA polymerase. Phillip Hollinger and his collaborators in the UK were the first to successfully use this process to prepare variants of Taq DNA polymerase with improved thermal stability and heparin resistance compared to the native polymerase (Ghadessy et al., 2001, 2007). The best-performing variant they discovered, T8, contained six mutations: F73S, R205K, K219E, M236T, E434D, and A608V. The first four of these mutations (F73S, R205K, K219E, and M236T) clustered within the 5'→3' exonuclease domain, extending from position 1 to position 288. In this regard, it should be noted that the Taq variant lacking the exonuclease domain (i.e., the Stoffel fragment) exhibits improved thermal stability. These two facts suggest that the exonuclease domain of Taq polymerase is less heat-resistant than the rest of the enzyme's structure, or may be the source of thermal instability. T8 appears to express 5'→3' exonuclease activity impairment, as it possesses four of the six mutations within the thermally unstable exonuclease fragment.

[0097] W / O emulsions: CSRs rely on the fact that it is possible to prepare water-in-oil reversed-phase emulsions that can compartmentalize individual bacteria derived from colonies within emulsion droplets, thus enabling the maintenance of the association between genotype and phenotype. Next, the most important aspect of CSRs is designing a W / O emulsion that has millions of aqueous emulsion droplets in a continuous oily medium [in contrast to common oil-in-water (O / W) emulsions], where the droplets are not interconnected in any chemical sense. In the case of Ghadessy et al. (2001), a slightly modified emulsion described by Sweasy et al. (1993) worked well in their system, however, the process may be complicated by the introduction of other components, such as organic solvents similar to those described herein. The problems associated with emulsion technology stem from the fact that, in most cases, the precise science behind creating stable emulsions remains shrouded in secrecy, as the products are often very expensive, for example, when incorporated into drug formulations. Even the definitions of terms such as micellar solution (Hartley), nanoemulsions (Graves, 2004), microemulsions (Schulman, 1940), swollen micelles (Adamson, 1969), and miniemulsions (Ugelstad, 1973) are used in a mixed manner without any consensus among users. If any broad generalization must be made, it is based on the particle size of the dispersed phase, using the diameter of the emulsion droplet, which provides rough guidance regarding its thermodynamic and kinetic stability. It is believed that when the particle size of the dispersed phase is less than 1 μm in diameter, the emulsion is thermodynamically stable, while emulsions larger than that size are metastable and can exhibit varying degrees of kinetic stability (McClements, 2012).Of course, the point at which thermodynamic stability ends and kinetic stability begins is not as clear as the inventors understand herein, but emulsions with dispersed phase particle sizes in the range of 15 μm to 50 μm are empirically stable, especially in light of the purpose for which they are designed.

[0098] The properties of an emulsion depend on a variety of factors, including, but not limited to, the properties of the nonpolar (oily) phase, the composition of the aqueous phase, the properties and composition of the emulsifier used (which can be various types of anionic, cationic, or nonionic surfactants), and the relative amounts of these three components. Therefore, defining an emulsion based on a specific criterion is extremely difficult. Of course, things become even more complicated when certain small organic compounds, especially short-chain aliphatic alcohols, are added to the system. Such alcohols (monools with directional polarity) have their own merits, exhibiting only molecular solubility in both the aqueous and / or oil phases, and when added to an emulsion system, they act as auxiliary surfactants. In fact, such short-chain alcohols (like butanol and octanol) are important components of microemulsions and often determine whether the system tends to form a water-in-oil (W / O) emulsion or an oil-in-water (O / W) emulsion (Prince, 1977). The essential components of the emulsion systems described herein consist of a specific low molecular weight organic solvent, which is not necessarily a unidirectional polar molecule like a monool, but nevertheless belongs to one of four chemical structural groups (amide, sulfoxide, sulfone, and diol), making this phenomenon of particular interest to this specification. In the presence of such a solvent, emulsion stability is placed in uncharted territory. Such mixed solvent systems have never been tested before and therefore require some deeper consideration.

[0099] While various theoretical models primarily dealing with colloidal systems are well-known and their development is constantly ongoing, they are not particularly practical when designing stable emulsions from complex mixtures of components. One theory that seems to have been helpful in this regard was developed by P.A. Winsor in the 1950s (known as Winsor's "R-theory of solubilization"). It remains the most famous and easily understandable theory, encompassing all phases of emulsion systems, including W / O on one side, O / W on the other, and intermediate open-end (and connected) liquid crystal structures (Winsor, 1948-1960).

[0100] Winsor's R-theory: Winsor's theory examines the intermolecular processes of attraction (both electrostatic and electrokinetic phenomena) between surfactants, oils, and water. Electrostatic interactions are the actions between ions and dipoles and contribute to hydrophilic properties. H This is expressed as follows. Electrokinetic interaction is caused by the motion of electrons within a molecule, and it is the familiar van der Wall interaction (involved in the attractive forces between nonpolar substances, such as hydrocarbon molecules, and therefore contributes to hydrophobic properties). Its ID is A L This is how it is expressed. For example, within a unit volume of a two-component solution, molecular interactions can be expressed as follows:

number

[0101] Interaction A AA Or A BB This promotes the clustering of molecule A or molecule B, and ultimately the phase separation. Interaction A AB This promotes the mixing of molecules A and B. However, both of these interactions are concentration and temperature dependent.

[0102] Winsor begins by assuming an equilibrium between three types of micelles: layered micelles (liquid crystal structure), spherical Hartley micelles (oil in water), and spherical inverse micelles (water in oil). In a three-component system (surfactant, oil, and water), the ratio R is then defined as follows: R = (Tendency of a single layer of surfactant to be convex toward oil) / (Tendency of the same layer to be convex toward water) This is shown in Figure 4.

[0103] A crucial condition for the stability of a liquid crystal solution is R=1, meaning that the surfactant monolayer does not exhibit a tendency to be convex or concave towards its oily or aqueous environment. The tendency of the surfactant monolayer to be convex towards the oil phase is supported by the interaction between the surfactant and oil molecules, and inhibited by the interaction between oil molecules. Similarly, the tendency of the surfactant layer to be convex towards water is supported by the surfactant-water interaction, and inhibited by the water-water interaction. The change in R associated with the composition is therefore given by the following equation

number

[0104] In dilute solutions of surfactants dissolved in water (where oil is present in limited amounts), R decreases due to the mass effect, resulting in spherical micelles (O / W Hartley micelles or oil-in-water micelle solutions) with the polar side facing outward toward the outer aqueous phase. At high concentrations of surfactant (where both water and oil are present in limited amounts), R=1, and the liquid crystal structure is stable. If the surfactant concentration is even higher (where water is present in limited amounts), spherical inverse micelles (W / O emulsions) with the hydrophobic side facing outward toward the oil phase are formed. When low molecular weight alcohols (which are not surfactants themselves) are introduced under conditions that contribute to the formation of layered or liquid crystal structures, an interesting case arises where such alcohols act as auxiliary surfactants and align themselves within the surfactant monolayer to form swollen micelles (often called microemulsions). Depending on the precise chain length or hydrophobicity of the hydrocarbon tail of the alcohol, layered micelles can transform into microemulsions containing spherical oil-in-water or water-in-oil emulsion droplets. Alcohols with shorter chain lengths (C3-C5) tend to form oil-in-water microemulsions, while alcohols with longer chain lengths (C6-C5) tend to form oil-in-water microemulsions. 10 ) tends to form water-in-oil microemulsions. Other factors, such as the salinity of the water, temperature, the exact properties of the oil (aliphatic, aromatic, mixture, or other), and the properties of the surfactant (anionic, cationic, zwitterionic, nonionic, and their various structural forms), also play a role in determining the exact properties of the emulsion. Figure 4 shows a highly simplified schematic diagram of Winsor's R theory.

[0105] The effect of short-chain alcohols suggests that the organic solvents described herein should have a strong effect on the formation and stability of the O / W emulsions that the inventors are pursuing, but this does not provide any specific guidance. As far as the inventors' systems are concerned, ionic surfactants may have additional interactions with the inventors' CSR reactants, so the inventors are reluctant to consider all surfactants other than nonionic surfactants. The inventors' emulsion compositions, consisting of a hydrocarbon as a nonpolar phase, an organic aqueous medium as a polar phase, and a nonionic surfactant as an emulsifier, are novel compositions, and they had to be designed to form water-in-oil emulsions (where the contents of polar droplets (containing organic solvents or biomolecules) cannot be substituted and / or shared between them).

[0106] Emulsifier: The emulsifier found to be useful in creating the W / O emulsion of the present invention belongs to a class of surfactants called nonionic surfactants. It may consist of one or more molecules belonging to the chemical group shown in Figure 5.

[0107] The nonionic surfactants that can be used as emulsifiers in the present invention may be nonionic fluorinated surfactants as shown in Figure 6. Such surfactants differ from conventional nonionic surfactants listed in Figure 5 in that they have a hydrophobic tail (R'') made of fluorocarbon. Examples include fluorinated surfactants of the Novac® brand from 3M, and three of the most common representative members are listed below: ● 3M (registered trademark) fluorinated surfactant FC-4430; ● 3M (registered trademark) fluorinated surfactant FC-4432; and ● 3M (registered trademark) fluorine-based surfactant FC-4434

[0108] Oil: In the emulsion of the present invention, the oil acting as a continuous outer phase is a hydrophobic liquid having low to medium viscosity. The oil may be an aliphatic hydrocarbon, an aromatic hydrocarbon, or a mixture of both. A common type is a mineral oil with low to medium viscosity, which is a mixture of refined paraffinic and naphthenic hydrocarbons with a boiling point above 200°C.

[0109] A particularly useful mineral oil for this purpose is diesel fuel (which has a minimum viscosity of 15 cP at 40°C, a specific gravity of 0.85 at 25°C, and a flash point (closed cup) of approximately 215°C).

[0110] An interesting class of oils that can be used to prepare the emulsions of the present invention is synthetic oils. Of particular note among synthetic oils are high-boiling point fluorinated hydrocarbons (PFCs), or mixtures of PFCs and perfluoropolyethers (PFPEs). As an alternative to these conventional fluorinated synthetic compounds, there is an engineered liquid (Novac® 7500 liquid from 3M). Novac® 7500 liquid, used as an emulsifier in combination with a fluorinated surfactant, is particularly useful when the emulsion is prepared using the μEncapsulator from Dolomite Microfluidics, UK (see below).

[0111] Mechanical energy: When preparing emulsions, not only is the appropriate selection of oils, aqueous systems, and emulsifiers required, but the application of mechanical energy is also necessary to facilitate the dispersion of the internal phases within the continuous phase. This can be achieved by mixing the two phases with the emulsifier by stirring manually with a stirring rod, using a mechanical device such as a conventional magnetic stirrer or a motorized blade stirrer (the blades may be made from metal, Teflon, or glass), using a so-called Hersberg stirrer (a device consisting of a wire connected to one end of a rod, with the other end connected to a motor), or using many other forms of stirrers commonly used in the laboratory.

[0112] While the above-mentioned types of stirrers may suffice for most emulsification tasks, highly sophisticated equipment is required when a very homogeneous emulsion containing monodisperse droplets is desired. One such instrument is the μEncapsulator, sold by Dolomite Microfluidics in the UK. In this specification, the inventors have used the μEncapsulator to create monodisperse emulsions with droplet sizes in the range of 15–30 μm with great success (Figure 9).

[0113] Emulsion Stability: Emulsions must maintain their integrity and, even at temperatures considerably higher than room temperature, must not chemically communicate with each other (i.e., their contents must not be exchanged) at least up to the denaturation temperature in PCR. This means that emulsion droplets should preferably maintain their identity and compositional integrity at all temperatures from room temperature to 100°C. Theory can be helpful in designing such systems, but they must ultimately pass rigorous tests to demonstrate such integrity.

[0114] In this specification, in which the inventors used 1,4-butanediol as an organic cosolvent in their reversed-phase emulsion formulation, the inventors found an unexpected coincidence: the same oil-surfactant combination that worked in the cases of Sweasy et al. (1993) and Ghadessy et al. (2001), and also in the inventors' case, resulted in a stable water-in-oil emulsion containing spherical emulsion droplets that were not interconnected, despite a wide distribution of droplet sizes (Figure 7). Further trial-and-error experiments were conducted using slightly different combinations of surfactant and oil, as well as mechanical equipment (Dolomite's μEncapsulator® system) to mix and disperse different phases, after which a more uniform droplet size distribution was achieved (Figure 9).

[0115] The emulsion compositions described herein are complex, but this offers greater freedom in combining components (oil, emulsifier, organic / aqueous phase) and their relative proportions. Using this freedom, those skilled in emulsion science can develop two or more combinations that provide non-cooperative water-in-oil emulsion droplets suitable for the present invention. Two examples of such emulsions are shown in Figures 7 and 9. Such emulsions vary not only in the mixing method but also in the composition of the oil and emulsifier. These and other emulsions described herein are distinguished from other emulsions by combining different proportions of cosolvents, water, oil, surfactants, and other essential reagents (all within the constraints imposed on them). What makes such emulsions novel compositions is precisely their composition of oil, water, a specific organic solvent, a surfactant selected from a defined group of structures, and other essential CSR reagents.

[0116] Schematic diagram of CSR: Figure 10 shows an overview of the CSR process. A diversified library of Taq DNA polymerase genes is incorporated into E. coli, and the bacterial pool is added to a reverse-phase water-in-oil emulsion. Each E. coli containing only one variant pol gene is incorporated into a single aqueous compartment of the emulsion. Also included in the aqueous compartment are PCR buffer containing dNTPs, adjacent primers, and an organic co-solvent (as described separately herein). The PCR reaction is carried out here in such emulsion under selective pressure. In this specification, the selective pressure used is a combination of an organic co-solvent and a gradually increasing temperature, the latter applied at the beginning of each round of the PCR cycle. In the first step of PCR, i.e., denaturation (or including an additional step of thermal cell wall lysis), the applied heat breaks down the cell wall, and the released polymerase enzyme and coding gene cause self-replication within the emulsion droplet. Replication does not occur in compartments containing bacteria with incompatible DNA polymerases (inactive or insufficiently active DNA polymerase variants). Such polymerase variants that cannot replicate under selective pressure conditions are therefore removed from the amplified pool. The surviving progeny polymerase genes are released and recloned for another cycle of CSR. If desired, additional mutational diversification can be incorporated between CSR cycles. Polymerases derived from individual clones can then be ranked for their suitability to the selection conditions by appropriate means.

[0117] Enriched CSRs: CSRs placed under selective pressure are best suited to generating a pool of polymerase variants that can survive the selective pressure; however, this pool may contain certain preferred mutants present in very small amounts, making their isolation and characterization difficult. Performing a few more CSR rounds on a pool of better-fitting variants, without changing the selection conditions, can be helpful in enriching trace amounts of mutants through the amplification process. Therefore, it may be advantageous for enriched CSRs to follow or be followed by selective CSR rounds.

[0118] Directed Evolution of DNA Polymerases: Other Examples: CSR has now become a standard selection method in the directed evolution of DNA polymerases. Since the publication of its research by the Holliger group (Ghadessy et al., 2001), several other groups have used this technique to design polymerases with other properties. Below is a selection of patents and published materials describing DNA polymerase variants developed using directed evolution. In each case, the protocol was modified to meet the needs of the system and the requirements of specific selection pressures.

[0119] Arezi and his collaborators at Agilent Technology, California, described a method for developing polymerases for rapid PCR from Taq polymerase using the CSR method. They used random mutagenesis (epPCR) to generate the initial diversity pool. After performing five rounds of selective CSR while gradually shortening the PCR extension time, they performed multi-site targeted mutations on each of the top eight fastest cycling clones to develop a combinatorial library. The best-performing combinatorial mutants had 35–90 times higher affinity (lower K) to the primed template (549bp GAPDH gene). d), and showed a very small (2x) increase in elongation rate compared to wild-type Taq. The top three mutant Taq polymerases had the following mutations in order of performance (cycling time): #1: G59W, V155I, L245M, E507K''; #2 G59W, V155I, L245M; L375V, E507K, E734G, E749I; #3 V155I, L245M, E507K, F749I. Note that all three of the top mutants had two common mutations (V155I and E507K) (Arezi et al., 2014).

[0120] Many thermostable archaeal family B DNA polymerases possess a uracil-binding pocket within their N-terminal domain (acting as a "read-ahead," halting DNA replication upon approaching a uracil residue). While uracil is not a standard component of DNA structure, the high temperatures used in the PCR denaturation step often cause cytosine deamination, generating trace amounts of uracil. Although uracil formation is undesirable (reducing product fidelity), it does not have significant practical implications for many diagnostic tests that utilize PCR. However, interruptions due to polymerization halt (stoppage) reduce the usefulness of such archaeal family B polymerases used in many routine applications. Using CSR-based directed evolution, Tubeleviciute et al. (2010) successfully knocked in uracil-binding properties into the archaeal ShIB DNA polymerase (derived from Thermococcus litoralis). They generated diversity through random mutation (epPCR) and applied selective pressure to gradually replace dTTP with dUTP in the dNTP mixture. After a fifth CSR selection round in which dTTP could be completely replaced by dUTP in the PCR reaction, a ShIB polymerase variant containing mutant P36H, which lacks "read-ahead" (or uracil-binding) function, was selected. The results are interesting in light of the fact that mutations that partially (Y7A) or completely (V93Q) eliminate uracil-induced DNA replication stall in homologous Pfu DNA (implemented by site-directed mutation) (Connolly et al., 2009) are not effective in the case of ShIB, suggesting the power of directed evolution (Tubeleviciute et al., 2010).

[0121] Bourn et al. (filed by KAPA Biosystems, Massachusetts, USA, U.S. Patents 8,481,685B2 and 10,457,968B2) used directed evolution to develop variants of Kofu and Taq polymerases. They introduced random diversity into the genes by using epPCR. A notable feature of their study is that they did not introduce any selective pressure, but instead used several PCR rounds while using standard conditions with minor modifications to the buffer composition to adapt to standard changes commonly used in PCR amplification. The rationale was that natural polymerases like Taq are inherently designed to function under natural conditions, and the in vitro conditions for the PCR reaction constitute selective pressure themselves. They also reason that small changes introduced into chimeric polymerases like Kofu by combining the functional regions of two natural polymerases (KOD and Pfu) do not alter their preference for natural conditions. After several rounds of PCR, more compatible variants survive the in vitro conditions, while less compatible ones are eliminated. Next, they subjected the surviving clones to three initial phenotypic tests: a) enzyme activity (increased or decreased); b) DNA binding affinity; and c) fidelity. Based on these results, they also tested other phenotypic characteristics and identified variants suitable for different applications that exhibit superior compatibility with in vitro conditions. In this way, they found Taq variants with better salt tolerance, increased heparin binding affinity, and variants more suitable for amplifying long DNA substrates (2 kilobases or more) compared to the wild type (US Patent No. 10,457,968B2). They also isolated variants of Kofu polymerase with better DNA binding affinity and enzyme activity changes, fidelity, processing capacity, extension rate, and stability compared to the parent Kofu polymerase (US Patent No. 8,482,685B2).

[0122] Combinations of Directed Evolution and Other Techniques: In some embodiments, next-generation sequencing (NGS), also known as deep sequencing, and gene synthesis were used to enhance the size and quality of our variant sequence pool. The unique combinations of these techniques, and the methods of using them, constitute novel approaches that we pursue to achieve our selected objectives. These will become apparent throughout this specification as they are discussed.

[0123] Next-generation sequencing (NGS): Also known as deep sequencing, NGS is a high-throughput sequencing method. It involves sequencing genomic regions multiple times (often thousands of times). This process allows researchers to detect rare clonal types that occupy as little as 0.1% of the original sample (Mardis, ER, 2011).

[0124] The fundamental principle of sequencing in NGS is the same as that of the linkage arrest sequencing method developed by Frederick Sanger (Sanger et al., 1977), except that NGS is a high-throughput method that, in the case of large DNA segments, cuts them into smaller pieces and then sequences hundreds of thousands of these fragments simultaneously in a massively parallel manner. Companies such as Qiagen, ThermoFisher, and Illumina each offer their own unique high-throughput sequencing platforms in various forms, although Illumina's platform, which uses its proprietary synthetic sequencing (SBS) platform, is the most popular.

[0125] Sanger sequencing utilizes 3' blocker chemistry. The method is based on performing a PCR reaction to amplify a gene, with the exception of introducing a chain termination nucleotide (ddNTP) into the reaction mixture, by adding it to the usual components (i.e., a set of primers, DNA polymerase, dNTPs, and standard PCR buffer). In the PCR chain elongation reaction, chain elongation occurs at the 3' hydroxyl group within the deoxyribonucleotide portion at the head of the growth chain. The ddNTP molecule lacks a 3' hydroxyl group; therefore, whenever a ddNTP molecule is introduced during chain elongation, the resulting chain cannot elongate any further. In a typical Sanger analysis, the DNA segment to be analyzed is amplified in five parallel tubes. One tube contains the usual PCR reaction mixture. Each of the other four tubes contains, in addition to the usual mixture, one of the four ddNTPs (ddATP, ddTTP, ddCTP, and ddGTP). After the amplification reaction is complete, the product is electrophoresed on a standard agarose gel that separates DNA molecules by molecular weight. When such a gel, derived from a ddNTP tube, is compared to a ddNTP-free gel, the positions of A, T, C, and G within the strand can be determined. As you can see, this is a redundant process and is far from high-throughput.

[0126] NGS possesses two extremely important features that make it high-throughput.

[0127] Firstly, instead of using standard ddNTPs, NGS uses fluorescently tagged ddNTPs, in which case each ddNTP (ddATP, ddTTP, ddCTP, and ddGTP) has a different fluorescent tag associated with a 4-pass / band-filter camera / sensor that records all nucleotide addition events for all four nucleotides. Newer versions use reversible fluorescently labeled dNPPs. The use of fluorescently labeled dNTPs (each with a different emission wavelength) also eliminates the need to perform four different reactions and to read chain arrest sites based on the gel.

[0128] A second high-throughput feature of NGS is the amplification of DNA on a solid surface of the flow cell, often called a chip. In the NGS Lumina platform, the DNA to be analyzed is broken down into small fragments (called amplicons) up to 500 nt in length. These DNA fragments are then spread onto the two-dimensional surface (chip) of the flow cell and ligated to it with the help of special small DNA molecules called adapters. Subsequent reactions are carried out on this surface.

[0129] There are five basic steps to developing an Illumina platform for NGS. 1. DNA fragment / DNA sample - Amplicon preparation. For NGS, long DNA fragments must all be randomly cut into smaller pieces for the amplicon library, with each segment being 500 nt or less in length. Such fragments can be generated by PCR using a duplicate primer set. The quality, size, and purity of the amplicons are crucial in determining the quality of the final NGS results. 2. Linking adapter molecules to DNA fragments: Adapters are small DNA molecules linked to both ends of a single-stranded DNA fragment using DNA ligation chemistry. These act as sticky ends of the fragment for hybridization into short complementary DNA on the flow cell (see next step). 3. Immobilization of the flow cell and short DNA segments. A pool of short ssDNA segments complementary to the adapter DNA molecule is anchored (immobilized) on the surface of an 8-channel flow cell. One end of such a molecule is anchored, and the other end is free. This acts as a primer in PCR extension when performing "bridge amplification" to form a cluster in the next step. The result is a loan of immobilized oligomeric DNA primers on the flow cell surface. 4. Cluster Generation / Bridge Amplification: The single-stranded amplicon containing the adapter is added to the flow cell. It hybridizes with its complementary oligo on the flow cell surface, which has a free 3' end at its adapter end. Using high fidelity DNA polymerase, the free 3' end of the hybridized oligo (acting here as a primer) is isothermally extended, thereby forming a full-length copy of the amplicon (anchored on the flow cell surface). This copy also copies the adapter molecule from the undhybridized end of the template amplicon. The amplicon template is separated here by denaturation. The newly formed DNA molecule wraps around (folds around) the surrounding tissue, and its free end (containing the adapter copy) hybridizes on the cell surface with another anchored complementary oligo, forming a bridge between the two immobilized oligos, with the formation of an inverted U-shape. As the loop elongates, another copy of the full-length amplicon with an adapter end is produced, and thus another inverted U-shape can be formed, linked to two other anchored complementary oligonucleotides after denaturation from the anchored loop. This process repeats itself until hundreds of thousands, if not millions, of looped copies of each template are formed. This is bridge amplification, and the multiple copies of the template amplicon become clusters of identical DNA. Thousands of such clusters are formed around the thousands of added amplicon DNA. Each cluster of the dsDNA bridge undergoes chemical denaturation, and the reverse strand is removed by specific base cleavage, leaving the forward DNA strand. The 3' ends of the DNA strand and cell-binding oligonucleotides are blocked to prevent interference with the sequencing reaction in the next step. 5. Sequencing Reaction / Synthetic Sequencing: Illumina's synthetic sequencing technology does not use ddNTPs as terminator nucleotides. Instead, the technology uses a method based on Illumina's proprietary reversible terminator, employing a total of four fluorescently tagged dNTPs, each with a unique emission wavelength. The reversible terminator property of the nucleotides means that only one base can be added at a time. A camera records the addition of each fluorescently labeled nucleotide—the emission wavelength and intensity are used to identify the base. The cycle is repeated n times to create a reading length of n bases. Sequencing is a fully automated operation, and once the process is started, there is little the operator can do. The actual process involves some more information regarding washing and reagent addition between steps, as well as other details that are confidential and proprietary by the technology provider. 6. Computational Analysis / Alignment / Data Analysis / Quality Score / Base Calling / Mutation Calling. The output from the sequencer is a set of "reads," the length of which depends on the specific platform used. Illumina platforms offer two or more read options, e.g., HiSeq, MiSeq, etc. In our research, we used MiSeq with a read length of 250 bp. A large number of reads, around 100,000, can be obtained from a single operation. The reads are raw data and therefore cannot be used directly without further transformation. Transformation is performed using bioinformatics software, which many companies keep as confidential proprietary information. The software aligns the reads to a reference sequence and identifies their unique sequences. This allows for the identification of single nucleotide polymorphisms (SNPs) or insertion-deletion (indels) within the reads and their frequency of occurrence. Occurrence can be ranked in relation to their frequency. The computer program assigns a quality score (Q score), called a Phred score, to each identified base. A higher Phred value indicates better quality prediction regarding base identity. Theoretically, a Phred score can range from 0 to infinity. However, in practice, an upper limit is set by the platform's reliable detection limit—in Illumina's case, this limit is 40. A Phred score of 10 means the probability of an incorrect base calling is 1 in 10, a score of 20 means it's 1 in 100, a score of 30 means it's 1 in 1,000, and a score of 40 means it's 1 in 10,000. Filters can remove scores below a certain number. Therefore, by setting the filter to 20, all scores below 20 can be blocked, thereby reducing the probability of incorrect calling to less than 1% for all base callings.

[0130] The inventors' objectives in using NGS consisted of two parts. The first was to identify and confirm mutations within variant genes that had survived selective pressure. This objective was to confirm mutations found within variant genes detected by qPCR screening of CSR products. This is not an insignificant task, as most molecular biology experiments are unlikely to be repeated to obtain precise details, and therefore alternative methods for confirming the results play an important role. The second objective of NGS was to detect mutations that may have been missed in standard screening processes. Such lost mutations can be particularly interesting if they are significantly enriched by CSR and may play some important role in enhancing the compatibility metrics that the inventors are striving to improve. The rarity of a mutation does not necessarily mean that it will be beneficial in the future. In fact, many of them may actually be detrimental to achieving the compatibility that the inventors are pursuing. The objective is, of course, not to miss rare mutations that may have beneficial properties. Conducting NGS experiments, however, is not as simple as the inventors have recognized. Because the process requires sophisticated equipment and specialized operators, the inventors used an external service provider, GeneWiz, to carry out part of the work. The inventors chose GeneWiz partly because its Amplicon-EZ Service uses the Illumina Technology Platform, which employs synthetic sequencing for NGS.

[0131] The process works roughly as follows: We begin with a variant pool of polymerase genes after several rounds of selected CSRs, and often after further CSR enrichment cycles on the selected products. We prepare small duplicate fragments (each 450–468 nt in length) of the variant pool using PCR and a suitable set of primers. Our service provider, GeneWiz (using the Illumina Technology Platform), then links them via ligation to 5' and 3' adapter molecules derived from the adapter pool. GeneWiz then performs the remaining NGS steps, namely ligating the amplicons onto the flow cell surface, generating clusters, and finally sequencing them using Illumina's synthetic sequencing method. The company provides us with the “reads” which we analyze using our proprietary software. The output of the analysis includes not only identification information for surviving mutations, but also their location and frequency of occurrence, expressed in a ranked format (expressed as the incidence rate (percentage) of a given mutation at an arbitrary site compared to the total number of mutations occurring at that site). Special software then converts the nucleotide positions in the gene sequence to the corresponding amino acid positions in the enzyme.

[0132] One problem with NGS is that it only provides a list of mutants and their locations, and in most cases provides little to no information about their presence in any particular gene sample. As observed in this invention, when mutations were detected at a particular location (and survived the selective pressures of organic solvents and temperature) by any method (whatever it may be) used herein, the substitution of existing amino acids was almost always due to only one specific amino acid out of 19 possible ones. However, some exceptions existed. NGS proved useful in detecting and confirming such rare exceptions.

[0133] When mutations detected by NGS are identical to those found within variant genes identified from selected CSRs via Sanger sequencing, it helps to confirm and reinforce the CSR selection data. However, when they are added to those found within selected sequences, it becomes particularly interesting, and the inventors here need to determine how the sequence order can achieve the beneficial properties (some of which are beneficial, if so) that we seek. The latter is a challenging task, and here the inventors used gene synthesis methods to prepare novel sequences, either by constructing novel polynucleotide chains nucleotide by nucleotide using conventional gene synthesis methods, or by site-directed mutagenesis of the parent polymerase (Taq in this case). Furthermore, the inventors used the enzyme's ΔΔG and ΔΔS to determine whether a particular mutation is beneficial or detrimental to its stability (robustness). vib The value (see below) was calculated.

[0134] Gene synthesis and site-directed mutagenesis: CSR yielded variant genes in which certain mutations were specifically arranged and configured within each variant. Shuffling variants derived from CSR yielded novel sequences in which the number and arrangement of mutations within the variant were reconfigured within a single sequence. NGS primarily yielded lists of desirable point mutations. In this specification, further diversified sequences were constructed by conventional gene synthesis or site-directed mutagenesis. For this purpose, the starting point was a list of mutations and their locations derived from CSR and NGS. To obtain novel sequences for phenotypic testing and to provide new insights into the importance of point mutations and their advantages or disadvantages when combined, such positional point mutations were arranged and configured in any desired number within a single sequence. Thus, sequences with highly desirable properties that could not be isolated using CSR or shuffling alone were generated.

[0135] In the case of conventional gene synthesis, the work was outsourced to the service provider GenScript. In the case of site-directed mutagenesis, the inventors either used their in-house capabilities or used an external service provider such as GenScript.

[0136] Various methods exist for performing site-directed mutagenesis (also known as site-specific mutagenesis or oligonucleotide-directed mutagenesis), and newer or modified methods are constantly being developed. Those skilled in the art are familiar with such developments. In one simple and original scheme, the method uses custom-designed primers to introduce a desired mutation at a specific site within a double-stranded DNA plasmid. It is a powerful technique for practically introducing any mutation (including single nucleotide substitutions, short deletions, or insertions) at any site. The basic concept is described below, although only an example is presented.

[0137] The target gene (in this case, the DNA polymerase gene) is first cloned into a single-stranded vector, such as phage M13. Next, oligonucleotide primers are chemically synthesized that are sequentially complementary to the cloned gene located at the desired mutation site, except that the primers contain one or two planned mismatches near the center presenting the desired mutation to be incorporated into the gene. The primers undergo annealing and extension by PCR, and the extended strands are ligated to form a circular loop. This double-stranded plasmid is cloned into bacteria to generate multiple copies of the gene containing the desired mutation. The method can be used to introduce multiple mutations into the same gene (Mathews et al., 1999). It should be noted that the above is only one approach. Other approaches for site-directed mutagenesis are available and are well known to those skilled in the art.

[0138] Construction of a list of point mutations that may help increase the suitability of DNA polymerases to organic aqueous media: Taq polymerase variants that survived the CSR selection process and ranked well in real-time qPCR screening inherently possessed better suitability for performance in organic aqueous media; however, it is difficult to confidently say that all point mutations within such variants positively contribute to such suitability (see also below). Some point mutations within such variants did indeed have a negative contribution to polymerase suitability, but this negative contribution may have been offset by other strong positive contributions. CSR enrichment and DNA shuffling can only improve the opportunity to detect variants and enhance variant diversity, respectively, and they were not able to identify point mutations that generally made a positive contribution to the enzyme suitability index.

[0139] Gene synthesis and / or site-directed mutagenesis provided not only means for creating new genes but also means for verifying the inventors' hypothesis about which mutations and their combinations are most beneficial. One method the inventors used with a synthetic approach was to verify a list of single mutations that are highly useful with respect to the inventors' desired properties. By doing so, we can confidently say that every rearrangement and combination of such mutations is beneficial, and that this eliminates the need to create a very large number of sequences to prove whether all rearrangements and combinations are beneficial.

[0140] Another approach used herein to eliminate mutations that may have a negative contribution to enzyme stability is to remove the enzyme's ΔΔG value or ΔΔS value. vib The task was to calculate the effect of mutations on [the subject]. This will be described in detail in another section below.

[0141] It was of great importance to generate a list of mutations that could generally make a positive contribution to the enzyme's compatibility index without having a detrimental effect on the enzyme's stability. For the first time in the history of PCR, we had the opportunity to randomly combine point mutations that could not only increase the compatibility index but also increase the enzyme's stability, particularly in organic aqueous media. It can be pointed out again that adding another mutation to a promising combination does not, and will not, constitute a new useful composition, as the sequence to which the new mutation is added already possesses a robust positive contribution that can offset any negative contribution from the new addition, and thus may still have positive performance.

[0142] In short, this specification uses various molecular biology techniques (epPCR, DNA shuffling, CSR selection, CSR enrichment, real-time qPCR screening, NGS, gene synthesis, site-directed mutation, ΔΔG / ΔΔS) to generate a list of point mutations that can only positively contribute to the suitability index of Taq DNA variants for performing in the organic aqueous media described herein. vib We have uniquely combined value calculation (see below) and phenotypic testing.

[0143] ΔΔG and ΔΔS resulting from point mutations vib Directed evolution is an optimization process that seeks to improve the overall suitability of an enzyme to environments different from the environment in which it evolved (or was designed). By applying specific selective pressures (solvent and temperature), the inventors induced evolution to achieve better thermal stability in the presence of certain organic solvents, but it was impossible, and was not attempted, to limit the evolution to just one such dimension. This is because optimization is necessarily a multidimensional task. In this case, optimization, in addition to stability at higher temperatures and in the presence of solvents, includes, to name just a few, enzyme activity, DNA binding affinity, processing capacity, ability to amplify long templates, and extension rate (V). maxThis means improvements in properties such as nucleotides per second (nucleotides / second) and fidelity. It is the combination of these properties that brings about a better overall compatibility of the selected variant with the organic aqueous media described herein. This means that the selected variant will likely have mutations that improve some of these other properties (even if this comes at the cost of providing the best possible thermal stability in the presence of the solvent). Thus, a variant sequence containing a combination of mutations that best survives the selection criteria is essentially a novel composition, either alone or in the presence of the organic aqueous media described herein.

[0144] In this specification, the inventors describe a Taq variant library that passed real-time qPCR screening performed using a special solvent / temperature filter. Clones were selected that best met the inventors' selection criteria (high thermal stability in organic aqueous media) and still yielded optimized compatibility characteristics. As described herein as appropriate, the inventors developed these variants by both directed evolution and synthesis.

[0145] While preferred variants were selected as described above, the inventors' goal was also to identify a list of these mutations, at least one of which must be present in any of the variants to satisfy the inventors' most important requirement: high thermal stability in organic aqueous media. No mutation that destabilizes the enzyme can be included in this list, even if its presence in the variant could add other characteristics of overall compatibility. To achieve this goal, the inventors employed a theoretical approach, namely, the change in folding free energy (ΔΔG) and the change in vibrational entropy (ΔΔS) when a specific mutation is introduced into the parent polymerase. vibThe advantages of calculating ) had to be incorporated. The inventors considered both indicators because each had its own merits and demerits, and the confidence level in the inventors' conclusions increased significantly when both indicators suggested substantially the same conclusions.

[0146] While the folded structure of enzymes is conventionally represented by a static structure, they are actually highly dynamic molecules, and the ability of enzymes to exhibit various dynamic conformations within their overall folded structure is crucial for their effective catalytic function. Within this flexibility, we also explore the robustness of the structure appropriate for stability (in thermodynamic terms, meaning a decrease in its folding free energy (Gibbs free energy) (<0), or a decrease in its vibrational entropy, both of which are readily applicable to computational calculations derived from fundamental molecular forces).

[0147] Structural robustness, or change in folding free energy (ΔΔG expressed in kcal / mol), or change in vibrational entropy (ΔΔS expressed in kcal / mol / K). vib While various known approaches exist to determine the effects of point mutations on protein dynamics and stability (each with its own merits and weaknesses), the inventors used DynaMut (a user-friendly and freely available web server (http: / / biosig.unimelb.edu.au / dynamut)) to analyze the effects of point mutations on protein dynamics and stability. DynaMut is an integrated computational method that uses two approaches (Bio3D and ENCoM) to carry out its operations (Rodrigues et al., 2018). This method has been tested and has been successful in explaining the effects of mutations in hardening (stabilizing) protein structures such as SIR2 enzymes, while also improving their catalytic function (Ondracek et al., 2017).

[0148] In summary, in light of our objectives, for a mutation to be included in the list of selected point mutations in Taq polymerase, it must pass the following two simultaneous tests: a) Firstly, the mutation must belong to a variant that can pass real-time qPCR screening performed with the application of selection pressure and several filters; and b) the mutation must decrease its Gibbs free energy to below that of the wild type (typically ΔΔG < 0), and its oscillatory entropy must not increase to above that of the wild type (typically ΔΔS vib <0). The first criterion is that the variant enzyme, even if it contains a specific mutation in its sequence, does not prevent it from achieving the overall compatibility index. The second criterion is to ensure that the variant enzyme has a positive effect on the enzyme's stability.

[0149] At the present level of knowledge of the inventors, ΔΔG or ΔΔS vib This calculation is only calculable for proteins in their natural (aqueous) environment, not in the organic aqueous media described herein. However, there is ample evidence that structural robustness, which translates to thermal stability, is a property that can be transferred from heat to solvent. Thus, enzymes engineered to achieve thermal stability have been shown to be resistant to organic solvents as well, as has been found in the case of lipases, sucrose phosphorylases, haloalkane dehalogenases, kanamycin nucleotidyltransferases, etc. (Reetz et al., 2010; Koudelakova et al., 2013; Liao, 1993). This transferability does not mean that in our case it can be completed solely by using temperature as the selective pressure. For one thing, the strictness of the selective pressure required for our purposes cannot be applied to temperature alone (boiling point constraint). On the other hand, optimizing suitability is a media-specific endeavor. However, the opposite suggestion, namely using evolutionary variants obtained from this experiment conducted only on heat resistance (without solvent), should work quite well.

[0150] How many mutations are there per gene? As previously pointed out, if all amino acids in a polymerase protein were substituted with 19 other possible amino acids, the number of variant enzymes would be staggeringly large, reaching billions. However, for CSR selection products, CSR enrichment, NGS, and theoretical studies (ΔΔG and ΔΔS) are being considered. vib By applying the calculation of values, the inventors identified a very limited number of unique mutations (shown in the examples) that are particularly suitable for carrying out PCR reactions in an organic aqueous medium in the present invention. Therefore, the mutations described in the present invention are not limited to a specific position, but also to a specific single unique amino acid that replaces an existing amino acid at the selected position. A very few exceptions to this rule were found, however, the number of possible amino acids that replace an existing amino acid at that position did not exceed two. Examples of such exceptions are as follows: ●D244 to D244E and D244V ●F413 to F413S and F413L ●A454 to A454E and A454L ●V586 to V586A and V586M ●H767 to H767L and H767R ● D732 to D732G and D732N ●E832 to E832K and E832X

[0151] In some embodiments, the maximum number of mutations within any given enzyme variant was 10. To demonstrate that point mutations can be randomly combined, various combinations of intrinsic mutations within a single gene were synthesized and tested, and it was found that the desirable properties expected from the combinations were generally preserved. Mutational loads exceeding 12 are considered undesirable based on considerations other than their individual contributions.

[0152] Exemplary variant polymerases are described in the table and examples below.

[0153] The designed Taq variant was developed to eliminate the shortcomings of the wild-type polymerase when used in the artificial organic aqueous media described herein, but its usefulness is by no means limited to such media. Rather than being limited to organic aqueous media, the Taq variant is inclusive for both standard aqueous and organic aqueous media. In this sense, such an evolved polymerase is far more versatile than its parent in in vitro applications of PCR reactions.

[0154] Other parent DNA polymerases besides Taq: In our experiments, we used Taq DNA polymerase as the parent to design variants that did not have the drawbacks of the parent when used in organic aqueous media. We determined a list of amino acid positions in the parent of 834 amino acids in which specific mutations resulted in the benefits we sought (primarily superior thermal stability in organic aqueous media without sacrificing other desirable properties of the parent when used in PCR amplification of genetic material). We assert that the principle of "corresponding positions" (determined by using 3D alignment of crystal structures) can be used to extract variants from other DNA polymerases that confer the same or similar beneficial performance in organic aqueous media.

[0155] Applications: Variant Taq DNA polymerase or other variants derived from the other parent polymerases listed above can be used in a variety of PCR amplification processes, including, but not limited to, i) standard PCR; ii) hot-start PCR; iii) touchdown PCR; iv) nested PCR; v) reverse PCR; vi) arbitrary primed PCR (AP-PCR); vii) RT-PCR; viii) RACE (rapid amplification of cDNA ends); ix) differential display PCR (DD-PCR); x) multiplex PCR; xi) Q / C PCR (quantitative / relative PCR); xii) inductive PCR; xiii) asymmetric PCR; xiv) in situ PCR; xv) TaqMan assay; xvi) quantitative PCR using SYBR Green; xvii) cold PCR (co-amplification at lower denaturation temperatures); xviii) error-prone PCR; and xix) NGS. Those skilled in the art should be familiar with these terms. Definitions can also be found in relevant books or on Google.

[0156] The products of the present invention can also be used in kit format for many of the above applications. The kit may contain buffers, organic solvents, dNTPs, primers, and other PCR reaction components in a suitable packaging form.

[0157] The primary objective of this specification is to provide a designed DNA polymerase, particularly one with better thermal stability, that is highly compatible with functioning in mixed organic aqueous media. In this specification, the inventors achieved their objective by a) identifying variants of existing polymerases (in this case, Taq DNA polymerase) via directional evolution based on CSR, and b) identifying specific mutations that can primarily provide tolerance to solvents and higher temperatures. To achieve the second objective, the inventors had to uniquely combine various techniques. This specification presents the largest, most multifaceted, and comprehensive trials ever conducted to develop DNA polymerases for artificial media. The various claims presented herein are the result of this multifaceted approach to solving complex problems. The following examples are provided to illustrate this unique endeavor. [Examples]

[0158] The following examples are incorporated to provide guidance to those skilled in the art for practicing representative embodiments relating to the subject matter of this disclosure. Given the general level of this disclosure and that of those skilled in the art, they will recognize that the following examples are intended to be illustrative only, and that a great many changes, modifications, and alterations may be adopted without departing from the scope of the subject matter of this disclosure. The following synthetic descriptions and specific examples are intended for illustrative purposes only and should not be considered in any way as a limitation on the preparation of the compounds of this disclosure by any other means.

[0159] Example 1 Preparation of a variant library of Taq DNA polymerase by epPCR and generation of expression cells. a) When preparing the diversity library of Taq DNA polymerase, codon-optimized WT Taq polymerase (synthesized by Genscript, NJ, USA for expression in E. coli) was used as the parent enzyme. Error-prone PCR (epPCR) was used to create the initial diversity library. For this purpose, the inventors used a diversity epPCR kit (purchased from Takara BIO USA Inc., CA, USA) and followed the manufacturer's recommended procedure. The thus generated epPCR-diversified gene library was digested with DpnI, followed by column purification using a PCR purification kit from Qiagen. The purified product was digested with XbaI and SalI, and then ligated into an XbaI and / SalI-digested pASK-IBA5C vector. The ligated product was electroporated and incorporated into E. coli TG1 cells. After recovery for 1 hour, 5 μL of cells were serially diluted and spread onto an LB-chloramphenicol (50 μg / ml) plate to evaluate the library size. b) To generate expression cells (see definition), the transformed library was inoculated into LB-chloramphenicol (20 mL) in a 50 mL conical flask by shaking overnight at 250 RPM and 37°C. After proliferation, the cell suspension (0.5 mL) was inoculated into LB-chloramphenicol (50 mL). The cells were induced with anhydrotetracycline (300 ng / ml) and OD 600 Taq polymerase was expressed until the pH reached 0.4-0.5 (approximately 4 hours). After 4 hours, cells were collected by centrifugation, washed, and resuspended in 1× Taq buffer [10 mM Tris-HCl (pH 8.5) containing 50 mM KCl, 1.5 mM MgCl2, and 1% Triton X-100].

[0160] Example 2 CSR selection experiment: partitioned self-replicating This experiment was conducted using various selective pressures. A schematic of the process is shown in Figure 10. Three examples are presented below. a) As a selective pressure, exposure to 5% 1,4-butanediol at 95°C for 6 minutes, or at 98.3°C for 1 minute and then at 95°C for 6 minutes:

[0161] The selection pressure was established as follows: In the presence of 5% 1,4-butanediol, wild-type Taq DNA polymerase was found to survive and be unable to produce any PCR product. In the absence of 1,4-butanediol, it survived (Figure 11). Based on several preliminary CSR experiments, the inventors further demonstrated that a pre-PCR exposure temperature of 98.3°C, and residence times of 1 minute at 98.3°C and 6 minutes at 95°C, provide a strict selection pressure for any Taq variant to survive in 5% 1,4-butanediol. Under these conditions, WT Taq does not survive, and only variants with robust resistance to 1,4-butanediol survive.

[0162] This experiment used a reversed-phase emulsion. In addition, a negative control using the same composition (but without dNTPs) was measured in parallel with the main experiment.

[0163] The emulsion was pre-incubated at 98.3°C for 1 minute and 95°C for 6 minutes (as a selective pressure and to lyse the cell wall), followed by CSR PCR. For CSR, 25 cycles of PCR were performed, using the following conditions for each cycle: denaturation at 94°C for 1 minute, primer annealing at 55°C for 1 minute, and chain extension at 72°C for 5 minutes. The primer set used in CSR PCR was as follows: Forward: CAGGAAACAGCTATGACAAAAATCTAGATAACGAGGGCAA (Sequence ID 6) Reverse: GTAAAACGACGGCCAGTAGCTTAGTTAGATATCAGAGACCATGGT (Sequence No. 7) After CSR-PCR, the reaction mixture was extracted with diethyl ether (to remove light petroleum and organic solvents), and the residue was purified using Qiagen's PCR purification kit. After gel purification of the PCR product, the following primer set was used: Forward: GAATAGTTCGACAAAAATCTAGATAACGAGGGCAAAAAATG (Sequence No. 8) Reverse: CCTG CAGG TCGA CTTA TTCT TTCG CGCT CAGC CAGTC (Sequence ID 9) The DNA was then re-amplified using high-fidelity Q5 DNA polymerase.

[0164] The re-amplified product was digested with XbaI and SalI, and then ligated into a pASK vector digested with the same restriction enzymes. The ligated product was transformed and seeded on LB-chloramphenicol petri dishes. Individual colonies were isolated and grown in 96 deep-well plates to rank them for their thermal stability and tolerability to selected organic solvents by screening using a real-time qPCR-based method, as shown in Example 5 below. b) As a selective pressure, exposure to 7% 1,4-butanediol at 98.3°C for 1 minute and at 95°C for 6 minutes:

[0165] In this experiment, a reversed-phase emulsion was used. In addition, a negative control using the same composition (but without dNTPs) was measured in parallel with the main experiment. For the CSR reaction, the conditions were the same as those in Example 2(a).

[0166] Example 3 CSR enrichment CSR enrichment experiments were performed only on CSR selection products (Examples 4a and 4b). As previously described, purified variants of the Taq DNA gene were incorporated into novel E. coli cells to prepare new expression cells (Example 1b). The procedure for the CSR enrichment experiment was the same as that used in Example 2. The product recovery and purification steps were also unchanged. This is called enriched CSR because the inventors did not use any further diversification or apply any more stringent (or novel) selection pressures during this CSR period.

[0167] Example 4 Preparation of a Taq DNA polymerase variant library by DNA shuffling (StEP PCR) DNA shuffling using alternating extension process PCR (StEP PCR) was used to further diversify the top-ranking Taq polymerase variants selected in Examples 2(a) and 2(b). This process is designed to provide additional diversity through shuffling mutants present in the start sequence and to generate novel sequences (some of which are thought to contain more mutants per sequence than those in the start sequence). Representative examples are presented below.

[0168] Nine top-ranking clones derived from Example 4(a) [CSR selection in 5% 1,4-butanediol], each containing one to four mutations, were shuffled. The mutations within each clone were as follows: Y116 stop; E832K; L365P; G12T-A61V-2494delGA; K206Q; T186A; D244V-K314R-V586A-S612R; P10S; and A54V.

[0169] Plasmids isolated from these clones were restricted digested with XbaI and SalI to generate StEP templates. The reaction mixture in 1×Thermopol buffer consisted of equimolar amounts of each fragment (total 0.15 pmole), 250 μM dNTPs, 1.5 units of Vent polymerase, and 25 pmole of the following primers (5'→3'): GAAT AGTT CGAC AAAA ATCT AGAT AACG AGGG CAAA AAAT G(41nt) (SEQ ID NO: 8) CCTG CAGG TCGA CTTA TTCT TTCG CGCT CAGC CAGT C(37nt) (SEQ ID NO: 9) It was composed of.

[0170] The PCR extension protocol was as follows: initial denaturation at 95°C for 5 minutes; 150 cycles of [95°C for 1 second; 55°C for 5 seconds; 72°C for 2 seconds]; and final extension at 72°C for 2.5 minutes. The PCR product (composition after shuffling) was treated with DpnI, precipitated with sodium acetate, digested with XbaI and SalI, and cloned into a pASK vector for the next round of CSR.

[0171] Example 5 Real-time qPCR screening of evolved clones (transformed organisms): List of top-ranked clones (hit clones) A real-time qPCR assay based on SYBR GREEN I was used to screen transformants obtained after CSR selection and CSR enrichment, as shown in Example 2 and / or 3.

[0172] Transformed colonies were isolated and inoculated into a 96-deep-well culture plate containing LB-chloramphenicol medium (500 μL). The cells were grown and then OD (Oral Dissociation). 600When the pH reached 0.4-0.5, polymerase expression was induced by inducing it with anhydrotetracycline. Next, the cells were collected by centrifugation and resuspended in 1× Taq buffer (10 mM Tris-HCl (pH 8.0), 50 mM KCl, 1.5 mM MgCl2, 0.1% Triton X-100) (200 μL) for a qPCR screening assay. The PCR mixture used to perform the real-time qPCR assay (conducted in a 96-well plate) contained 10 μL of cell suspension and 40 μL of master mixture. The master mixture consisted of 1,4-butanediol (5% v / v or 7% v / v), 0.25 mM dNTP, 1 mg / mL BSA, 3.5 mM MgCl2, 0.5× SYBR GREEN I, and 0.5 μM of the following primers (5'→3'). GGTCACCCGTTCAACCTGAACAG(23nt)(Sequence ID 10) GTCAACCGCCTTCACGCGGAAC(22nt)(Sequence ID 11)

[0173] Using the Bio-Rad CFX96® real-time PCR detection system, qPCR was performed on a 5% 1,4-butanediol master mixture using the following program: 6 minutes at 95°C followed by 16 cycles [30 seconds at 94°C, 30 seconds at 57.8°C, and 30 seconds at 72°C]. For the extremely rigorous 7% 1,4-butanediol experiment, the qPCR conditions were 1 minute at 98.3°C, 6 minutes at 95°C, followed by 16 cycles [30 seconds at 94°C, 30 seconds at 57.8°C, and 30 seconds at 72°C]. Melting curve analysis was performed at a melting rate of 0.1°C / second between 55°C and 95°C.

[0174] The melting peak was visualized by plotting the absolute relative fluorescence (RFU) of the first derivative against temperature. The peak area was calculated using GraphPad Prism software, and the peak area was normalized to the cell number to rank the clones. To demonstrate the correlation between Taq DNA amount and melting curve area, the Taq DNA polymerase gene was amplified in bulk using the same primers as the two primers described above. The PCR product was column purified, dissolved in Taq buffer, and subsequently quantified using a TeCan instrument. Melting curves and their peak areas were generated using amplicons of different concentrations mixed with SYBR Green to demonstrate the correlation (which showed a linear correlation) (Figure 12). Once the linear correlation was demonstrated, thousands of clones were screened using this method. The quality and specificity of the PCR product were evaluated by agarose gel electrophoresis. To minimize ranking bias and variability, top-hit clones obtained from 96-well plates were re-inoculated, grown in a single plate, and the screening assay was repeated.

[0175] The top 50 clones based on melting curve peak area are shown in the table below. Since the list includes results from several experiments, all melting curve peak areas have been normalized. The normalized results are shown in ranked order in Table A. For the remaining clones, a precise ranking method across the clones could not be established, so they are shown in Table B without rank. While the results are presented in ranked order in Table A, it should be noted that the ranking should be considered only a rough ranking. The primary objective here is to select only the top clones for further investigation.

[0176] The samples shown in the table belong to the following series. Various processes are described in Examples 2, 3, and 4. N Series:

[0177] Wild-type Taq Pol - (error-prone PCR) → N-epPCR → 1 round of selective CSR using 5% BD → N-1st. Seven rounds of CSR are applied to the same library without altering diversity or selection pressure.

[0178] After the enrichment round, clones were screened and named N-round#-plate#-well#. For better visual clarity, the above flowchart can also be represented by the block diagram below.

[0179] TIFF2023059361000010.tif28170

[0180] This library may be referred to as "Library #1" below. L Series:

[0181] Wild-type Taq Pol - (error-prone PCR) → N-epPCR - (selective CSR performed using 5% BD) → N-1st → Top 11 clones selected from screening - (StEP PCR / shuffling) → L-StEP - (selective CSR performed using 7% BD) → L-1st → Screening.

[0182] The same library is subjected to five rounds of CSR without altering diversity or selection pressure.

[0183] Clones after screening are named L-round#-plate#-well#.

[0184] For better visual clarity, the above flowchart can also be represented by the block diagram below.

[0185] TIFF2023059361000011.tif36169

[0186] This library may be referred to as "Library #2" below. T8 Series:

[0187] It is a variant of wild-type Taq polymerase possessing the following unique mutations: F73S, R205K, K219E, M236T, E434D, and A608V; it exhibits superior thermal stability in aqueous media compared to wild-type Taq polymerase (Ghadessy et al., 2001). T8 polymerase was subjected to error-prone PCR, followed by one round of CSR selection, and then screening similar to that of the N-series library described above.

[0188] The inventors define "generation" in terms of the number of times diversity has been introduced into the original epPCR library—for example, when the WT sequence is first diversified by random mutagenesis, it is called the "first generation"—while the "round" number represents the number of times the library has undergone CSR—for example, after the first CSR round means that the library has been selected after one round of CSR.

[0189] Hit clones obtained from screening were represented using the following notation: library# - round# - plate# - well#. For example, N-7-1-E10 refers to a clone isolated from the epPCR library (N) after seven CSR rounds on plate 1 in well E10; on the other hand, L-1-14-H10 refers to a clone isolated from the shuffling library (L) after one CSR round on plate 14 in well H10.

[0190] The inventors first present the screening results obtained from the first round of CSR enrichment for the epPCR library and the shuffled library.

[0191] A. First-round clones ranked based on NMPA within 5% and 7% 1,4-butanediol. [NMPA = Peak area of ​​the melting curve after normalization; BD = 1,4-butanediol] [Table 1]

[0192] B. First-round clones that were not ranked based on NMPA within 5% and 7% 1,4-butanediol [Table 2] TIFF2023059361000014.tif43164

[0193] Next, we present the screening results obtained from the post-shuffling and subsequent enrichment rounds of the epPCR library. [Note: Some of these clones are N-7] th and L-5 th It was synthetically generated based on the combination of mutations identified in the library series. See Example 7A for details.

[0194] Ranked clones after the fifth round of enrichment CSR for the library after C. shuffling. [Table 3] TIFF2023059361000016.tif229162TIFF2023059361000017.tif140164

[0195] Ranked clones after the 7th round of enriched CSR for the D. epPCR library. [Table 4] TIFF2023059361000019.tif206159 Finally, the best clones obtained from all libraries and rounds were screened together for comparison.

[0196] E. Combined screening of epPCR libraries and shuffled libraries. Top hit clones derived from the 1st, 7th CSR epPCR, and 5th CSR SL1 prime library clones were grown in a single 96-well plate to compare PCR performance under two different conditions. Cells were grown as described in "Materials and Methods". PCR was performed in 5% BD according to a program of 6 minutes at 95°C followed by 16 cycles of 30 seconds at 94°C, 30 seconds at 57.8°C, and 30 seconds at 72°C, or in 7% BD according to a program of 1 minute at 98.3°C, 6 minutes at 95°C followed by 16 cycles of 30 seconds at 94°C, 30 seconds at 57.8°C, and 30 seconds at 72°C. In both cases, a final extension was performed at 72°C for 2 minutes, followed by retention at 4°C. The inventors used PCR product (30 μL) and mixed it with 1×SYBR GREEN I to measure the melting curve and determine its area. The melting curve area was normalized by the total number of cells. (Note: SPC refers to a synthetic clone constructed based on a combination of mutations identified after the first round of CSR on an epPCR library; see Example 7 for details.) [Table 5]

[0197] Conclusion: 1. The samples screened in Example 5 were not purified. Therefore, a high NMPA score for 1,4-butanediol inevitably suggests a highly desirable clone. Such clones are L-1-36-A08, L-1-17-A09, L-1-23-H10, N-1-1-D5, L-1-15-A07, and L-1-14-H10 in Table A. They are essentially effective clones. 2. Other clones in Table A, and unranked clones in Table B, are superior only in the sense that they survived the CSR selection process. Their relative merits await further evaluation. Individual mutations present in these clones will also be critically evaluated in Example 11 (a composite list of mutations in the 1,4-butanediol tolerability variant of Taq polymerase). 3. The individual mutations detected in the clone of Example 5 are: i) Incorporation of Example 7 into the synthetic clone, and ii) Theoretical calculations in Example 8 (ΔΔG and ΔΔS vib ) Implementation This also serves as the origin for selecting mutations with a specific purpose in mind. 4. The variant sequence selected for phenotypic testing in Example 11 was also selected from Example 5. 5. The library for NGS (Example 6) also incorporated some modifications derived from the colony in Example 7. 6. After 7 rounds of enrichment CSR in the epPCR library, higher NMPA was observed in the 5% BD compared to the WT (Table C). The same was true for clones obtained after 5 rounds of enrichment CSR in the shuffled library (Table D). 7. To compare the extent to which enrichment improved the performance of clones, the inventors selected clones from each library and screened them in the presence of 5% and 7% BD (Table E). The data clearly show that clones obtained after enrichment have higher NMPA at both BD concentrations.

[0198] Example 6 NGS experiments to identify mutants within Taq polymerase variants To perform the NGS experiments, Taq variant libraries that had undergone three CSR selections and CSR selection / CSR enrichment were used. The steps (diversification and CSR selection) carried out to arrive at these libraries are listed below. The libraries used for NGS were introduced based on Example 5. Note: 1. BD = 1,4 - butanediol. 2. The selection PCR and enrichment PCR are described in Examples 2 and 3. 3. The StEP PCR or DNA shuffling procedure is described in Example 4. 4. Each of Library #1 and #2 contains various sub - libraries corresponding to the number of rounds of enrichment applied. In the case of Library #1, there are 7 rounds of enrichment, while in the case of Library #2, 5 rounds of enrichment were applied. 5. The T8 Taq variant is a variant of wild - type Taq polymerase that contains the following unique mutations: F73S, R205K, K219E, M236T, E434D, and A608V (Ghadessy et al., 2001). Library #3 is based on 1 round of CSR performed on an error - prone library while using T8 as the parental sequence.

[0199] Next - generation sequencing (NGS) using the Illumina technology platform with MiSeq read options was individually performed on each of the above variant libraries.

[0200] As a first step, each DNA within the variant pool was segmented into 6 fragments - 5 of which were each measured to be 450 bp and the 6th was measured to be 468 bp. This was carried out using a high - fidelity DNA polymerase (Q5 obtained from New England Biolab), a standard mixture of dNTPs, and the following 6 sets of overlapping primers. NGS R1: FWD (AAA TCT AGA TAA CGA GGG CAA AAA) (SEQ ID NO: 12) REV(GTC TGC GGT CAG AAT ACG)(SEQ ID NO: 13) NGS R2:FWD(GAG AAA GAA GGT TAC GAG GTT)(SEQ ID NO: 14) REV(ACC GAA CTC CAG ACG TTC)(SEQ ID NO: 15) NGS R3:FWD(CTG CGT GCG TTC CTG)(SEQ ID NO: 16) REV(ACC CCA CAG GTT CGC)(SEQ ID NO: 17) NGS R4:FWD(CTG AGC GAA CGT CTG TTC)(SEQ ID NO: 18) REV(GGT ACG CGG GTG AAT CAG)(SEQ ID NO: 19) NGS R5:FWD(GAC CCG CTG CCG GAC)(SEQ ID NO: 20) REV(GTA ACG TTC GAT GAA CGC TTG)(SEQ ID NO: 21) NGS R6:FWD(GCG ATT CCG TAC GAG GAA)(SEQ ID NO: 22) REV(CCC CTG CAG GTC GAC)(SEQ ID NO: 23)

[0201] Based on the Sanger sequencing data, the inventors have roughly grasped the locations of most of the mutants present within the Taq variant they used. To prevent the loss of mutant data due to primer hybridization during DNA amplification, primers were designed to generate six overlapping regions. Using a quality filter with a Phred score of 13, mutations were identified across the full length of the Taq DNA polymerase gene through alignment of the six overlapping regions against the template Taq polymerase gene.

[0202] The cycling conditions for PCR were as follows: 30 seconds at 98°C + 29 cycles [5 seconds at 98°C, 15 seconds at 55°C, 15 seconds at 72°C] + 2 minutes at 72°C.

[0203] Note that the combined length of the six segments is 2,718 bp. The WT Taq gene is 832 amino acids long, equivalent to a 2,496 bp gene. The difference between the two numbers (2,718 and 2,496) is a result of overlap that occurred while segmenting the gene by PCR.

[0204] The six segments corresponding to each variant pool were sent to GeneWiz for next-generation sequencing using Illumina's synthetic sequencing platform. The "reads" provided by GeneWiz were analyzed in-house using the inventors' proprietary software.

[0205] NGS is a statistical method. To improve the reliability of the results, it is important to increase the diversity of the samples. In this case, this was done by using three variant libraries. NGS also generates a large amount of data. A complete analysis of such data is beyond the scope of this patent specification and is left to one or more subsequent academic publications. In this specification, only top single mutations detected by NGS are considered. Here again, since there is no standard or generally accepted method for prioritizing findings, the inventors used the mutation frequency (F) [percentage of the total (%)] as a general indicator of the significance of the mutation, and in a limited number of cases, also used the enrichment factor (Fe) as an indicator of the rarity of the detected mutation. Here again, in order to keep the number of data within manageable limits, only the top 50 ranked frequencies in each library (or sublibrary) and a smaller number of enrichment factors (Fold-enrichment) in each case were considered. In the case of enrichment factors, the reported data are limited to enrichment factors of 10 or greater. Since very low frequencies can lead to very high enrichment ratios, the inventors considered both frequency and enrichment ratio. If the frequency of a mutation was measurable after NGS, but this mutation was not found in the sample before NGS (pre-deletion), it could mean a very high enrichment ratio, and therefore such mutations are listed.

[0206] Frequency (of occurrence) is defined as the percentage of any particular mutation relative to the total number of mutations. Enrichment factor refers to the enrichment of a particular mutation caused by NGS. It is measured by dividing the frequency of occurrence of that mutation after NGS by its frequency before NGS.

[0207] The results of the NGS tests conducted within the above constraints are shown in Tables A, B, and C below. These results themselves constitute only a part of the information that the inventors used to find the significance of any particular mutation. Its significance becomes more pronounced when used in light of other indicators, as noted in Example 9 (a composite list of mutations in the 1,4-butanediol tolerable Taq polymerase variant).

[0208] A. NGS results table [Top mutations originating from individual libraries]: a) The three spaces within the frequency and enrichment ratio columns represent results obtained from three corresponding libraries. b) In the frequency and enrichment ratio columns, the first space represents the 7th CSR enrichment round of Library #1, the second space represents Library #3 (and several combinations of Library #1, #2, and #3 when considering frequency-enhancement), and the third space represents the 5th enrichment round of Library #2. c) In the case of frequency, a blank (-) means not detected. (A blank (-) means a frequency below the cutoff.) d) A blank (-) means that, in the case of enrichment ratio, it is not selected for measurement (or FE < 10). e) "Missing in Pre" in the enrichment ratio column may indicate a very high enrichment ratio. [Table 6] TIFF2023059361000022.tif225168TIFF2023059361000023.tif225169TIFF2023059361000024.tif228170TIFF2023059361000025.tif227170TIFF2023059361000026.tif227167TIFF2023059361000027.tif228169TIFF2023059361000028.tif227169TIFF2023059361000029.tif228170TIFF2023059361000030.tif226169TIFF2023059361000031.tif230170TIFF2023059361000032.tif229169TIFF2023059361000033.tif229169TIFF2023059361000034.tif196168

[0209] B. Ranking of frequencies (>0.8) - Table of top NGS mutations from all libraries, excluding T8 mutations. NGS + List

Table 7

[0210] C. Table of top NGS mutations from all libraries based on having both high frequency (>5.0) and high amplification factor (>10). NGS ++ List

Table 8

[0211] Conclusion: 1. The following mutations (in the frequency column of Table A) may originate from the T8 variant used as the parent polymerase in library #3 and therefore should not be considered for further evaluation: F73S, R205K, K219T, E434D, and A608V. However, the presence of such mutations in the NGS reaction products of libraries prepared from T8 polymerase confirms the robustness of the NGS method in high-throughput sequencing. 2. Table "A" lists hundreds of mutations. These were aggregated from the top highest-frequency ranked mutations (the most stringent evaluation criteria for 1,4-butanediol tolerability, due to the selection process applied) originating from the three libraries (and, in some cases, their sublibraries). The number obtained from all libraries is smaller than the sum of the top mutations from all three libraries due to overlap. This is also because the list is too large, and the lowest frequency listed in Table "A" is 0.25%. 3. Table "B" is created from Table "A" to reduce this list to a more reasonable number, and to develop the list using a single frequency number (highest of three) for each mutation, and similarly using a single frequency enhancement number (highest of three). In this table, mutations originating from the Taq T8 variant were removed. Using these constraints, Table B provides a list of the top 52 mutations, with a lowest frequency of 0.8%. This list, along with lists obtained by other methods that assess the importance of various mutations resulting in organic solvent tolerance, is incorporated into the Table in Example 9 (a composite list of mutations in the 1,4-butanediol-tolerant Taq polymerase variant). The list in Table B is "NGS + It is named "List". 4. Table "C" is generated by combining frequencies (>5%) that are still higher than those used in Table B, and by imposing another restriction of high-frequency reinforcement (>10) to give importance to the relative rarity of mutations. This list (Table "C") contains fewer mutations - the most important mutations detected by NGS. This list is "NGS ++ This is named the "list." This is also shown in the table of Example 9. Unless strongly refuted by theoretical calculations (ΔΔG calculation - see Example 8), these mutations should be considered highly desirable for conferring stability to the Taq variant in organic aqueous media. 5. The mutations listed in the two tables (Table B and Table C) are adequately suited to the two main objectives of NGS as described herein: namely, to confirm the presence of mutations that strongly contribute to solvent tolerance within selected Taq variants, and namely, to detect rare mutations that provide the same properties but may have been missed by other methods. 6. The most important function of NGS was to identify mutations that might otherwise have been missed. Such mutations (even those with high frequency and enrichment ratios) can be both beneficial and detrimental to enzyme stability, as will be discussed later, but theoretical calculations (ΔΔG and ΔΔS) are used to evaluate their beneficial or detrimental effects. vib This provided a reasonable number of mutations that would allow for such a result. Without such a screening method, theoretical calculations would be faced with billions of possible mutations.

[0212] Example 7 [A?] Variants of the Taq polymerase gene created by synthesis: Evaluation of variant enzymes by real-time qPCR To create more variants from individual mutations that survived the selection pressure, individual mutations derived from selected clones (Examples 5 and 6) were selectively combined to result in 5 to 7 mutations per gene. The objective was to determine whether variants containing multiple mutations could be constructed while possessing the desirable properties derived from the selected mutations. Therefore, the design included not only combinations expected to provide superior tolerance to solvents and temperatures, but also combinations that would result in inferior tolerance. The latter group (expected to have inferior tolerance) was included to provide a negative control and to validate the inventors' design strategy. The proposed combinations were synthesized at Genscript. A total of 14 such genes (denoted SPC1 to SPC14 for clarity when tracking the synthetic clones tested in this example) were synthesized based on mutations identified after one round of CSR on an epPCR library (derived from screening, NGS analysis, or both). The synthetic genes were cloned between the Xbal and Sall restriction sites in a pASK vector. (Many additional synthetic clones were generated based on mutations identified from both screening and NGS analysis in subsequent rounds; see Example 7A). For simplicity of notation, synthetic clones derived from these rounds, including those with mutations observed in screening and NGS, are represented using the conventional format, as described in Example 5, based on the library, CSR round # (from which constitutive mutations are identified), and the plate / well # of the post-screened clones with the largest number of mutations reflected in the synthetic sequence.

[0213] A real-time qPCR assay based on SYBR GREEN I was used to screen clones following the same procedure as in Example 5. In short, colonies were isolated and inoculated into a 96-deep-well culture plate containing LB-chloramphenicol medium (500 μL). The cells were grown and then OD 600When the pH reached 0.4–0.5, polymerase expression was induced by inducing it with anhydrotetracycline. Next, the cells were collected by centrifugation and resuspended in 1× Taq buffer (10 mM Tris-HCl (pH 8.0), 50 mM KCl, 1.5 mM MgCl2, 0.1% Triton X-100) (200 μL) for a real-time qPCR screening assay. The PCR mixture used to perform the real-time qPCR assay in a 96-well plate contained the cell suspension (10 μL) and the master mixture (40 μL). The master mixture consisted of 1,4-butanediol (5% (v / v, or 7% v / v), 0.25 mM dNTP, 1 mg / mL BSA, 3.5 mM MgCl2, 0.5× SYBR GREEN I, and 0.5 μM of the following primers (5'→3'). GGTCACCCGTTCAACCTGAACAG(23nt)(Sequence ID 10) GTCAACCGCCTTCACGCGGAAC(22nt)(Sequence ID 11)

[0214] Using the Bio-Rad CFX96® real-time PCR detection system, qPCR was performed for a 5% 1,4-butanediol master mixture using the following program: 6 minutes at 95°C, followed by 16 cycles [30 seconds at 94°C, 30 seconds at 57.8°C, and 30 seconds at 72°C]. For the extremely rigorous 7% 1,4-butanediol experiment, the qPCR conditions were 1 minute at 98.3°C, 6 minutes at 95°C, followed by 16 cycles [30 seconds at 94°C, 30 seconds at 57.8°C, and 30 seconds at 72°C].

[0215] Melting curve analysis was performed at a melting rate of 0.1°C / second between 55 and 95°C. The melting curves were analyzed as described in Example 7. The melting curve peak areas are shown in the table below. Note that one of the genes (SPC11) was discarded due to a failure in purifying the test enzyme. Four genes—SPC10, 12, 13, and 14—were designed with the expectation that they would not survive qPCR screening (and they did not). A negative control was introduced and this was done to confirm the validity of the positive results.

[0216] [NMPL = Normalized Melt Curve Peak Area; BD = 1,4-Butanediol] [Table 9] TIFF2023059361000042.tif117168

[0217] Conclusion: 1. When the characteristics of the selected mutations for synthesis, derived from the table in Example 9 (see below), were analyzed, the synthesized genes exhibited the expected performance. Thus, we gained confidence that by developing a list of desirable mutations that can be randomly combined, we could obtain solvent-tolerant and temperature-tolerant variants for Taq polymerase. 2. By combining mutations within SPC3, 4, 6, 7, 8, and 9, Taq variants with superior solvent and temperature resistance compared to the parent Taq polymerase were obtained.

[0218] Example 7A [B?] Taq polymerase gene variants obtained from NGS analysis To evaluate and determine optimized mutant sequences, four computational approaches were applied to terminal library mutants identified from libraries N-7th and L-5th. Based on previous results, mutants were selected using previously determined frequencies, cumulative enrichment, and FoldX & Maestro energy calculations obtained for terminal N1-7th libraries and top variants actively manually screened. Using this pooled data collected from both manual and digital screening, the four selection approaches were designed to select individual intrinsic mutants, which were then used to generate random combinations (further digitally screened using energy prediction software FoldX and Maestro). All four selection approaches were designed to maximize the opportunity to select mutants that, when combined, yield variants that provide the greatest improvement in 1,4-butanediol resistance and activity. The four approaches are detailed below.

[0219] Selection approach 1 is the terminal N-7 performed by the inventors. th Or L-5 th This study is based on top cumulative enrichment mutants identified by NGS digital screening of the library. The top 50 most highly cumulatively enriched endemic species present in each of the six regions were analyzed for their predicted effects on protein stability using two tools (FoldX and Maestro). Combinatorial sequences were generated by comprehensively combining endemic species predicted to stabilize Taq polymerase and handle high cumulative enrichment ratios.

[0220] Selection approach 2 is based on the highest frequency mutants with the highest cumulative enrichment, as measured by NGS digital screening of terminal N-7th or L-5th libraries performed by the inventors. Two sets of tables were generated for the Taq polymerase region, one set containing the top 10 most cumulatively enriched intrinsic mutants, and the other containing the top 10 highest frequency (mutants) for each given library series. Combinatorial sequences were generated for each library series by comprehensively combining the intrinsic species found in both the high-frequency table and the high-cumulative enrichment table.

[0221] Selection approach 3 is N-7 th and L-5 th Manual screening was performed using activity assays of variants derived from the library, and the top-performing sequences identified were used as the basis. The scores given by the activity assays (see the "Methods" section) are normalized peak areas (NPAs). Screened sequences that yielded higher NPA scores than those determined for positive control T8s were pooled, and mutations within these sequences were compared against cumulative enrichment data collected by next-generation sequencing. Endemic species identified for each region were compared against their determined cumulative enrichment values, and only those with positive values ​​were retained and comprehensively combined to generate combinatorial sequences.

[0222] Selection Approach 4 is similar to Approach 1, except for the following modifications. Endemic species excluded from the selection of Approach 1 were re-evaluated if the top three endemic species were identified as having high cumulative enrichment and the lowest ΔΔG predicted values ​​greater than zero were exhaustively combined. The top 10 variants with the smallest positive values, as determined by FoldX predictions, were retained and exhaustively combined to obtain N-7 th and L-5 th Combinatorial sequences were generated for both library series.

[0223] The sequence diversity of the comprehensively generated combinatorial sequences is prioritized by clustering the sequences within 10 subgroups based on their sequence similarity to approaches 1, 2, 3, and 4. The most stable members of these subgroup clusters are retained, resulting in a diverse set of combinations that sample the majority of the initial mutation pool. The 10 retained members derived from each selection system and their predicted stability values ​​are shown in Tables A and B.

[0224] The table below does not constitute an exhaustive list of synthetic clones generated by this procedure. It includes N-7 clones, including mutations observed during screening and NGS. th and L-5 th In the library series, synthetic clones generated based on mutations identified from both screening and NGS analysis were represented using the conventional format described in Example 5, based on the library, CSR round # (from which constitutive mutations are identified), and the plate / well # of the post-screening clone that reflected the maximum number of mutations in the synthetic sequence. In addition, some of these synthetic clones included several mutations within NGS regions 2-5 that were not highly ranked in NGS but were randomly selected from a list of mutations observed in NGS in these regions, because mutations in such regions did not show as much enrichment compared to regions 1 and 6 (i.e., some of the observed mutations in regions 2-5 were more neutral in terms of compatibility).

[0225] Table A: Combination variants based on the (N-7)th element [Table 10] TIFF2023059361000044.tif176169

[0226] Table B:L-5 th Alvariant combination based on [Table 11] TIFF2023059361000046.tif57169 All Combinatorial Mutations L5Q,P10S,V14A,A23P,A29T,G32D,G38D,K53R,D58Y,S72N,F73S,K82I,A9 7T,V103A,A109V,R110L,A118T,A141P,R205K,E210D,S213G,K219E,R223P ,L224Q,M236T,D238E,D244E,A246P,A259P,E274K,G304D,D320N,A326V,R 328H,V332I,E337D,E363D,G364D,P382S,N384D,G389D,T399A,A414G,N41 5Y,N415D,E434D,A454E,L461R,A472G,V474I,A478V,H480R,R492L,A502T ,A521T,A454V,E507K,S543I,P548R,D551N,R556G,A568G,V586A,E602D,L 606M,V607I,A608V,S612R,E626D,E626V,L657M,H676Y,Q680R,E708K,D73 2G,E734G,S739G,E745K,F749I,F749V,K762R,K767R,K793R,E832N,E825K

[0227] Example 8 Selected mutations ΔΔG and ΔΔS vib value A specific point mutation in wild-type Taq DNA polymerase causes a change in Gibbs free energy, i.e., folding free energy ΔΔG, and a change in vibrational entropy ΔΔS. vib These were determined using the DynaMut and ENCoM methods. A total of 87 point mutations were selected for such calculations. They were then selected from a list of top clones selected by real-time qPCR screening of CSR products (Example 5).

[0228] ΔΔG value and ΔΔS vibEach value represents the effect of mutation on structural robustness, which in turn suggests thermal stability; the more robust the structure, the higher its thermal stability. Robustness is represented by the negative value of such a function—the larger the negative value, the more stable the structure. While those skilled in the art generally prefer to use only ΔΔG (or its negative value) as an indicator of stability, the inventors of this invention use ΔΔS for the same purpose without denying the importance of ΔΔG. vib We also used this. The reason behind this is that calculations do not necessarily represent reality, and therefore our argument is that the calculated ΔΔG value and ΔΔS vib When all the values ​​point in the same direction, the inventors' confidence in the validity of the conclusion increases. When this occurs, the inventors can freely use ΔΔG for quantitative prediction purposes.

[0229] The calculated values ​​of point mutations related to enzyme stability are shown in three tables. The first table (Table A) shows ΔΔG and ΔΔS vib List only point mutations that yield negative values ​​for both functions. The second table (Table B) lists mutations where one function is negative (indicating stabilization) and the other is positive (indicating destabilization). The third table (Table C) lists ΔΔG and ΔΔS vib List mutations that have positive values ​​for both (indicating destabilization in both cases).

[0230] A. Effects of point mutations on Taq polymerase stability: ΔΔG and ΔΔS vib Both of these suggest stability (robustness). [Table 12] TIFF2023059361000048.tif230163TIFF2023059361000049.tif93169

[0231] B. Effects of point mutations on the stability of Taq polymerase: ΔΔG and ΔΔS vib This exhibits the opposite effect on stability. For clarity, destabilization is marked in red. [Table 13]

[0232] Effect of point mutations on the stability of C. Taq polymerase: ΔΔG and ΔΔS vib Both substances exhibit an enzyme-destabilizing effect. For easier comparison, the destabilized state is shown in red. [Table 14]

[0233] Conclusion: 1. ΔΔG index and ΔΔS vib Based on both indicators, 52 out of 87 tested sites (Table A) showed a stabilizing effect. Of these, 7 mutations (A29T, A61V, T186I, D244V, A608V, S612R, and E832K) showed a very strong stabilizing effect, with ΔΔG values ​​close to or greater than -1.0 kcal / mol. The state of the mutation P10S is ΔΔS vib It is unique in that it has the largest negative value among all of them (-1.053 kcal / mol / K), but its ΔΔG value is only moderately negative (-0.372 kcal / mol). Therefore, the inventors have placed P10S within the group of seven other compounds that have a strong stabilizing effect. 2. Of the 87 mutations tested, 17 other mutations listed in Table A (A23P, P87Q, P89S, K171T, E201K, M236T, D244E, R261H, L287Q, V310L, H333R, L351M, S543G, D551N, Q592R, H676L, and D732N) exhibited a strong stabilizing effect, with ΔΔG values ​​ranging from -0.5 to -1.0 kcal / mol(ΔΔG). 3. Most of the dual stabilization sites (Table A) had a single amino acid substitution, but there were four sites (D244, A454, V586, and H767) that had two amino acid substitutions. 4. Of the 87 sites tested, 35 (Tables B and C) showed an overall destabilizing effect. Of these 35 destabilizing mutations, at least 8 (V14A, L106Q, I163V, L254Q, E277P, L376V, T644G, and F749V) had a very strong destabilizing effect. These mutations should be avoided in particular if the goal is to improve the thermal stability of the enzyme. 5. Mutations A206Q (Table B), V586A (Table A), E687K (Table A), and K709N (Table A) have ΔΔG (<±0.1 kcal / mol) and ΔΔS that are too small to have any meaningful effect on enzyme stability. vib It has a value of (<±0.1 kcal / mol / K).

[0234] Example 9 A composite list of mutations in 1,4-butanediol tolerable Taq polymerase variants. This table lists the mutations found in the top clones during qPCR screening (Example 5), NGS analysis (Example 6), synthesis (Example 7), and ΔΔG calculation (Example 8), in order to assess their importance in developing suitability for PCR reactions in organic aqueous media, particularly when the organic component of the medium is 1,4-butanediol. For definitions of clonal nomenclature, please refer to Examples 5, 6, and 7.

[0235] [Table 15] TIFF2023059361000053.tif226169TIFF2023059361000054.tif225169TIFF2023059361000055.tif224169TIFF20230593610 00056.tif223169TIFF2023059361000057.tif226170TIFF2023059361000058.tif229169TIFF2023059361000059.tif231169 TIFF2023059361000060.tif229170TIFF2023059361000061.tif227169TIFF2023059361000062.tif224169TIFF20230593610 00063.tif230169TIFF2023059361000064.tif228169TIFF2023059361000065.tif222169TIFF2023059361000066.tif127169 1 Among all the samples analyzed, the most negative (highly stabilized) ΔΔS vib It has (-1.053). 2 Strong stabilization ΔΔS vib It has a value of (-0.676). 3 Strong stabilization ΔΔS vib It has a value of (-0.658). 4 ΔΔS vib (-0.305) suggests stabilization. 5 ΔΔS vib (-0.121) suggests mild stabilization. 6 ΔΔS vib (-0.174) suggests mild stabilization. 7 ΔΔS vib (-0.231) suggests a moderate anti-stabilizing effect. 8 ΔΔS vib (-0.793) suggests a strong anti-stabilizing effect. 9 ΔS vib (-0.176) suggests a nearly identical stabilizing effect. 10 ΔS vib (-0.143) suggests a stabilizing effect.

[0236] conclusion 1. Considering all factors, the following mutations (when present in the Taq variant) provide excellent compatibility for PCR and excellent stability in organic aqueous media: (Note: Selection criteria for amino acid position: present in the list). aa must satisfy the following conditions - A) The location must be within two or more independent clones in any case. B) Must have a high NGS frequency (NGS in Example 9) ++ ) C)ΔΔG is stable. D) Any of the above TIFF2023059361000067.tif69169 Note: Underlined -N-Series 7th Round Enrichment CSR-derived unique aa. Italics - A unique mutation derived from the L-series 5th round CSR, The preferred mutations within the above list are as follows: L5Q, F8L, P10S, L16P, A23P, A29T, K31R, G38D, A61V, P89S, A97T, A118V, L162P, K171T, T186I, E201K, R205K, K206Q, G208S, K219E, N220D, I228V, M236T, D244E, D244V, R261H, D273G, L287Q, S290G, V310L, H333R, K346R, L351M, P382T, E388D, E434D, A454E, L461Q, L461R, V474I, F482I, I503T, E507K, S515N, A521V, Q534R, S543G, D551G, D551N, Q592R, L606M, A608V, S612R, H676L, Q680R, K702R, D732N, E734G, S739G, E742K, F749I, F749V, F749L, K762R, K767R, L768M, Q782H, and E832K. And more desirable mutations are as follows: L5Q, F8L, P10S, L16P, A23P, A29T, T186I, K31R, G38D, A97T, A118V, L162P, R205K, G208S, K219E, N220D, I228V, D273G, S290G, K346R, P382T, E388D, E434D, A454E, L461Q, L461R, V474I, F482I, I503T, E507K, S515N, A521V, Q534R, D551G, L606M, A608V, S612R, Q680R, K702R, E734G, S739G, E742K, F749V, F749I, F749L, K762R, K767R, L768M, Q782H, and E832K. And the most desirable mutations are as follows: L5Q, P10S, A23P, A29T, T186I, L461R, E507K, A608V, S612R, E742K, F749L, F749I, K762R, K767R, and E832K. E) Mutations that are detrimental to the stability of the Taq variant in organic aqueous media, Examples include R13H, V14A, L30P, R85S, L106Q, A126G, I163V, Y182H, K187R, G200S, K219E, L254Q, E277P, E365P, L376V, E434D, T664G, and F749I. This does not mean that they cannot be present within the preferred variant; the presence of preferred mutations may outweigh the harmful effects of undesirable ones. The four mutations present in F)T8 were also analyzed in the table above. These are F73S, K219E, M236T, and E434D. Of these, only F73S and M236T are favorable for compatibility in organic aqueous media; the other two, K219E and E434D, are detrimental to compatibility in organic aqueous media. G) The inventors can divide their mutants into two groups: 1) those located within the 5'→3' exodomain; and 2) those belonging to the polymerase domain (Figures 1 and 2). Since deletion of amino acids 1-288 at the N-terminus (similar to the Stoffel fragment) results in a more heat-stable polymerase domain, mutants within category 1), such as P10, P30, A54, A61, F73, I186, etc., may affect thermal stability. In this regard, it has been reported that the Stoffel fragment has a half-life of approximately 2× of WT Taq polymerase at 97.5°C, while the inventors' engineered polymerases showed a half-life up to 7× longer at the same temperature. Even within the exodomain, no thermal stability hotspot is defined. The inventors propose their mutants as part of an increasing list of residue positions that contribute to thermal stability. Interestingly, Reetz and his collaborators have shown a positive correlation between thermal stability and organic solvent tolerance of enzyme activity; for example, lipase mutants exhibiting higher thermal stability also show increased tolerance of enzyme activity to organic solvents. On the other hand, mutants within Category 2 (belonging to the polymerase domain) are particularly concentrated around the substrate binding pocket (Figures 1 and 2). The role of this residue in enhancing the thermal stability of polymerase is not clear, however its effect on catalytic reactions, such as DNA binding and dNTP recognition, is demonstrated as discussed above. Although the molecular determinants of organic solvent tolerance in Taq polymerase are unknown, several other classes of industrial enzymes, including hydrolases, oxidoreductases, transferases, and lyases, have been genetically engineered to be tolerant of organic solvents. Apart from other minor structural elements, most involve surface residues (improvement of interactions within the hydrophobic core or modification of the substrate binding pocket). Water-miscible organic solvents tend to penetrate the active site of enzymes, alter their conformation, and remove water molecules bound to proteins, potentially affecting their activity.The inventors hypothesize that the same mechanism may be at play in Taq polymerase when it encounters hydrophilic organic solvents, such as 1,4-butanediol (for example, BD may affect the thermal stability and activity of the enzyme through its effect on surface residues and / or by substituting water near active site residues). The polymerase domain residues identified in this report may resist local environmental changes and counteract the inhibitory effect of solvents on activity.

[0237] Example 10 Preparation and purification of selected sequences for phenotypic testing For determining the functional properties, ranking Taq variants were prepared and purified in large quantities according to the following procedure.

[0238] To add a His tag to the amplified gene, selected clones were amplified by PCR using a Q5 site-directed mutagenesis kit (NEB) with the following primers. Forward: CACCACCACCGTGGTATGCTGCCGCTG (Sequence ID 24) Reverse: ATGATGATGCATTTTTTGCCCTCGTTATCTAGATTTTTGCT (Sequence ID 25)

[0239] The amplified gene containing the His tag was digested with Xbal and Sall, ligated into a pASK vector, and then digested with the same vector. The ligated product was transformed as previously described. Single colonies expressing WT Taq polymerase or its variant were grown overnight in LB-chloramphenicol (5 mL) at 37°C.

[0240] The culture, after being grown overnight, was re-inoculated in LB-chloramphenicol (200 mL). OD 600When the pH reached 0.4–0.5, protein expression was induced with anhydrotetracycline (300 ng / ml). Cells were collected by centrifugation, washed with buffer (50 mM Tris-HCl (pH 7.9), 50 mM dextrose, 1 mM EDTA, 1 mM PMSF), and resuspended in the same buffer (2.5 mL). The cell suspension was partially lysed by two freeze-thaw cycles. The partially lysed cells were incubated with 1 mg / mL lysozyme at room temperature for 15 minutes. After incubation, an equal volume of lysis buffer (10 mM Tris-HCl (pH 7.9), 50 mM KCl, 1 mM EDTA, 1 mM DTT, 1 mM PMSF, 0.5% Tween-20, 0.5% Nonidet P40) was added; the samples were kept on ice for 30 minutes. The unpurified lysate was then incubated at 75°C for 30 minutes, followed by centrifugation to collect the supernatant. Nucleic acids were precipitated from the supernatant by slowly adding a 20% streptomycin sulfate solution (in 10 mM Tris-HCl (pH 7.90)) at 4°C with constant stirring until the streptomycin concentration reached 4% and nucleic acid precipitation was complete (Upadhyay et al., 2010). The solution was centrifuged, and the supernatant was loaded onto an IMAC column. The column was washed with equilibrium buffer (10 mM Tris-HCl (pH 7.9), 50 mM KCl, 20 mM imidazole), and then eluted with 10 mM Tris-HCl (pH 7.9), 50 mM KCl, 300 mM imidazole. The proteins were dialyzed against a dialysis buffer containing 20 mM Tris-HCl (pH 8.0), 1 mM DTT, 0.1 mM EDTA, 100 mM KCl, 0.5% NP40, 0.5% Tween-20, and 50% glycerol. DNA polymerase was quantified using Biorad's DC protein assay. Protein purity was confirmed by low-resolution SDS-PAGE.

[0241] To determine whether the His tag affects the functional properties of polymerase, in a separate experiment, a protease-cleavable His tag was introduced to the N-terminus of WT polymerase using the following primer set. Reverse: TCGTGGTGGTGATGATGATGCATTTTTTGCCCTCGTTATCTAGATTTTTGTC (Sequence No. 26) Forward:GAACCTGTACTTCCAGTCCCGTGGTATGCTGCCGCTG (Sequence ID 27)

[0242] The remaining procedure was the same as before. Following the vendor's recommendation, the cleavable purified protein was treated with TEV protease (NEB). The cleaved His tags were removed by loading onto an IMAC column, the flow-through was collected, and dialysis followed. Up to 90% of the His tags were removed, as confirmed by InVision His-Tag In-Gel Stain (Invitrogen).

[0243] Example 11 Phenotype testing The following performance tests were performed using purified enzymes (Taq polymerase variants) prepared according to the procedure described in Example 10.

[0244] Example 11a qPCR assay of purified polymerase To evaluate the PCR efficiency of the purified enzyme, fragments of a Taq open reading frame with a length of 531 nucleotides were amplified using wild-type and mutant strains of equal activity. Each variant of equal activity was examined to ensure that each variant had the same active site when co-solvent and at high temperatures. The enzyme was subjected to two different PCR programs: 6 minutes at 95°C in the presence of 5% BD, followed by 16 cycles of 30 seconds at 94°C, 30 seconds at 57.8°C, and 30 seconds at 72°C; and 1 minute at 98.3°C in the presence of 7% BD, followed by 6 minutes at 95°C, followed by 16 cycles of PCR of 30 seconds at 94°C, 30 seconds at 57.8°C, and 30 seconds at 72°C. The number of PCR cycles was limited to restrict the amount of product formed in order to evaluate the effectiveness of the enzyme. The inventors applied very strict criteria and reasoned that mutants with the same or lower Cq values ​​as the wild type (without co-solvent) were variants that performed better when co-solvent was present. Since the Cq value correlates with amplification efficiency, a real-time PCR assay can be employed to confirm the screening rank and identify mutants with better performance. In this assay, the inventors observed a WT Taq polymerase Cq value of 10.61 ± 0.76 cycles in the absence of BD, while this value increased to 11.85 ± 1.34 in the presence of 5% BD. Following the inventors' reasoning, ten mutant clones (in 5% BD) with Cq values ​​similar to or lower than wild-type Taq polymerase (in 0% BD) were identified. These clones originated from both initial epPCR (first generation, first round) and StEP diversification (second generation, first round) libraries, and the data are shown in Table A. Representative qPCR results are shown in Figure 14A. Cq values ​​for WT and SPC clones are the mean ± SD of quadruple measurements. Otherwise, Cq values ​​are the average of two independent experiments. The Cq value for the WT enzyme was 15.5 cycles within 7% BD, suggesting that the enzyme was not efficient up to 16 cycles.Wild-type Taq polymerase showed a Cq very close to the end of PCR, and our mutants continued to amplify the target sequence; therefore, we compared the Cq values ​​of the wild-type and mutants within 7% BD. By doing so, we identified 11 clones capable of amplifying the target DNA in the presence of 7% BD. On the other hand, the synthesized polymerase clones SPC3, 4, 5, and 9 were tolerable up to 7% BD and continued to amplify the target DNA. Overall, 0.37% of the clones derived from the epPCR library performed better than the wild-type, while this ratio was 0.13% and 40% for SL1' and the synthesized clones, respectively.

[0245] We further tested some of our mutants in amplifying high-GC template c-jun fragments by qPCR assay. Experiments were conducted in the presence of different BD concentrations ranging from 0% to 8% (Figure 14B). The WT enzyme was unable to amplify the c-jun template at 0% BD, while substantial amounts of product were produced at BD concentrations of 1–5%. Amplification of the template by WT was negligible at 7% BD and above. However, the mutants were able to amplify the template in the presence of up to 7% BD and produced significant amounts of product at 8% BD. Our data further demonstrate that the PCR efficiency of the mutants is better than that of WT (lower Cq) even at 5% BD. Overall, we conclude that we have engineered organic solvent-tolerant Taq polymerases suitable for amplification of GC-enriched targets at BD levels of at least up to 7–8%.

[0246] A total of seven superior mutants exhibiting better resistance to temperature and BD were selected for real-time PCR analysis from three different generations and rounds of CSR screening (Pipeline 1 - 1st generation library - 1 clone from the 1st enrichment, and 3 clones from the 7th enrichment round, 2 clones from the 2nd generation after the 5th enrichment CSR round, and a synthetic clone derived from the 1st enrichment, represented as SPC9) (see Example 5 for the integrated screening results used to select these mutants). The Cq values ​​of PCR of these mutants were compared to WT-Taq under two different initial temperature treatments (95°C for 6 minutes, or 98.3°C for 1 minute + 95°C for 6 minutes) before PCR cycles using pASK-Taq (Figure 14C) and c-jun (Figure 14D) as a template with various concentrations of BD. All data are shown in Table B. Based on their Cq values, the mutants (L-5-2-F01 and L-5-26-D04) (derived from the second generation after the 5th enrichment CSR round) exhibited better temperature and BD tolerance than the remaining mutants in both pASK-Taq and c-jun template backgrounds. Of these seven mutants, L-5-2-F01-2 was able to produce PCR products even at 10% BD.

[0247] A. Assessment of high-ranking clones by real-time PCR assay. To evaluate the enzyme amplification efficiency on nucleotide fragments with a Taq open reading frame length of 531, WT and engineered polymerases of equal activity were used. Cq values ​​for WT and SPC clones are the mean ± SD of quadruple measurements. Otherwise, Cq values ​​are the mean of two independent experiments, and SDw was calculated according to Synek (2008). NA = No Cq values ​​above background were observed; * = Cq values ​​of polymerase were equal to or less than wild-type Cq values ​​in the absence of 1,4-butanediol within 5% BD, while † = polymerase had better Cq than WT within 7% BD.

[0248] [Table 16]

[0249] B. Efficacy of top-ranked Taq mutants against templates with different GC content in the presence of various BD concentrations. Equal activity WT and engineered polymerases were used to evaluate enzyme amplification efficiency. Cq values ​​were calculated for top clones selected from WT and pipeline 1-first generation libraries (1 clone from the first enrichment round and 3 clones from the seventh enrichment round), 2 clones from the second generation after the fifth enrichment CSR round, and synthetic clone SPC9 (Table 2). (A) Amplification of pASK-Taq template: 6 minutes at 95°C, 30 seconds at 94°C for 17 cycles, 30 seconds at 57.8°C for 30 seconds, and 60 seconds at 72°C, using Q1 and Q2 primers. (B) Amplification of c-Jun template: 6 minutes at 95°C, 30 seconds at 94°C for 16 cycles, 30 seconds at 57.8°C for 30 seconds, and 60 seconds at 72°C, using J1 and J3 primers. (C) Amplification of pASK-Taq template For amplification: Using Q1 and Q2 primers, PCR programs were used at 98.3°C for 1 minute, 95°C for 6 minutes, 94°C for 30 seconds, 57.8°C for 30 seconds, and 72°C for 60 seconds for 17 cycles. For amplification of (D)c-Jun templates: Using J1 and J3 primers, PCR programs were used at 98.3°C for 1 minute, 95°C for 6 minutes, 94°C for 30 seconds, 57.8°C for 30 seconds, and 72°C for 60 seconds for 17 cycles. Cq values ​​are derived from the triple experiment. NA = No Cq values ​​were observed up to the 17th PCR cycle.

[0250] [Table 17] TIFF2023059361000070.tif214169TIFF2023059361000071.tif46169

[0251] The inventors measured the PCR efficiency of one of the best-performing engineered polymerases in BD, also in two other extremely potent organic cosolvents (2-pyrrolidone and sulfolane) (Table C). C. Efficiency of WT-Taq and mutant L-5-2-F01 in the presence of cosolvents of different concentrations (2-pyrrolidone or sulfolane).

[0252] To evaluate the amplification efficiency of the enzyme, WT-Taq and the superior mutant L-5-2-F01-2 with equal activity were used. Cq values ​​are the mean ± SD of the triple measurement. No Cq values ​​above NA = background were observed. The following PCR program was used to amplify the upper stage (2-pyrrolidone; c-jun template) and the lower stage (2-pyrrolidone; c-jun template). In both cases, PCR was performed at 95°C for 6 minutes, or at 98.3°C for 1 minute + 95°C for 6 minutes, followed by 16 cycles at either 94°C for 30 seconds, 57.8°C for 30 seconds, or 72°C for 60 seconds.

[0253] [Table 18] TIFF2023059361000073.tif43160

[0254] Conclusion: 1. For an enzyme to be suitable for PCR applications, PCR efficiency is one of the most important parameters, after specificity and fidelity. A highly efficient polymerase produces high-yield amplicons within the minimum number of PCR cycles. PCR efficiency was evaluated based on Cq values ​​in a non-optimized buffer, limiting PCR to 16 cycles. The identified mutants were portions of both the pPCR library and the post-shuffling library. In addition, four of our synthetic clones that performed better than the WT in terms of PCR efficiency at either 5% BD or 7% BD were also identified in this manner. Our mutants are suitable not only for general PCR applications but can also be used for amplification of GC-enriched target DNA. We demonstrated that some of our mutants can efficiently amplify c-jun templates (templates that cannot be amplified by Taq polymerase in the absence of PCR additives) in the presence of up to 8% BD (Figure 14B), and other mutants in the presence of up to 10% BD (Figure 14C; Table B). In addition to enabling amplification of highly GC-enriched sequences, it should be noted that the fact that the amplification efficiency is nearly equal for two templates with very different GC content, as demonstrated in Figure 14, has significant implications for reducing GC bias in NGS enrichment workflows. Other examples of highly GC-enriched templates include those related to the diagnosis of triplet repeat disorders, such as fragile X syndrome. 2. As shown in Table C, this engineered polymerase was tolerant to such cosolvents at concentrations 7 to 10 times higher than that of WT Taq polymerase.

[0255] Example 11b Thermal stability test The thermal stability (tests) of some of the higher-ranked variants was performed using fluorescence-based methods (Chakrabarti, 2002, 2003). The method was slightly modified by using Eva Green instead of Pico Green (as this dye (Eva Green) is more heat-resistant, although slightly less sensitive). For this experiment, 10 mU lots (in 2 μL) of DNA polymerase were incubated in 1 × Taq buffer at either 95°C or 97.5°C for a) 0, 1, 3, 5, 10, 20, 40, or 60 minutes in the presence of 5% 1,4-butanediol; and b) 0, 5, 10, 20, 40, 60, or 90 minutes in the absence of 1,4-butanediol. The heat-treated sample was kept on ice until the reaction was initiated by adding a substrate mixture (x μl) containing 3 mM MgCl2, 250 μM of each dNTP, 1 × Evagreen in 1 × buffer, and 100 nM of the following SATP primers (Upadhyay et al., 2010): Tagcgaaggatgtgaacctaatccc TGCTCCCGCGGCCGatctgcCGGCCGCGGGAGCA (Sequence No. 28) (Underlined lowercase segments form overhangs for primer extension.)

[0256] After determining the residual activity, the percentage of residual activity (%) at a specific temperature is plotted against the heat exposure time to determine the half-life (t 1 / 2 ) was calculated.

[0257] The results are shown in the table below (t 1 / 2 (The time is shown in minutes.) [Table 19] TIFF2023059361000075.tif48167

[0258] The inventors' goal in this project was to design a Taq variant with better composite compatibility with organic aqueous media, but the most important criterion set by the inventors for their selective pressure was thermal stability in the presence of organic solvents. In this example, we present 12 variants that exhibit significantly higher thermal stability than WT Taq (both in the presence and absence of organic solvents).

[0259] Conclusion: 1. In the samples tested above (these are also included in the preferred selections - see conclusions in Example 9), each resulted in better thermal stability both in the presence and absence of the organic cosolvent 1,4-butanediol. The 18 mutations present in the sequence were P10S, G12T, L30P, A61V, A64V, Y116Stop, T161I, T186I, D244V, K314R, E434D, E520G, V586A, S612R, V730I, F749V, 2493ΔA, and 2493ΔG. These also correspond to the best mutations in that they provide the best overall compatibility for PCR in aqueous organic media, as determined by the composite list in Example 9. 2. The samples that were more tolerant to solvent and temperature, i.e., the five clones that had a half-life of more than twice as long under any test conditions (0% to 7% 1,4-butanediol and 95°C to 97.5°C), had two or more mutations from the following list of 10 mutations: P10S, L30P, E434D, E520G, V586A, S612R, V730I, F749V, 2493ΔA, and 2494ΔG. 3. One sample containing a single mutation F749V (clone N-1-5-E9) exhibited the second-best thermal stability under all conditions.

[0260] The thermal stability of wild-type Taq polymerase and its variants containing mutations was also measured at 95°C in the presence of 0.565 moles of 1,4-butanediol, according to the procedure described above, for a broader spectrum of the polymerase. The overall results are shown in the table below.

[0261] [Table 20]

[0262] While thermal stability assays are well-established and widely accepted for evaluating the heat resistance of polymerases, they depend on the enzyme's primer extension ability. Organic solvents can independently affect both activity and thermal stability. To demonstrate this, thermal stability was assayed independently of activity. Using nanoDSF, the thermal fusion temperatures of His-tagged and non-His-tagged WT and mutant polymerases were profiled. 5 μM of each enzyme was used, and unfolding was monitored while gradually adding heat.

[0263] Melting temperature (T) of wild-type and engineered polymerases M) The profiles were also tested. Thermal unfolding experiments of wild-type polymerase and variants (dissolved in 5 μM in 20 mM Tris-HCl (pH 8.0), 1 mM DTT, 0.1 mM EDTA, 100 mM KCl, 0.5% Nonidet P40 (or Ipegal CA-630), 0.5% Tween-20, and 50% glycerol) were performed on a Prometheus NT.48 instrument using nanoscale differential scanning fluorimetry (nanoDSF), a technique previously applied to evaluate protein stability, with a high-temperature package and backscattering optics that enabled analysis of thermal unfolding and aggregation up to 110 °C. The thermal denaturation of each protein was determined by measuring the change in fluorescence at 330 and 350 nm while varying the temperature from 30 °C to 110 °C at a heating rate of 1 °C / min and a sensitivity setting of 10% (fluorescence excitation power). This measurement was completed in the absence and presence of 1%, 2%, 3%, 5%, 7.5%, and 10% 1,4-butanediol. The nanoDSF high-sensitivity glass capillary was sealed to prevent evaporation of the buffer and sample. The experiment was performed in triplicate at 2Bind GmbH (https: / / 2bind.com, Regensburg, Germany). The 350 / 330 nm ratio and scattering data were analyzed using PR stability analysis software (v.1.1, Nanotemper Technologies, Munich, Germany).

[0264]

Table 21

[0265] T of the clones derived from the 7th round CSR enrichment M To evaluate, the protein was subjected to a slightly modified nanoDSF protocol. Instead of measuring T in 50% glycerol in the presence of surfactant, the inventors excluded the surfactant and limited the glycerol concentration to 5%. This resulted in two distinct Ts that were consistent with previously reported data. M M ​A peak was obtained. In this case, T M,1 While T corresponds to the 5'-3' exonuclease domain, M,2 This represents the stability of the polymerase domain.

[0266] [Table 22]

[0267] Conclusion: 1. As shown in the table above, the His tag does not affect thermal stability; the T of the polymerase domain of His-tagged and non-His-tagged Taq polymerases M These values ​​are 104.30±0.34℃ and 104.39±0.08℃, respectively. 2. Our data are consistent with the two-domain unfolding patterns previously reported for polymerases. We observed that both His-tagged and non-His-tagged Taq polymerases denature in a two-domain manner, meaning the 5'-3' exonuclease domain is unfolded at an earlier temperature compared to the polymerase domain. 3. In engineered Taq polymerases in water and mixed aqueous organic media, there is a general agreement between the half-life and melting temperature results. 4. In most of the top mutants, superior thermal stability was observed compared to the wild type (WT), and the margin of improvement increased with a higher percentage of BD (Block Derived).

[0268] Example 11c Nucleic acid extension time-long, high-speed PCR applications and enhanced processing capacity. Eighteen samples (Example 5) derived from the top-ranking variants (Round 1) that survived the qPCR screening test were tested for nucleotide elongation rates in the PCR reaction.

[0269] To determine the extension time of the selected mutant, a series of PCR reactions were performed in a 96-well PCR plate. The experiment consisted of two parts: (1) Preparation of mutant somatic cell cultures (2) PCR reaction mixture based on the desired PCR program Part 1: Preparation of mutant-hit cell cultures Step 1. The overnight culture (35 μL) of each mutant (serial number 1-20) was transferred to an LB / CPL 96 deep-well plate (500 μL) and grown at 250 rpm, 37°C for 2 hours and 15 minutes. Step 2. Induce the reaction with anhydrotetracycline at a final concentration of 300 ng / mL. Step 3. After incubation for 4 hours, collect the cells by centrifugation at 4,000 rpm and 4°C for 15 minutes. Step 4. Resuspend the cell pellet in 1× Taq buffer + 0.1% Triton (200 μL). Then, place on ice until ready to use. Part 2: Preparation of the PCR reaction mixture PCR was performed in a total volume (50 μL) containing 1 mg / mL BSA, 0.25 mM dNTP mixture, 0.5 μM forward and reverse primers (P1 and P2), and 10 μL of cell suspension (derived from Part 1 Step 4) in 1× Taq buffer containing 0.1% Triton.

[0270] The PCR program and cosolvent conditions used were as follows: [Table 23]

[0271] PCR amplification products were detected using DNA gel. The target amplicon size was approximately 2.5 kb. The extension time for each mutant could be determined. As a control, WT Taq polymerase with an extension rate of 1 minute / kbd (for a 2.5 kb amplicon, the extension time was calculated to be 2.5 minutes) was used. The results were as follows (Table A).

[0272] A. Nucleic acid elongation rate [Table 24]

[0273] Next, we characterized the polymerase specific activity of the top-ranked clones across all rounds (Table B).

[0274] Effect of 1,4-butanediol on the specific activity of B.WT and engineered polymerases. The proteins were purified and quantified. Equal amounts of the proteins were used to evaluate the primer extension activity of the enzyme in the absence and in the presence of 5% BD.

[0275] [Table 25] TIFF2023059361000082.tif212167

[0276] In addition, the processing capacity of the engineered polymerases was evaluated using the methods described above with some modifications. To ensure that the polymerase extended the template in a single binding event, the inventors included heparin as a trapping agent. The processing capacity of the WT was recorded as 14 nucleotides per binding event, which is close to the reported value. Commercial Taq polymerases from NEB were also used to verify and compare the inventors' findings. Generally, the processing capacity of the mutants was similar to that of the WT Taq polymerase, however, mutants such as SPC5, SPC9, N-7-3-C8, and N-7-1-F6 showed higher processing capacity. Common mutations found in these sequences include A61V, T186I, K314R, E520G, A608V, S612R, and F749I. These mutations have been previously identified in the sequences shown in this study as enhancing thermal stability and resistance to inactivation by 1,4-butanediol. Furthermore, it can be noted that the mutations identified here occur in both proximal and distal regions relative to the polymerase active site. Previous studies have shown that increased DNA binding strength of polymerase results in increased processing capacity and resistance to PCR inhibitors. These results suggest that such mutations may enhance other polymerase properties, leading to improved DNA binding or processing capacity. The results obtained from this analysis are shown in the table below (Table C).

[0277] Processing capacity of D.WT and selected Taq mutants. Median elongation lengths and determined processing capacity values ​​for various sequences.

[0278] [Table 26]

[0279] Conclusion: 1. To determine the significance of the above speed for high-speed PCR, we examined a Taq variant possessing one of the fastest extension rates reported to date in high-speed PCR (Arezi et al., 2014). This variant contains a single unique mutation E507K; for a target of Taq gene size (length 2.5kb), it had a calculated extension time of up to 1.25 minutes. Arezi et al. used a 549bp amplicon to measure the extension rate. With our available technology, the small amplicon size used by Arezi et al. did not provide reliable guidance—it showed an unnaturally fast extension rate. For this reason, we used a considerably longer 2.5kb amplicon. Therefore, our results are rather conservative.

[0280] 2. The elongation rate appears to be slightly faster in the presence of the organic cosolvent (1,4-butanediol), as can be seen from the elongation rates of WT Taq in and without 1,4-butanediol.

[0281] 3. As can be seen from the results above (Table A), all 18 variants tested showed significantly faster elongation rates than the parent WT Taq, suggesting that the improvement in compatibility index includes a faster nucleotide integration rate.

[0282] 4. Of the 18 samples tested, at least 8 samples had extension times of less than 1.5 minutes in the absence of 1,4-butanediol and less than 2.1 minutes in the presence of 1,4-butanediol for a full-length 2.5kb Taq gene. These can be considered suitable and preferable for use in most rapid PCR reactions, both in and out of the presence of organic solvents. The clones are as follows: ●Clone N-1-1-D12 containing unique mutations A454L, F749Y, and 2494ΔG ●N-1-4-H07 containing unique mutations H676R and D732G ● N-1-1-E11 containing unique mutations E687K and 2494ΔG ● N-1-1-G10 containing unique mutations A29T and V737D ● N-1-5-E09 containing unique mutations V740A and F749Y ●N-1-5-H07 containing a single unique mutation F749Y ●N-1-1-E08 containing a single unique mutation V310L ●N-1-1-C10 containing a single unique mutation 2494ΔG

[0283] 5. The following clones are more preferable for FAST PCR 9, achieving extension of less than 1.5 minutes in the absence of 1,4-butanediol and less than 1.7 minutes in its presence: ●N-1-5-H07 containing a single unique mutation F749Y ●N-1-1-E08 containing a single unique mutation V310L ●N-1-1-C10 containing a single unique mutation 2494ΔG Each of these more desirable clones possesses a single unique mutation.

[0284] 6. Regarding specific activity, another characteristic contributing to high-speed PCR (Table B), we obtained polymerases ranked higher than WT Taq in 5% BD from screening, with over 30 top-ranking polymerases. Total enrichment clones derived from the final round exhibited better performance than WT.

[0285] 7. From the processing capacity analysis (Table C), it was found that at least eight polymerases had better processing capacity than the WT, and most of the others showed values ​​similar to the WT values. Processing capacity is also related to fast PCR, as higher processing capacity leads to a faster completion of the extension step, especially with long templates.

[0286] Example 11d Fidelity Assay (Note: This test was planned and conducted in the absence of the organic cosolvent 1,4-butanediol.) The fidelity of the wild-type and its mutant derivatives was evaluated using the method described by Barnes and collaborators (Kermekchiev, Tzekov, and Barnes, 2003). Equal amounts of wild-type and mutant polymerases were used to copy the entire LacZ gene along with portions of the kanamycin and ampicillin genes derived from plasmid pWB407. The following PCR program was used to copy the lacZ gene: 3 minutes at 94°C, 45 seconds at 94°C for 16 cycles, 57 seconds at 30°C, and 6 minutes at 72°C. Final extension was performed at 72°C for 10 minutes. The PCR products were purified, restricted digested with ScaI and PstI, and religated to pWB407 digested with the same restriction enzymes. After transforming electrocompetent SS320 E. coli cells, the transformants were seeded onto LB plates containing kanamycin (50 mg / mL), ampicillin (100 mg / mL), and X-gal (20 μg / mL). Following previous documentation (Barnes, 1992), the following formula:

number

[0287] In the formula, F is the percentage of blue colonies; 1000 is the estimated number of non-expression-restricted target sites within the LacZ gene; E is the apparent error rate of the polymerase (error per incorporated nucleotide); m is the number of PCR cycles, and quantity m-1 is used under the hypothesis that it is recessive to the wild-type strand and that errors occurring in the final cycle are not expressed. The results are shown in the table below.

[0288] Fidelity of wild-type Taq polymerase and mutant derivatives. The fidelity of top clones selected from the first-generation library, two clones derived from the seventh enrichment round, two clones derived from the second generation after the fifth enrichment CSR round, and the synthetic clone SPC9 was determined. The apparent error was calculated using the formula above. As a control, a multi-purpose kit polymerase (amplifying the entire LacZ gene and producing 3.5 mutations / kb) was included, followed by an apparent error rate assessment. T and W represent the number of total colonies and white colonies, respectively. The apparent error (Error app.) is given as a change per incorporated base pair, and for example, in the case of WT, it is expected to be one nucleotide change per 19173 incorporated nucleotides. *- represents a clone obtained from the first round of epPCR CSR, **- represents a synthetic polymerase clone, §- represents a clone obtained after the seventh round of epPCR enrichment, l- represents a clone obtained after the fifth round of CSR of shuffling library enrichment.

[0289] [Table 27]

[0290] Conclusion: 1. To determine whether the modified enzyme (Taq polymerase variant) of the present invention has inherently superior or inferior fidelity, the above experiment was systematically carried out in the absence of an organic cosolvent. 2. The results clearly suggest that the selection process carried out by the inventors (Examples 2, 3, and 4) resulted in enzymes that exhibited not only excellent solvent-temperature tolerance (Examples 5 and 11b) but also superior fidelity. 3. Almost all of the samples tested above exhibited better fidelity than the parent WT Taq polymerase, and at least five samples (N-1-2-G2, N-1-1-B1, N-1-2-G4, N-1-1-G5, and N-1-1-G11) showed an improvement in fidelity of more than 25%. Three of these samples (N-1-1-B1, N-1-2-G4, and N-1-1-G5) had a single mutation (A54V, T186I, and E832K). These three mutations (A54V, T186I, and E832K) also qualify as selected single mutations because they possess excellent solvent-temperature tolerance as an individual merit (see conclusions in Example 11). 4.N-1-2-G02 likely has a lower error rate (higher fidelity) than the WT because this clone has two mutations (V586A and S612R) that interact with the substrate. Our findings are consistent with the original concept of CSR (overall compatibility, in which the enzyme must evolve and enrich variants without compromising essential properties, such as fidelity).

[0291] Example 11e Amplification of GC-enriched templates using engineered polymerases A wide range of GC-enriched templates derived from genomic DNA (Table A) were PCR-amplified in the presence of 1,4-butanediol (BD) using WT Taq and one previously reported engineered polymerase (L-5-2-F01). Two different PCR cycling conditions (high denaturation temperature and moderate denaturation temperature) were employed, using various BD concentrations.

[0292] Figure 16 shows that, under high denaturation temperatures, WT Taq cannot amplify any of the seven GC-enriched templates shown in the presence of 5% BD. The target was also unamplable at 1–4% BD when using WT Taq. In contrast, the engineered polymerase variant efficiently amplifies five of the seven GC-enriched templates shown under high denaturation temperatures and in the presence of 7% BD.

[0293] Figure 17 shows that, under high denaturation temperatures, WT Taq still cannot amplify any of these seven GC-enriched templates, even in the presence of 7% BD. On the other hand, by increasing the BD concentration to 10%, the engineered polymerase variant is able to amplify all seven GC-enriched templates (however, it exhibits some nonspecificity, particularly for CD5R2 and DACT3, which have the highest GC content of 64% average / 88% and 79% average / approximately 100%, respectively; Table A and Figure 19). Additional GC-enriched templates, including BAIP3 and LF14 (GC content: 64% average / 80% and 72% average / 90%, respectively; Table A and Figure 19), were also tested under these conditions using the engineered polymerase. The BAIP3 template showed strong amplification with only one nonspecific band, while KLF14 showed significantly lower specificity under these conditions.

[0294] Finally, PCR amplification was compared using this polymerase under lower denaturation temperatures (higher thermal stability of the polymerase) (Figure 18). For WT Taq, this allows the use of higher BD concentrations (7%). Four highly GC-enriched templates were tested, including DACT3 and KLF14 derived from Figure 17, as well as CDN1C and PO3F3 (GC content 77% / max 98% and mean 78% / max 93%, respectively; Table A and Figure 19). Under these conditions, KLF14 was strongly amplified with high specificity when using the engineered polymerase. CDN1C and PO3F3 were also efficiently amplified with high specificity. Note that DACT3 (which had the highest GC content of all templates tested) could not be amplified under these conditions, but its amplification was successfully achieved using the engineered polymerase, as already shown in Figure 17 (higher denaturation temperature and 10% BD). Under these cycling conditions, WT Taq was only able to slightly amplify one of the four templates (KLF14) in 7% BD; for KLF14, the amplification yield was significantly lower than that of the engineered polymerase under identical conditions. Thus, all GC-enriched sequences tested were efficiently amplified using the engineered polymerase, but almost all of them were not efficiently amplified using WT Taq. Only limited PCR optimization was required to achieve these results.

[0295] A. Target Tm and GC content [Table 28]

[0296] Conclusion: 1. Overall, the results above demonstrate that PCR, when using WT Taq, cannot be optimized to amplify many GC-enriched genes, whereas the reported engineered polymerase has the ability to amplify templates with almost any GC content. Therefore, such cosolvent-tolerant engineered polymerases solve, to some extent, the long-term GC-enriched template problem in PCR. 2. In addition, these results are in complete agreement with the first-principles model for PCR amplification presented in the literature (which explains amplification yield in terms of the thermal stability / activity of the enzyme and the effect of the co-solvent on DNA fusion). This allows for the rational optimization of GC-enriched template amplification. 3. When using WT Taq with a higher %BD, a lower denaturation temperature is required (due to the decreased thermal stability of WT Taq within the BD), and when using WT Taq at a higher temperature, a lower %BD is required. Therefore, WT Taq cannot amplify most of the GC-enriched templates tested. Furthermore, regardless of the denaturation temperature used, the cosolvent has a very large inhibitory effect on WT Taq enzyme activity, thus limiting the maximum usable %BD. In contrast, engineered polymerases overcome these limitations that hinder robust GC-enriched template amplification. In particular: 4. In Figure 16 (high denaturation temperature), it is impossible to increase the BD concentration at this denaturation temperature to achieve a sufficient decrease in template Tm to amplify most GC-enriched templates, while this is possible with engineered polymerases. Thus, 5% BD lowers the template Tm by approximately 3-4°C (Figure 13), which may be sufficient for some templates when the denaturation temperature is set above 97°C (due to the fact that only 50% of the template denatures at this Tm). However, it should be noted that the thermal stability of WT Taq is negligibly low when present in 5% BD at such temperatures. 5. In Figures 17 (high denaturation temperature, high %BD) and 18 (lower denaturation temperature, high %BD), it is impossible for WT Taq polymerase to find a temperature-%BD combination that amplifies DACT3 (or other GC-enriched templates) by sufficiently reducing Tm without excessively impairing thermal stability at the selected denaturation temperature. In contrast, in Figure 17, the template Tm can be reduced by approximately 6-7°C by using 10%BD (allowing significant template denaturation at 98°C—a temperature that the engineered polymerase can withstand), and DACT3 was successfully amplified using the engineered polymerase. Note that while the thermal stability of WT Taq at 95°C is reasonable at 5%BD, 5-7%BD only reduces the template Tm by 4-5°C, and therefore even a sufficient reduction in template Tm is not sufficient to amplify highly GC-enriched genes such as PO3F3 and DACT3. In addition, WT Taq activity is significantly reduced in 5% BD. 6. This analysis demonstrates that the polymerase characterization described above has broad and important practical value, as it can be used to predict optimal conditions and enable the amplification of GC-enriched sequences that are otherwise difficult to amplify. 7. Finally, it should be noted that DACT3 and PO3F3 have regions with over 90% GC enrichment, and therefore many other highly GC-enriched sequences are likely to be amplified using such polymerases. Reference materials

[0297] All published materials, patent applications, patents, and other reference materials described herein suggest the level of skill of those skilled in the art to which the subject matter of this disclosure relates. All published materials, patent applications, patents, and other reference materials are incorporated herein by reference to the same extent as each individual published material, patent application, patent, and other reference material is specifically and individually indicated as being incorporated by reference. Although some patent applications, patents, and other reference materials are cited herein, such reference materials are not to be understood to constitute an acknowledgment that any part of this document forms part of the common general knowledge in the art. 6,949,368, September 2005, Chakrabarti, et al. 7,276,357 B2, October 2007, Chakrabarti, et al. 7,514,210 B2, April 2009, Hollinger, et al. 7,772,383 B2, August 2010, Chakrabarti, et al. 8,481,685 B2, July 9, 2013, Bourn, et al. 10,457,968 B2, October 29, 2019, Bourn, et al. Winsor, et al., 1948 Trans. Faraday Soc. vol. 44, p 376; 1950 Trans. Faraday Soc. vol. 46, p 762; 1953 J. Phys. Chem. vol. 57, p 889; 1955 J. Colloid Sc. vol. 10, p 88; and 1960Chemistry and Industry, p. 645. Lindahl et al., Rate of depurination of native deoxyribonucleic acid. 1972Biochemistry, vol. 11, pp 3610-3618. Lindahl et al., Heat-induced deamination of cytosine residues in deoxyribonucleic acid. 1974 Biochemistry, vol. 13, pp 3405-31410. Kuntz et al., Hydration of proteins and polypeptides. 1974 Adv. Protein Chem., vol. 28, pp 239-1974. Sanger et al., DNA Sequencing with chain-terminating inhibitors. 1977 Proc. Natl. Acad. Sci. USA, vol. 74(12), pp 5463-5477. Prince, L.M. Ed. Microemulsions - Theory and Practice, Academic Press, New York, 1977 (pp 21- 32). Eigen et al., Evolutionary molecular engineering based on RNA replication. 1984 Pure Appl. Chem., vol. 56, pp p67-978. Saiki et al., Enzymatic amplification of beta-globin genomic sequences and restriction site analysis for diagnosis of sickle cell anemia. 1985Science, vol. 230, pp 1350-1354. Saiki et al., Analysis of enzymatically amplified beta-globulin and HLA-DQ alpha DNA with allele-specific oligonucleotide probes. 1986 Nature, vol. 324, pp 163-166. Saiki et al., Primer Directed enzymatic amplification of DNA with thermostable DNA polymerase. 1988 Science, vol. 239, pp 487-491. Lawyer et al., Isolation, characterization, and expression in Escherichia coli of the DNA polymerase gene from Thermus aquaticus, 1989 J. Biol. Chem., Vol. 264, pp 6427-6437. Sarkar et al., Formamide can dramatically improve the specificity of PCR. 1990 Nucleic Acids Research, vol. 18, p. 7465 Pomp et al., 0rganic solvents as facilitators of polymerase chain reaction. 1991 Biotechiques, vol. 10, pp 58-59. Fry et al., A DNA polymerase alpha pause site is a hot spot for nucleic acid misinsertion. 1992 Proc. Natl. Acad. Sci. USA, vol. 89, pp 763-767. Barnes, W.M. The Fidelity of Taq Polymerase catalyzing PCR is improved by N-Terminal deletion. 1992Gene, vol. 112, pp 29-35. Sweasy et al., Detection and characterization of mammalian DNA polymerase beta mutants by functional complementation in Escherichia coli. 1993 Proc. Natl. Acad. Sci. USA, vol. 90, pp 4626-4630. Arnold, F.H. Engineering proteins for unusual environments. 1993 FASEB J., vol. 7, pp 744-749. Liao, H.H. Thermostable mutants of kanamycin nucleotidyltransferase are also more stable to proteinase K, urea, detergents, and water-miscible organic solvents. 1993 Enzyme Microb. Technol. vol. 15, pp 286-292 Stemmer, W.P., DNA Shuffling by random fragmentation and reassembly: in vitro recombination for molecular evolution. 1994 Proc Natl Acad Sci USA, vol. 91, pp 10747-10751. Innis, Gelfand & Sninsky, ed., PCR Strategies. Academic Press, New York 1995. Koskinen and Klibanov Ed., Enzymatic Reactions in Organic Media, Balckie Academic and Professional, New York, 1996. Eom et al., Structure of Taq Polymerase with DNA at the polymerase active site. 1996Nature, vol. 382, pp 278-281. Henke et al., Betaine improves the PCR amplification of GC- rich DNA Sequences. 1997 Nucleic Acids Research, vol. 25, pp. 3957-3958. Tawfik et al., Man-made cell-like compartments for molecular evolution. 1998Nature Biotechnol, vol. 16, pp 652-656. Shao et al., Random-priming in vitro recombination: an effective tool for directed evolution. 1998 Nucleic Acids Research, vol. 26(2), 681-683 Mathews, Van Holden and Ahern, Biochemistry, Third Edition, Paerson Prentice Hall, Saddle River, NJ 1999. Klibanov, A.M. Improving enzymes by using them in organic solvents. 2001 Nature, vol. 409, pp 241-246 Ghadessy et al., Directed evolution of polymerase unction by compartmentalized self-replication, 2001 Proc. Natl. Acad. Sci. USA, vol. 98, No. 8, pp 4552-4557. Chakrabarti et al., The enhancement of PCR amplification by low molecular-weight amides. 2001 Nucleic Acids Research, vol. 29, No. 11, pp. 2377-2381. Chakrabarti et al., The enhancement of PCR amplifications by low molecular-weight sulfones. 2001 Gene, vol. 274, pp 293-298. Chakrabarti et al., Novel Sulfoxides Facilitate GC-Rich Template Amplification. 2002 BioTechniques, vol. 32, No. 4, pp 866-874. Chakrabarti, R., PCR Enhancement by Organic Solvents - Progress Toward the Development of Chemical PCR. Dissertation, Princeton University, June 2002. Kermekchiev et. al., Cold-sensitive mutants of Taq DNA polymerase provide a hot start for PCR. 2003 Nucleic Acid Research, vol. 31, pp 6139-6147 Chakrabarti, R., “Novel PCR-Enhancing Compounds and Their Modes of Action”PCR Technology Current Innovations, T. Weissensteiner, H.G. Griffin, and A. Griffin, Ed., Chapter 6, pp 51-63, CRC Press, New York, 2004. Wang et al., A novel strategy to engineer polymerase for enhanced processivity and improved performance in vitro. 2004 Nucleic Acids Res. Vol. 32, pp 1197-1207. Zhao et al., “In vitro ‘sexual’ evolution through the PCR-based staggered extension process (StEP)”2006 Nature Protocols, Doi: 10.1038 / nprot.2006.309. Ghadessy et al., “Compartmentalized Self-Replication” Methods in Molecular Biology, vol. 352: Protein Engineering Protocols, K.M. Arndt and K.M. Muller Ed., Chapter 14, pp. 237-248, Humana Press, Totowa, NJ, 2007. Reetz, et al., Iterative saturation mutagenesis (SM) for rapid directed evolution of functional enzymes. 2007 Nature Protocols, vol. 2(4), pp 891-903. Connolly et al., Recognition of deaminated bases by archaeal family-B DNA polymerases, 2009 Biochem. Soc. Trans., vol. 37, pp 65-68. Tubeleviciute et al., Compartmentalized self-replication (CSR) selection of Thermocococcus Litoralis ShIB DNA polymerase for diminished uracil binding. 2010Protein Engineering, Design & Selection, vol. 31, pp 589-597. Reetz, et al., Increasing the stability of an enzyme toward hostile organic solvents by directed evolution based on iterative saturation mutagenesis using the B-FIT method. 2010 Chem Commun. vol 46, pp 8657-8658. Mardis, E.R., A decade’s perspective on DNA sequencing technology. 2011 Nature, vol. 470, pp 198-203. McClements, D. J. Nanoemulsions versus microemulsions: terminology, differences, and similarities. 2012 The Soft Matter(Journal of the Royal Society of Chemistry), vol. 8, pp 1719-1729. Koudelakova, et al., Engineering Enzyme Stability and Resistance to Organic Cosolvents by Modification of Residues in the Access Tunnel. 2013 Angew. Chem., Int. Ed., vol. 52, pp. 1959-1963. Arezi et al., Compartmentalized self-replication under fast PCR cycling conditions yields Taq Polymerase mutants with increase DNA-binding affinity and blood resistance. 2014 Frontiers in Microbiology, vol. 5, Article 408, 1-10 pages. Yamagami et al., Mutant Taq Polymerases with improved elongation ability as a useful reagent for genetic engineering. 2014Frontiers in Microbiology, vol. 5, Article 46, pp. 1-10. Ondracek et al., Mutations that allow SIR2 orthologs to function in a NAD + -depleted environment. 2017 Cell Reports, vol. 18, pp 2310-2319. Rodriogues et al., DynaMut: predicting the Impact of mutations on protein conformation, flexibility and stability. 2018 Nucleic Acids Research, Web Server issue, vol. 46, pp W350-W355.

[0298] This disclosure includes the following embodiments. Embodiment 1 a) G3D、M4I, L5Q, F8L, E9V, P10S, V14A, L16P, H21R, A23P, L22M, F27S, A29T, G32D, G38D, K53N, A54V, L55P, A61V, D67G, P71L, R74L,R74H,R74C, K82N, G84D, A86V, P87Q, P89S, E90D, A97T, V103A, D104G, A109V, R110Q, P114S, G115D, E117D, A118V, A118T, K128R, V136A, L149P, L162P, K171T, A180V, R183H, T186I, G187S, D191N, L193R, G195S, G200S, E201K, K202R, R205H, K206Q, G212D, S213N, S213G, N220D, L224Q, I228V, H235Y, D237G, W243R, D244E, D244V, L254P, K260N, F258S, R261H, P264S, E267K, E277G, L287Q, S290G, K292N, P302L, P302S, V310L, L311M, D320N, A326V, R328H, H333R, K346R, L351M, E363D, L365Q, P382T, N384D, E388D, T399A, A414S, A454E, A454L, A454V, A458V, L461Q, F482I, L461R, V474I, G499D A502T, I503T, E507K, S515N, S515G, A516G, E520G, A521V, I528T, K531R, Q534R, T539A, S543G, D551N, D551G, V586A, V586M, Q592R, L606M, A608T, S612R, I665V, F667Y, H676L, H676R, H676Y, Q680R, E681K, K702R, A705V, V720L, V730I, D732G, D732N, E734G, V737D, V737A, S739G, V740A, V740I, E742K, F749V, F749I, F749L,A modified Taq DNA polymerase having the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) having one or more amino acid changes selected from the group consisting of K762R, K767R, L768M, E773K, L781P, E797G, E797Q, V799A, P812Q, Q782H, A814V, L813M, E825Q, and E832K, b) A PCR buffer containing one or more low molecular weight organic solvents selected from the group consisting of amides, sulfoxides, sulfones, and diols, wherein one or more low molecular weight organic solvents are present in a concentration in the range of about 0.05 to about 3.0 moles, and A composition containing the following: Embodiment 2 The composition according to Embodiment 1, wherein one or more of the low molecular weight organic solvents are present in the PCR buffer at a concentration of about 0.1 to about 1.0 moles. Embodiment 3 The above one or more amino acid changes are L5Q, F8L, P10S, L16P, A23P, A29T, K31R, G38D, A61V, P89S, A97T, A118V, L162P, K171T, T186I, E201K, R205K, K206Q, G208S, K219E, N220D, I228V, M236T, D244E, D244V, R261H, D273G, L287Q, S290G, V310L, H333R, K346R, L351M, P382T, E388D, E434D, A454E, L461Q, L461R, V474I, F482I, A composition according to Embodiment 1 or 2, selected from the group consisting of I503T, E507K, S515N, A521V, Q534R, S543G, D551G, D551N, Q592R, L606M, A608V, S612R, H676L, Q680R, K702R, D732N, E734G, S739G, E742K, F749I, F749V, F749L, K762R, K767R, L768M, Q782H, and E832K. Embodiment 4 The above one or more amino acid changes are L5Q, F8L, P10S, L16P, A23P, A29T, T186I, K31R, G38D, A97T, A118V, L162P, R205K, G208S, K219E, N220D, I228V, D273G, S290G, K346R, P382T, E388D, E434D, A454E, L461Q, L461R, V474I, F482I, I503T, E507K, S515N, A521V, Q534R, D551G, L606M, A608V, S612R, Q680R, K702R, E734G, S739G, A composition according to any one of Embodiments 1 to 3, selected from the group consisting of E742K, F749V, F749I, F749L, K762R, K767R, L768M, Q782H, and E832K. Embodiment 5 The composition according to any one of Embodiments 1 to 4, wherein the one or more amino acid changes are selected from the group consisting of L5Q, P10S, A23P, A29T, T186I, L461R, E507K, A608V, S612R, E742K, F749L, F749I, K762R, K767R, and E832K. Embodiment 6 The composition according to any one of Embodiments 1 to 5, wherein at least one of the aforementioned amino acid changes is selected from the group consisting of P10S, L16P, A29T, K31R, G38D, A61V, A118V, L162P, T186I, G208S, N220D, I228V, D244V, D273G, S290G, K346R, L351M, E388D, A454E, L461Q, L461R, F482I, I503T, S515N, A521V, Q534R, D551G, L606M, A608V, S612R, Q680R, E734G, S739G, F749V, F749I, L768M, and E832K. Embodiment 7 The composition according to any one of Embodiments 1 to 6, wherein at least one of the aforementioned amino acid changes is selected from the group consisting of F8L, P10S, L16P, A29T, K31R, G38D, A61V, A97T, and L162P. Embodiment 8 The composition according to any one of Embodiments 1 to 7, wherein at least one of the aforementioned amino acid changes is selected from the group consisting of A186I, D244V, R205K, G208S, K219E, N220D, I228V, D273G, S290G, K346R, P382T, E388D, E434D, A454E, L461Q, L461R, V474I, F482I, I503T, E507K, S515N, A521V, Q534R, D551G, and L606M. Embodiment 9 The composition according to any one of Embodiments 1 to 8, wherein at least one of the aforementioned amino acid changes is A608V. Embodiment 10 The composition according to any one of Embodiments 1 to 9, wherein at least one of the aforementioned amino acid changes is selected from the group consisting of S612R, Q680R, K702R, S739G, E742K, L768M, F749I, F749V, K762R, K767R, and Q782H. Embodiment 11 The composition according to any one of Embodiments 1 to 10, wherein at least one of the aforementioned amino acid changes is E832K. Embodiment 12 A composition comprising a modified Taq DNA polymerase suitable for PCR reactions in an organic aqueous medium, wherein the organic aqueous medium comprises one or more low molecular weight organic solvents selected from the group consisting of amides, sulfoxides, sulfones, and diols, and the amino acid sequence of the modified Taq DNA polymerase is L30P, A54V, E434D, K206Q, S612R, V730I, and F749V (Sequence ID 42); P10S, A61V, T186I, D244V, K314R, E520G, V586A, S612R, V730I, and F749V (Sequence ID 43); G12T, A54V, T186I, D244V, F667Y, and F749V (Sequence ID 44); P10S, A61V, F73S, T186I, R205K, K219E, M236T, A608V, S612R, and 2494ΔG (Sequence ID 45); P10S, L30P, A61V, L365P, V586A, S612R, and E832K (Sequence ID 46); P10S, A61V, D244V, S612R, and E832K (Sequence ID 47); L30P and 2494ΔG (Sequence No. 48); A29T, G200S, D237G, and F749I (Sequence ID 49); L16P, F73S, E388D, Q680R, and F749I (Sequence ID 50); F73S, K346R, A454E, and F749V (Sequence ID 51); F73S, A118V, and F749I (Sequence ID 52); A23P, L162P, I228V, L461R, A521V, E734G, F749I, and L768M (Sequence ID 53); K31R, F482I, Q534R, A608V, and F749I (Sequence ID 54); A23P and F749I (Sequence ID 55); G38D, F73S, A454V, and F749V (Sequence ID 56); N220D, I503T, S515N, and F749V (Sequence ID 57); A29T, F73S, S290G, L461R, D551G, L606M, S739G, and F749I (Sequence ID 58); E434D, A608V, and K762R (Sequence ID 59); E434D, E507K, and K762R (sequence number 60); E434D, E507K, E742K, and F749I (sequence number 61); P10S, P382T, E434D, and E507K (Sequence ID 62); R205K, K219E, E434D, V474I, A608V, inS661R, E742K, and F749I (Sequence ID 63); A97T, A608V, K702R, and K762R (Sequence ID 64); F8L, P10S, E434D, E507K, K762R, and K767R (Sequence ID 65); P10S, E507K, Q680R, and K762R (Sequence ID 66); E507K, A608V, Q782H, and F749I (SEQ ID NO: 67); E434D, A608V, E742 K, and F749I (Sequence ID 68); E520G, V586A, S612R, and 2493ΔA (Sequence ID 69); P10S, V730I, and 2493ΔA (Sequence ID 70); V586A, S612R, S674S, and 2494ΔGA (SEQ ID NO: 71); E434D and 2494ΔGA (SEQ ID NO: 72); Y116Stop2494ΔG (Sequence ID 73); A54V (Sequence ID 74); A61V (Sequence ID 75); F749V (Sequence code 76); E832K (sequence number 77); T186I, V586A, S612R, and 2494ΔG (Sequence ID 78); A64V and 2493ΔA (Sequence ID 79); D244V, K314R, V586A, and S612R (Sequence ID 80); A61V, T161I, V586A, S612R, and 2494ΔG (Sequence ID 81); G12T, A61V, and 2494ΔG (Sequence ID 82); A29T, K53R, R205K, K219E, D320N, A326V, N415D, L461R, E602D, and A608V (Sequence ID 83); A29T, K53R, R205K, K219E, D244E, D320N, A326V, N415D, L461R, and A608V (Sequence ID 84); A29T, K53R, R223P, D320N, A326V, N415D, L461R, E602D, and A608V (Sequence ID 85); A29T, D238E, R328H, L461R, A608V, E745K, and F749I (Sequence ID 86); A29T, F73S, D238E, R328H, D551N, A608V, E745K, and F749I (Sequence ID 87); A29T, D238E, R328H, D551N, A608V, and F749V (Sequence ID 88); A109V, L224Q, T399A, A502T, A608V, and F749I (Sequence ID 89); A109V, L224Q, T399A, A502T, A608V, S739G, and F749I (Sequence ID 90); A29T, L224Q, T399A, A454E, A608V, S739G, and F749I (Sequence ID 91); K53R, F73S, A141P, P382S, A472G, R556G, and F749I (Sequence ID 92); R110L, K219E, M236T, E274K, R492L, A608V, E626D, K767R, and E825K (Sequence ID 93); R110L, K219E, M236T, N415Y, R492L, A608V, K767R, and E832N (Sequence ID 94); K82I, K219E, M236T, N415Y, R492L, A608V, E626V, and K793R (Sequence ID 95); P10S, F73S, K219E, M236T, E337D, E507K, A608V, and K767R (Sequence ID 96); P10S, F73S, K219E, E337D, E434D, V474I, A608V, and K767R (Sequence ID 97); P10S, F73S, K219E, E337D, E434D, A608V, and K767R (Sequence ID 98); P10S, V14A, R205K, K219E, M236T, N384D, V474I, A608V, S612R, and K762R (Sequence ID 99); P10S, V14A, K219E, N384D, E434D, V474I, A608V, S612R, K762R, and K767R (Sequence ID 100); P10S, V14A, R205K, K219E, N384D, V474I, A608V, S612R, and F749I (Sequence ID 101); and R110L, R205K, K219E, N415Y, S543I, A608V, E626D, K767R, and E825K (Sequence ID 102) A composition having an amino acid sequence that is at least 90% identical to that of wild-type Taq DNA polymerase (SEQ ID NO: 41) having an amino acid change selected from the group consisting of the above. Embodiment 13 A composition comprising one or more DNA polymerases having increased thermal stability compared to wild-type Taq DNA polymerase in a PCR buffer containing 0-10% by weight of one or more organic cosolvents, wherein the one or more DNA polymerases are P10S, G12T, L16P, A23P, A29T, L30P, K31R, G38D, A61V, A64V, F73S, Y116Stop, A118V, T161I, L162P, T186I, G200S, N220D, I228V, D237G, D244V, S290G, K3 A composition comprising a modified Taq DNA polymerase having an amino acid sequence containing the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) having one or more amino acid changes selected from the group consisting of 14R, K346R, E388D, E434D, A454E, A454V, L461R, F482I, I503T, S515N, E520G, A521V, Q534R, D551G, V586A, L606M, A608V, S612R, Q680R, V730T, E734G, S739G, F749I, F749V, L768M, 2493ΔA, and 2494ΔG. Embodiment 14 The composition according to Embodiment 13, wherein one or more of the aforementioned amino acid changes are selected from the group consisting of P10S, A29T, L30P, K31R, F73S, A118V, G200S, G237G, K346R, S434D, A454E, F482I, E520G, Q534R, V586A, A608V, S612R, V730I, F749I, F749V, 2493ΔA, and 2494ΔG. Embodiment 15 The aforementioned one or more DNA polymerases F749V (Sequence code 76); F30L and 2494ΔG; E520G, V586A, S612R, and 2493ΔA; E434D and 2494Δ (Sequence ID 72); P10S, V730I, and 2493ΔA (Sequence ID 70); Y116Stop and 2494ΔG (Sequence ID 73); A64V and 2493ΔA (Sequence ID 79); T186I, V586A, S612R, and 2494ΔG (Sequence ID 78); V586A, S612R, and 2494ΔG; D244V, K314R, V586A, and S612R (Sequence ID 80); A61V, T161I, V586A, S612R, and 2494ΔG (Sequence ID 81); G12T, A61V, and 2494ΔG (Sequence ID 82); A29T, G200S, D237G, and F749I (Sequence ID 49); L16P, F73S, E388D, Q680R, and F749I (Sequence ID 50); F73S, K346R, A454E, and F749V (Sequence ID 51); F73S, A118V, and F749I (Sequence ID 52); A23P, L162P, I228V, L461R, A521V, E734G, F749I, and L768M (Sequence ID 53); K31R, F482I, Q534R, A608V, and F749I (Sequence ID 54); A23P and F749I (Sequence ID 55); G38D, F73S, A454V, and F749V (Sequence ID 56); N220D, I503T, S515N, and F749V (Sequence ID 57); and A29T, F73S, S290G, L461R, D551G, L606M, S739G, and F749I (Sequence ID 58) The composition according to Embodiment 13, having an amino acid sequence that is at least 90% identical to the amino acid sequence of a wild-type Taq DNA polymerase (SEQ ID NO: 41) having an amino acid change selected from the group consisting of the above. Embodiment 16 The composition according to Embodiment 13, wherein the organic cosolvent is selected from the group consisting of low molecular weight amides, low molecular weight sulfoxides, low molecular weight sulfones, and low molecular weight diols. Embodiment 17 A composition comprising one or more DNA polymerases exhibiting increased fidelity compared to wild-type Taq DNA polymerase in a PCR buffer containing 0-10% by weight of one or more organic cosolvents, wherein the one or more DNA polymerases are P10S, G12T, A23P, K31R, A54V, A61V, F73S, Y116Stop, A118V, L162P, T186I, K206Q, I228V, D244V, K314R, L461R A composition comprising a modified Taq DNA polymerase having an amino acid sequence comprising the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) having one or more amino acid changes selected from the group consisting of F482I, A521V, Q534R, V586A, A608V, S612R, E734G, F749I, L768M, E832K, 2494ΔG, A23P, K31R, L162P, I228V, L461R, F482I, A521V, E734G, F749I, and L768M. Embodiment 18 The composition according to Embodiment 17, wherein one or more of the amino acid changes are selected from the group consisting of K31R, A54V, F73S, A118V, T186I, K206Q, D244V, K314R, F482I, Q534R, V586A, A608V, S612R, F749I, E832K, and 2494ΔG. Embodiment 19 The composition according to Embodiment 17, wherein one or more of the amino acid changes are selected from the group consisting of A54V, T186I, and E832K. Embodiment 20 The aforementioned one or more DNA polymerases A54V (Sequence ID 74); T186I (Sequence ID 103); E832K (sequence number 77); D244V, K314R, V586A, and S612R (Sequence ID 80); K206Q and 2494ΔG (Sequence ID 104); G12T, A61V, and 2494ΔG (Sequence ID 82); P10S (Sequence code 105); K31R, F482I, Q534R, A608V, and F749I (Sequence ID 54); F73S, A118V, and F749I (sequence number 52); and A23P, L162P, I228V, L461R, A521V, E734G, F749I, and L768M (Sequence ID 53) The composition according to Embodiment 17, having an amino acid sequence that is at least 90% identical to the amino acid sequence of a wild-type Taq DNA polymerase (SEQ ID NO: 41) having an amino acid change selected from the group consisting of the following. Embodiment 21 The composition according to Embodiment 20, wherein the organic cosolvent is selected from the group consisting of low molecular weight amides, low molecular weight sulfoxides, low molecular weight sulfones, and low molecular weight diols. Embodiment 22 A composition comprising one or more DNA polymerases in a PCR buffer containing 0-10% by weight of one or more organic cosolvents, wherein the nucleotide integration rate is increased and the processing capacity is increased compared to wild-type Taq DNA polymerase, wherein the one or more DNA polymerases comprises a modified Taq DNA polymerase having an amino acid sequence that includes the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 1) having one or more amino acid changes selected from the group consisting of A29T, V310L, A454L, H676R, E687K, D732G, V737D, V740A, F749V, and 2494ΔG. Embodiment 23 The composition according to Embodiment 22, wherein one or more of the amino acid changes are selected from the group consisting of V310L, F749Y, and 2494ΔG. Embodiment 24 The aforementioned one or more DNA polymerases F749V (Sequence code 76); F310L (Sequence ID 106); 2494ΔG; A454L, F749V, and 2494ΔG (Sequence ID 107); H676R and D732G (Sequence ID 108); E687K and 2494ΔG (Sequence ID 109); A29T and V737D (sequence number 110); and V740A and F749V (Sequence ID 111) The composition according to Embodiment 22, having an amino acid sequence that is at least 90% identical to the amino acid sequence of a wild-type Taq DNA polymerase (SEQ ID NO: 41) having an amino acid change selected from the group consisting of the above. Embodiment 25 The composition according to Embodiment 24, wherein the organic cosolvent is selected from the group consisting of low molecular weight amides, low molecular weight sulfoxides, low molecular weight sulfones, and low molecular weight diols. Embodiment 26 The amide is selected from the group consisting of formamide, N-methylformamide, N,N-dimethylformamide (DMF), acetamide, N-methylacetamide, N,N-dimethylacetamide, propionamide, isobutylamide, 2-pyrrolidone, N-methylpyrrolidone (NMP), N-hydroxyethylpyrrolidone (HEP), N-formylpyrrolidine, N-formylmorpholine; δ-valerolactam, ε-caprolactam, and 2-azacyclooctanone. The sulfoxide is selected from the group consisting of dimethyl sulfoxide (DMSO), n-propyl sulfoxide, n-butyl sulfoxide, methyl sec-butyl sulfoxide, and tetramethylene sulfoxide. The sulfone is selected from the group consisting of dimethyl sulfone, diethyl sulfone, di(n-isopropyl) sulfone, tetramethylene sulfone (sulfolane), 2,4-dimethylsulfolane, and butadiene sulfone (sulfolene). The composition according to any one of Embodiments 1 to 25, wherein the diol is selected from the group consisting of 1,2-propanediol, 1,3-propanediol, 1,2-butanediol, 1,3-butanediol, 1,4-butanediol, 1,2-pentanediol, 2,4-pentanediol, 1,5-pentanediol, 1,2-cyclopentanediol, 1,2-hexanediol, 1,6-hexanediol, and 2-methyl-2,4-pentanediol. Embodiment 27 The composition according to Embodiment 26, wherein the amide solvent is N,N-dimethylformamide (DMF) at a concentration of about 0.5 to about 1.5 molars; isobutylamide at a concentration of about 0.1 to about 1.0 molars; 2-pyrrolidone at a concentration of about 0.1 to about 1.0 molars; or N-methylpyrrolidone at a concentration of about 0.1 to about 1.0 molars. Embodiment 28 The composition according to Embodiment 26, wherein the sulfoxide is dimethyl sulfoxide (DMSO) at a concentration of about 0.5 to about 3.0 molars, or tetramethylene sulfoxide at a concentration of about 0.1 to about 1.0 molars. Embodiment 29 The composition according to Embodiment 26, wherein the sulfone is tetramethylene sulfone (sulfolane) at a concentration of about 0.1 to about 1.0 moles. Embodiment 30 The composition according to Embodiment 26, wherein the diol is 1,3-propanediol at a concentration of about 0.5 to about 3.0 molars; 1,4-butanediol at a concentration of about 0.5 to about 2.0% molars; or 1,5-pentanediol at a concentration of about 0.5 to about 1.0% molars. Embodiment 31 a) At least one modified Taq DNA polymerase as described in Embodiments 1, 2, 3, 4, or 5, b) A buffer suitable for use in a PCR reaction, which may optionally contain one or more organic cosolvents described in Embodiments 1 and 26, and A kit that includes this. Embodiment 32 a) One or more DNA polymerases comprising a modified Taq DNA polymerase, wherein the amino acid sequence of the modified Taq DNA polymerase is L30P, A54V, E434D, K206Q, S612R, V730I, and F749V (Sequence ID 42); P10S, A61V, T186I, D244V, K314R, E520G, V586A, S612R, V730I, and F749V (Sequence ID 43); G12T, A54V, T186I, D244V, F667Y, and F749V (Sequence ID 44); P10S, A61V, F73S, T186I, R205K, K219E, M236T, A608V, S612R, and 2494ΔG (Sequence ID 45); P10S, L30P, A61V, L365P, V586A, S612R, and E832K (Sequence ID 46); P10S, A61V, D244V, S612R, and E832K (Sequence ID 47); L30P and 2494ΔG (Sequence No. 48); E520G, V586A, S612R, and 2493ΔA (Sequence ID 69); P10S, V730I, and 2493ΔA (Sequence ID 70); V586A, S612R, S674S, and 2494ΔGA (SEQ ID NO: 71); E434D and 2494ΔGA (SEQ ID NO: 72); Y116Stop2494ΔG (Sequence ID 73); A54V (Sequence ID 74); A61V (Sequence ID 75); F749V (Sequence code 76); E832K (sequence number 77); T186I, V586A, S612R, and 2494ΔG (Sequence ID 78); A64V and 2493ΔA (Sequence ID 79); D244V, K314R, V586A, and S612R (Sequence ID 80); A61V, T161I, V586A, S612R, and 2494ΔG (Sequence ID 81); G12T, A61V, and 2494ΔG (Sequence ID 82); T186I (Sequence ID 103); K206Q and 2494ΔG (Sequence ID 104); P10S (Sequence code 105); F310L (Sequence ID 106); 2494ΔG; A454L, F749V, 2494ΔG (Sequence ID 107); H676R and D732G (Sequence ID 108); E687K and 2494ΔG (Sequence ID 109); A29T and V737D (sequence number 110); V740A and F749V (Sequence ID 111); A29T, K53R, R223P, D320N, A326V, N415D, L461R, E602D, and A608V (Sequence ID 85); A29T, D238E, R328H, L461R, A608V, E745K, and F749I (Sequence ID 86); A29T, F73S, D238E, R328H, D551N, A608V, E745K, and F749I (Sequence ID 87); A29T, D238E, R328H, D551N, A608V, and F749V (Sequence ID 88); A109V, L224Q, T399A, A502T, A608V, and F749I (Sequence ID 89); A109V, L224Q, T399A, A502T, A608V, S739G, and F749I (Sequence ID 90); A29T, L224Q, T399A, A454E, A608V, S739G, and F749I (Sequence ID 91); K53R, F73S, A141P, P382S, A472G, R556G, and F749I (Sequence ID 92); R110L, K219E, M236T, E274K, R492L, A608V, E626D, K767R, and E825K (Sequence ID 93); R110L, K219E, M236T, N415Y, R492L, A608V, K767R, and E832N (Sequence ID 94); K82I, K219E, M236T, N415Y, R492L, A608V, E626V, and K793R (Sequence ID 95); P10S, F73S, K219E, M236T, E337D, E507K, A608V, and K767R (Sequence ID 96); P10S, F73S, K219E, E337D, E434D, V474I, A608V, and K767R (Sequence ID 97); P10S, F73S, K219E, E337D, E434D, A608V, and K767R (Sequence ID 98); P10S, V14A, R205K, K219E, M236T, N384D, V474I, A608V, S612R, and K762R (Sequence ID 99); P10S, V14A, K219E, N384D, E434D, V474I, A608V, S612R, K762R, and K767R (Sequence ID 100); P10S, V14A, R205K, K219E, N384D, V474I, A608V, S612R, and F749I (Sequence ID 101); R110L, R205K, K219E, N415Y, S543I, A608V, E626D, K767R, and E825K (Sequence ID 102) K31R, F482I, Q534R, A608V, F749I; F73S, A118V, F749I; A23P, L162P, I228V, L461R, A521V, E734G, F749I, L768M; A29T, G200S, D237G, F749I; L16P, F73S, E388D, Q680R, F749I; F73S, K346R, A454E, F749V; A23P, F749I; G38D, F73S, A454V, F749V; N220D, I503T, S515N, F749V; and A29T, F73S, S290G, L461R, D551G, L606M, S739G, F749I The amino acid sequence is at least 90% identical to that of wild-type Taq DNA polymerase (SEQ ID NO: 41) having an amino acid change selected from the group consisting of the following: One or more DNA polymerases, b) A buffer suitable for use in a PCR reaction, which may optionally contain one or more organic cosolvents described in Embodiments 1 and 26, and A kit that includes this. The above description of illustrative embodiments of the Disclosure is provided for illustrative and explanatory purposes only. The above description is not intended to be exhaustive or to limit the Disclosure to any specific form disclosed, and modifications and alterations may be made, in view of the above teaching, or derived from the practice of the Disclosure. The embodiments have been selected and described as practical applications of the Disclosure to illustrate the principles of the Disclosure and to enable those skilled in the art to utilize the Disclosure in various embodiments, with various modifications to suit specific uses. The scope of the Disclosure is intended to be defined by the claims and equivalents appended herein.

Claims

1. a) A modified Taq DNA polymerase having the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) containing a set of amino acid changes selected from the second column of the following table, Table 1 b) A PCR buffer containing one or more low molecular weight organic solvents selected from the group consisting of amides, sulfoxides, sulfones, and diols, wherein the one or more low molecular weight organic solvents are present in a concentration range of about 0.05 to about 3.0 molars, and A composition containing the following:

2. The composition according to claim 1, wherein one or more of the low molecular weight organic solvents are present in the PCR buffer at a concentration of about 0.1 to about 1.0 molars.

3. A composition comprising a modified Taq DNA polymerase suitable for PCR reactions in an organic aqueous medium, wherein the organic aqueous medium comprises one or more low molecular weight organic solvents selected from the group consisting of amides, sulfoxides, sulfones, and diols, the amino acid sequence of the modified Taq DNA polymerase is at least 90% identical to the amino acid sequence containing the sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41), and has a set of amino acid changes based on SEQ ID NO: 41, selected from the second column of the table in claim 1.

4. A composition comprising one or more DNA polymerases having increased thermal stability compared to wild-type Taq DNA polymerase in a PCR buffer containing 0 to 10% by weight of one or more organic cosolvents, wherein the one or more DNA polymerases are as follows: Table 2 A composition comprising a modified Taq DNA polymerase having an amino acid sequence comprising the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) comprising a set of amino acid changes selected from the second column.

5. The composition according to claim 4, wherein the organic cosolvent is selected from the group consisting of low molecular weight amides, low molecular weight sulfoxides, low molecular weight sulfones, and low molecular weight diols.

6. A composition comprising one or more DNA polymerases having increased fidelity compared to wild-type Taq DNA polymerase in a PCR buffer containing 0 to 10% by weight of one or more organic cosolvents, wherein the one or more DNA polymerases are as follows: Table 3 A composition comprising a modified Taq DNA polymerase having an amino acid sequence comprising the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) comprising a set of amino acid changes selected from the second column.

7. A composition comprising one or more DNA polymerases having increased fidelity compared to wild-type Taq DNA polymerase in a PCR buffer containing 0 to 10% by weight of one or more organic cosolvents, wherein the one or more DNA polymerases have an amino acid sequence that is at least 90% identical to the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41), and have a set of amino acid changes based on SEQ ID NO: 41, selected from the second column of the table in Claim 6.

8. The composition according to claim 6 or 7, wherein the organic cosolvent is selected from the group consisting of low molecular weight amides, low molecular weight sulfoxides, low molecular weight sulfones, and low molecular weight diols.

9. A composition comprising one or more DNA polymerases in a PCR buffer containing 0-10% by weight of one or more organic cosolvents, wherein the nucleotide integration rate and processing capacity are increased compared to wild-type Taq DNA polymerase, wherein the one or more DNA polymerases are as follows: Table 4 A composition comprising a modified Taq DNA polymerase having an amino acid sequence comprising the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) comprising a set of amino acid changes selected from the second column.

10. A composition comprising one or more DNA polymerases in a PCR buffer containing 0 to 10% by weight of one or more organic cosolvents, wherein the nucleotide integration rate is increased and the processing capacity is increased compared to wild-type Taq DNA polymerase, wherein the one or more DNA polymerases have an amino acid sequence that is at least 90% identical to the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41), and have a set of amino acid changes based on SEQ ID NO: 41, selected from the second column of the table in Claim 9.

11. The composition according to claim 9 or 10, wherein the organic cosolvent is selected from the group consisting of low molecular weight amides, low molecular weight sulfoxides, low molecular weight sulfones, and low molecular weight diols.

12. The amide is selected from the group consisting of formamide, N-methylformamide, N,N-dimethylformamide (DMF), acetamide, N-methylacetamide, N,N-dimethylacetamide, propionamide, isobutylamide, 2-pyrrolidone, N-methylpyrrolidone (NMP), N-hydroxyethylpyrrolidone (HEP), N-formylpyrrolidine, N-formylmorpholine; δ-valerolactam, ε-caprolactam, and 2-azacyclooctanone. The sulfoxide is selected from the group consisting of dimethyl sulfoxide (DMSO), n-propyl sulfoxide, n-butyl sulfoxide, methyl sec-butyl sulfoxide, and tetramethylene sulfoxide. The sulfone is selected from the group consisting of dimethyl sulfone, diethyl sulfone, di(n-isopropyl) sulfone, tetramethylene sulfone (sulfolane), 2,4-dimethylsulfolane, and butadiene sulfone (sulfolene). The composition according to any one of claims 1 to 3, 5, 8 and 11, wherein the diol is selected from the group consisting of 1,2-propanediol, 1,3-propanediol, 1,2-butanediol, 1,3-butanediol, 1,4-butanediol, 1,2-pentanediol, 2,4-pentanediol, 1,5-pentanediol, 1,2-cyclopentanediol, 1,2-hexanediol, 1,6-hexanediol, and 2-methyl-2,4-pentanediol.

13. The composition according to claim 12, wherein the amide solvent is N,N-dimethylformamide (DMF) at a concentration of about 0.5 to about 1.5 molars; isobutylamide at a concentration of about 0.1 to about 1.0 molars; 2-pyrrolidone at a concentration of about 0.1 to about 1.0 molars; or N-methylpyrrolidone at a concentration of about 0.1 to about 1.0 molars.

14. The composition according to claim 12, wherein the sulfoxide is dimethyl sulfoxide (DMSO) at a concentration of about 0.5 to about 3.0 molars, or tetramethylene sulfoxide at a concentration of about 0.1 to about 1.0 molars.

15. The composition according to claim 12, wherein the sulfone is tetramethylene sulfone (sulfolane) at a concentration of about 0.1 to about 1.0 molar.

16. The composition according to claim 12, wherein the diol is 1,3-propanediol at a concentration of about 0.5 to about 3.0 molars; 1,4-butanediol at a concentration of about 0.5 to about 2.0% molars; or 1,5-pentanediol at a concentration of about 0.5 to about 1.0% molars.

17. a) At least one modified Taq DNA polymerase according to claim 1, 3, 4, 6, 7, 9, or 10, b) A buffer suitable for use in PCR reactions, which may optionally contain one or more organic cosolvents, and A kit that includes this.

18. a) One or more DNA polymerases comprising a modified Taq DNA polymerase, wherein the amino acid sequence of the modified Taq DNA polymerase is at least 90% identical to the amino acid sequence comprising the sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41), and has a set of amino acid changes based on SEQ ID NO: 41, selected from the second column of the table described in claim 1; b) A buffer suitable for use in PCR reactions, which may optionally contain one or more organic cosolvents, and A kit that includes this.

19. The kit according to claim 17 or 18, wherein the one or more organic cosolvents are selected from the group consisting of amides, sulfoxides, sulfones, and diols.

20. The amide is selected from the group consisting of formamide, N-methylformamide, N,N-dimethylformamide (DMF), acetamide, N-methylacetamide, N,N-dimethylacetamide, propionamide, isobutylamide, 2-pyrrolidone, N-methylpyrrolidone (NMP), N-hydroxyethylpyrrolidone (HEP), N-formylpyrrolidine, N-formylmorpholine; δ-valerolactam, ε-caprolactam, and 2-azacyclooctanone. The sulfoxide is selected from the group consisting of dimethyl sulfoxide (DMSO), n-propyl sulfoxide, n-butyl sulfoxide, methyl sec-butyl sulfoxide, and tetramethylene sulfoxide. The sulfone is selected from the group consisting of dimethyl sulfone, diethyl sulfone, di(n-isopropyl) sulfone, tetramethylene sulfone (sulfolane), 2,4-dimethylsulfolane, and butadiene sulfone (sulfolene). The diol is selected from the group consisting of 1,2-propanediol, 1,3-propanediol, 1,2-butanediol, 1,3-butanediol, 1,4-butanediol, 1,2-pentanediol, 2,4-pentanediol, 1,5-pentanediol, 1,2-cyclopentanediol, 1,2-hexanediol, 1,6-hexanediol, and 2-methyl-2,4-pentanediol. The kit according to claim 19.

21. a) L5Q, F8L, P10S, L16P, A23P, A29T, T186I, K31R, G38D, A97T, A118V, L162P, R205K, G208S, K219E, N2 20D, I228V, D273G, S290G, K346R, P382T, E388D, E434D, A454E, L461Q, L461R, V474I, F482I, I503T A modified Taq DNA polymerase having the amino acid sequence of wild-type Taq DNA polymerase (SEQ ID NO: 41) containing one or more amino acid changes selected from the group consisting of E507K, S515N, A521V, Q534R, D551G, L606M, A608V, S612R, Q680R, K702R, E734G, S739G, E742K, F749V, F749I, F749L, K762R, K767R, L768M, Q782H, and E832K, b) A PCR buffer containing one or more low molecular weight organic solvents selected from the group consisting of amides, sulfoxides, sulfones, and diols, wherein the one or more low molecular weight organic solvents are present in a concentration range of about 0.05 to about 3.0 molars, and A composition containing the following:

22. The composition according to claim 21, wherein the amino acid change is selected from the group consisting of L5Q, P10S, A23P, A29T, T186I, L461R, E507K, A608V, S612R, E742K, F749L, F749I, K762R, K767R, and E832K.

23. The composition according to claim 21, wherein one or more low molecular weight organic solvents are present in the PCR buffer in a concentration range of about 0.1 to about 1.0 molar concentration.

24. The amide is selected from the group consisting of formamide, N-methylformamide, N,N-dimethylformamide (DMF), acetamide, N-methylacetamide, N,N-dimethylacetamide, propionamide, isobutylamide, 2-pyrrolidone, N-methylpyrrolidone (NMP), N-hydroxyethylpyrrolidone (HEP), N-formylpyrrolidine, N-formylmorpholine; δ-valerolactam, ε-caprolactam, and 2-azacyclooctanone. The sulfoxide is selected from the group consisting of dimethyl sulfoxide (DMSO), n-propyl sulfoxide, n-butyl sulfoxide, methyl sec-butyl sulfoxide, and tetramethylene sulfoxide. The sulfone is selected from the group consisting of dimethyl sulfone, diethyl sulfone, di(n-isopropyl) sulfone, tetramethylene sulfone (sulfolane), 2,4-dimethylsulfolane, and butadiene sulfone (sulfolene). The diol is selected from the group consisting of 1,2-propanediol, 1,3-propanediol, 1,2-butanediol, 1,3-butanediol, 1,4-butanediol, 1,2-pentanediol, 2,4-pentanediol, 1,5-pentanediol, 1,2-cyclopentanediol, 1,2-hexanediol, 1,6-hexanediol, and 2-methyl-2,4-pentanediol. The composition according to claim 21.