Engineered nucleases with high salt tolerance
Engineered Serratia marcescens nuclease A with mutations at specific sites enhances salt tolerance, allowing effective DNA and RNA degradation in high-salt conditions, addressing the inhibition issue and expanding its application in high-salt environments.
Patent Information
- Application Number
- JP2025536323
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-20
- Filing Date
- 2023-12-20
- Publication Date
- 2026-01-19
AI Technical Summary
Serratia marcescens nuclease A is inhibited by high salt concentrations, limiting its use in applications requiring high-salt conditions, such as virus production, necessitating additional buffer exchange steps.
Engineered polypeptides derived from Serratia marcescens nuclease A with specific amino acid mutations, particularly at sites S74, N79, A95, T98, N101, A102, S137, D138, Q141, and D149, enhancing salt tolerance to maintain nuclease activity at ionic strengths up to 500 mM.
The engineered nuclease polypeptides exhibit at least 70% nuclease activity at ionic strengths greater than 200 mM, enabling effective DNA and RNA degradation in high-salt environments.
Smart Images

Figure 2026501926000008 
Figure 2026501926000009 
Figure 2026501926000010
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the fields of protein engineering and recombinant expression technology, and in particular to altering the biochemical properties of a nuclease derived from Serratia marcescens by combining amino acid residue mutations in the primary sequence of the protein to enhance the enzymatic activity of the mutant under high salt conditions. [Background technology]
[0002] Nucleases are a type of enzyme that cleaves phosphodiester bonds in nucleic acids in an endo / exo manner. Serratia marcescens nuclease A (Uniprot: P13717) has been identified as a nonspecific nuclease that catalyzes the hydrolysis of single-stranded / double-stranded DNA and RNA by cleaving phosphodiester bonds. The crystal structure of this enzyme shows that it is a dimeric Mg 2+ It is a nucleotide-dependent nuclease that requires water clusters for its activity. It is widely used for nucleic acid removal in protein production and pharmaceutical virus production. This enzyme is also marketed by EMD Millipore Corp. under the name Benzonase®.
[0003] Although Serratia marcescens nuclease A is widely used in industry and research and development, inhibition by high salt concentrations is a significant drawback of this enzyme. The high-salt-sensitive nature of this enzyme limits its use and requires additional buffer exchange steps in the production of some pharmaceuticals. Enzyme activity is significantly inhibited at solution ionic strengths above 200 mM. However, recent studies have shown that high-salt buffers are more suitable for virus production. Therefore, high-salt-tolerant Serratia marcescens nuclease A should be more useful for numerous applications under high-salt conditions. Summary of the Invention
[0004] The present disclosure includes polypeptides having nuclease activity (hereinafter also referred to as "nucleases" or "nuclease polypeptides"), polynucleotides comprising the coding sequences for these polypeptides, and methods for making and using these polypeptides and polynucleotides. Also provided herein are compositions and kits comprising one or more nuclease polypeptides, one or more nuclease-encoding polynucleotides, and any combination thereof disclosed herein.
[0005] The present disclosure includes synthetic or recombinant polypeptides that are derived from Serratia marcescens nuclease A and have high salt tolerance. In some embodiments, the polypeptides contain one or more mutations that result in more positively charged surface area in the three-dimensional structure of the polypeptide.
[0006] In some embodiments, the polypeptide has at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the nuclease activity at a solution ionic strength of greater than 200 mM compared to a nuclease having the sequence of SEQ ID NO:1 or the mature polypeptide thereof. In some embodiments, the ion is a monovalent ion or a divalent ion. In some embodiments, the solution ionic strength is greater than 300 mM. In some embodiments, the solution ionic strength is greater than 400 mM. In some embodiments, the solution ionic strength is greater than 500 mM.
[0007] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 1 or the mature polypeptide thereof. In some embodiments, the polypeptide has nuclease activity, provided that the polypeptide comprises mutations at one, two, three, four, five, six, seven, or more or all of S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, D156. Mutations include amino acid modifications, substitutions, or deletions.
[0008] In some embodiments, the polypeptide is (a) an amino acid sequence having at least 70% identity to SEQ ID NO:1, or the mature polypeptide thereof, and having mutations at one, two, three, four, five, six, seven, or more or all of S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, D156; or (b) An amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the sequence of (a), and having a mutation at the site described in (a), and retaining nuclease activity.
[0009] In some embodiments, the mutation of N79 is to a nonpolar amino acid such as alanine, valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine, or to an uncharged polar amino acid such as serine, threonine, glutamine, tyrosine, or cysteine, or to a positively charged polar amino acid such as histidine, lysine, or arginine.
[0010] In some embodiments, the mutation of A95 is to a nonpolar amino acid such as valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine, or to an uncharged polar amino acid such as serine, threonine, glutamine, asparagine, tyrosine, or cysteine, or to a positively charged polar amino acid such as histidine, lysine, or arginine.
[0011] In some embodiments, the mutation of A102 is to a nonpolar amino acid such as valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine, or to an uncharged polar amino acid such as serine, threonine, glutamine, asparagine, tyrosine, or cysteine, or to a positively charged polar amino acid such as histidine, lysine, or arginine.
[0012] In some embodiments, the mutation of D149 is to a nonpolar amino acid such as alanine, valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine, or to an uncharged polar amino acid such as serine, threonine, glutamine, asparagine, tyrosine, or cysteine, or to a positively charged polar amino acid such as histidine, lysine, or arginine.
[0013] In some embodiments, the mutations at S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156 are to a positively charged polar amino acid, preferably histidine, lysine, or arginine, more preferably lysine or arginine.
[0014] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has mutations at at least four sites selected from the group consisting of S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156.
[0015] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1, or a mature polypeptide thereof, and has a mutation at A95 or D149 and at least three mutations selected from the group consisting of S74, N79, A95, T98, N101, A102, S137, D138, Q141, and D156. Preferably, the polypeptide has mutations at A95 and D149 and at least two mutations selected from the group consisting of S74, N79, T98, N101, A102, S137, D138, Q141, and D156.
[0016] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has mutations at at least three sites selected from the group consisting of N79, A95, A102 and D149, and optionally has a mutation at at least one site selected from the group consisting of S74, T98, N101, S137, D138, Q141, D156.
[0017] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and: (1) has mutations at A95, A102, and D149, and optionally has a mutation at at least one site selected from the group consisting of S74, N79, T98, N101, S137, D138, Q141, D156; (2) has mutations at N79, A102, and D149, and optionally has a mutation at at least one site selected from the group consisting of S74, A95, T98, N101, S137, D138, Q141, D156. (3) having mutations at N79, A95, and D149, and optionally having a mutation at at least one site selected from the group consisting of S74, T98, N101, A102, S137, D138, Q141, and D156; or (4) having mutations at N79, A95, and A102, and optionally having a mutation at at least one site selected from the group consisting of S74, T98, N101, A102, S137, D138, Q141, D149, and D156.
[0018] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1, or the mature polypeptide thereof, and has mutations at N79, A95, A102, and D149, and optionally has a mutation at at least one site selected from the group consisting of S74, T98, N101, S137, D138, Q141, and D156. In some embodiments, the polypeptide has mutations at N79, A95, A102, and D149, and optionally has a mutation at one or two sites selected from the group consisting of Q141 and N101.
[0019] In some embodiments, the polypeptide comprises an amino acid sequence at least 70% identical to SEQ ID NO:1, or the mature polypeptide thereof, and has mutations at A95, N101, Q141, D149, and optionally mutations at at least three sites selected from the group consisting of S74, N79, T98, A102, S137, D138, and D156. In some embodiments, the polypeptide has mutations at A95, N101, Q141, D149, and optionally mutations at S74, N79, and T98.
[0020] In some embodiments, the polypeptide may be a polypeptide comprising an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 1 or a mature polypeptide thereof, provided that the polypeptide has nuclease activity at high solution ionic strength, and the polypeptide is selected from the group consisting of: (1) N79K or N79R; (2) A95K or A95R; (3) A95K or A95R; (1) the nucleotide sequence of the present invention is selected from the group consisting of (1) S74K, (2) A95K or A102R, (3) D149K or D149R, (4) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R; preferably, the nucleotide sequence of the present invention is selected from the group consisting of (1) S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0021] In some embodiments, the polypeptide is (i) comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof, and having a mutation selected from the group consisting of (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R, preferably having a mutation selected from the group consisting of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, D156R; or (ii) An amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to the sequence of (i), and having the mutations described in (i), and retaining nuclease activity.
[0022] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and has at least four mutations selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R, preferably at least four mutations selected from the group consisting of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0023] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof, and contains (a) an A95K or A95R mutation and at least three (or at least four, at least five, at least five) selected from the group consisting of: (1) N79K or N79R, (2) A102K or A102R, (3) D149K or D149R, (4) S74K, (5) T98K, (6) N101K, (7) S137K, (8) D138K, (9) Q141K, and (10) D156R, preferably selected from the group consisting of S74K, N79K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R. or (b) a D149K or D149R mutation and at least three (or at least four, at least five, at least six, at least seven or more) mutations selected from the group consisting of (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) S74K, (5) T98K, (6) N101K, (7) S137K, (8) D138K, (9) Q141K and (10) D156R, preferably S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K and D156R. More preferably, the polypeptide has the mutations (i) A95K or A95R, and (ii) D149K or D149R, as well as at least two (or at least three, at least four, at least five, at least six or more) selected from the group consisting of (1) N79K or N79R, (2) A102K or A102R, (3) S74K, (4) T98K, (5) N101K, (6) S137K, (7) D138K, (8) Q141K and (9) D156R, preferably S74K, N79K, T98K, N101K, A102K, S137K, D138K, Q141K and D156R.
[0024] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has at least three mutations selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, and (4) D149K or D149R, and optionally has mutations selected from the group consisting of S74K, T98K, N101K, S137K, D138K, Q141K, and D156R.
[0025] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1, or the mature polypeptide thereof, and has (i) N79K, A95K, A102K, D149K mutations, or (ii) N79R, A95R, A102R, and D149R mutations, and optionally has mutations selected from the group consisting of S74K, T98K, N101K, S137K, D138K, Q141K, and D156R. In some embodiments, the polypeptide has (i) N79K, A95K, A102K, and D149K mutations, or (ii) N79R, A95R, A102R, and D149R mutations, and optionally has mutations selected from the group consisting of Q141K and N101K.
[0026] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1, or the mature polypeptide thereof, and has the following mutations: A95K, N101K, Q141K, D149K, and optionally, a mutation selected from the group consisting of S74K, N79K, T98K, A102K, S137K, D138K, and D156R. In some embodiments, the polypeptide has the following mutations: A95K, N101K, Q141K, D149K, and optionally, S74K, N79K, and T98K.
[0027] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has the following mutations: S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0028] In some embodiments, the polypeptide has an optimum temperature of 30°C to 60°C. In some embodiments, the polypeptide has an optimum pH of pH 4 to pH 11. In some embodiments, the polypeptide does not comprise a signal sequence. In some embodiments, the polypeptide further comprises a signal sequence. In some embodiments, the signal sequence is a heterologous sequence or a native signal sequence.
[0029] The present disclosure also includes compositions or kits comprising one or more polypeptides disclosed herein. In some embodiments, the composition is a reaction mixture, a detergent composition, a detergent additive, a food, a food supplement, a feed supplement, a feed, a pharmaceutical composition, a fermentation product, a fermentation intermediate, a fermentation downstream reaction mixture, or a combination thereof. In some embodiments, the reaction mixture comprising one or more polypeptides is for the expression or purification of a protein or a viral vector. For example, the polypeptide may be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or the mature polypeptide thereof, provided that the polypeptide has nuclease activity at high solution ionic strength and the polypeptide comprises one or more of the following mutations: S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0030] The present disclosure also includes a nucleic acid comprising (1) a sequence encoding any one of the polypeptides disclosed herein, or a complementary sequence thereof, or (2) a sequence having at least 50%, 60%, 70%, 80%, or 90% identity to (1).
[0031] The present disclosure also includes nucleic acid constructs comprising the nucleic acid sequences disclosed herein. In one or more embodiments, the nucleic acid construct is a cloning vector, an expression vector, or a recombinant vector. In one or more embodiments, the nucleic acid sequence is operably linked to an expression control sequence. In some embodiments, the expression vector comprises a viral vector, a plasmid, a phage, a phagemid, a cosmid, a fosmid, a bacteriophage, an artificial chromosome, or a combination thereof.
[0032] The present disclosure includes recombinant cells comprising one or more polypeptides disclosed herein, one or more nucleic acids encoding any one of the polypeptides disclosed herein, one or more nucleic acid constructs comprising the polynucleotide sequence of a nucleic acid, or a combination thereof. In some embodiments, the nucleic acid is part of a chromosome of the recombinant cell. In some embodiments, the recombinant cell is a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, an insect cell, or a plant cell. For example, the polypeptide may be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or a mature polypeptide thereof, provided that the polypeptide has nuclease activity at high solution ionic strength, and the polypeptide comprises one or more mutations of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0033] The present disclosure includes methods for producing a polypeptide having nuclease activity at high solution ionic strength. In some embodiments, the method includes expressing a nucleic acid encoding any one of the polypeptides disclosed herein under conditions that allow expression of the polypeptide, thereby producing a recombinant polypeptide having nuclease activity, provided that the nucleic acid is operably linked to a promoter. In some embodiments, the nucleic acid is present in an expression vector. In some embodiments, the nucleic acid is present in a host cell to allow expression of the polypeptide. In some embodiments, the nucleic acid is present in a chromosome of the host cell. In some embodiments, the host cell is a cell derived from an organism selected from the group consisting of Pichia pastoris, Bacillus subtilis, Pseudomonas fluorescens, Myceliopthora thermophile fungus, Trichoderma reesei, Escherichia coli, Bacillus licheniformis, Aspergillus niger, Schizosaccharomyces pombe, and Sacaramyces cerevisiae. In some embodiments, the nucleic acid is present in an in vitro expression system. The polypeptide may be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO:1 or the mature polypeptide thereof, provided that the polypeptide has nuclease activity at high solution ionic strength, and the polypeptide comprises one or more mutations of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0034] The present disclosure includes methods for degrading polynucleotides, comprising contacting a polynucleotide molecule with one or more polypeptides disclosed herein, thereby degrading the polynucleotide molecule. For example, the polypeptide may be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO:1 or a mature polypeptide thereof, provided that the polypeptide has nuclease activity at high solution ionic strength and the polypeptide comprises one or more of the following mutations: S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R. In some embodiments, the polynucleotide molecule is a DNA molecule or an RNA molecule. In some embodiments, the contacting is performed at a pH between 4 and 11. In some embodiments, the temperature of the reaction mixture is between about 10°C and about 70°C. In some embodiments, the contacting is performed between 30°C and 60°C.
[0035] The present disclosure also includes a method for degrading DNA or RNA during cell lysis. In some embodiments, the method includes lysing a host cell of interest and adding a polypeptide disclosed herein under conditions that allow the polypeptide to degrade the DNA or RNA. For example, the host cell may be a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, a plant cell, or an insect cell.
[0036] The present disclosure also includes methods for degrading DNA or RNA during viral vector production. In some embodiments, the methods include culturing host cells containing the viral vector of interest and expressing or adding one or more polypeptides disclosed herein under conditions that allow the one or more polypeptides to degrade DNA or RNA other than the viral vector. In some embodiments, the host cells are bacterial cells, mammalian cells, fungal cells, yeast cells, plant cells, or insect cells. In some embodiments, the polypeptides are expressed from an expression vector present in the host cells, or the polypeptides are encoded by a nucleic acid sequence in a chromosome of the host cells. In some embodiments, the virus is an adeno-associated virus (AAV) or a lentivirus.
[0037] In some embodiments, the polypeptide is expressed by cells that do not express the viral vector of interest. In some embodiments, expression of the viral vector of interest and the polypeptide is inducible or non-inducible. In some embodiments, one or more polypeptides disclosed herein are added exogenously. For example, the polypeptide may be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or a mature polypeptide thereof, provided that the polypeptide has nuclease activity at high solution ionic strength and the polypeptide comprises one or more of the following mutations: S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0038] The present disclosure also includes methods for degrading DNA or RNA during protein production. In some embodiments, the methods include culturing host cells containing nucleic acid encoding a protein of interest and expressing or adding one or more polypeptides disclosed herein under conditions that allow the one or more polypeptides to degrade the DNA or RNA. In some embodiments, the host cell is a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, a plant cell, or an insect cell. In some embodiments, the polypeptide is expressed from an expression vector present in the host cell, or the polypeptide is encoded by a nucleic acid sequence in a chromosome of the host cell.
[0039] In some embodiments, the polypeptide is expressed by a cell that does not express the protein of interest. In some embodiments, expression of one or more proteins and polypeptides of interest is inducible or non-inducible. In some embodiments, one or more polypeptides disclosed herein are exogenously added. For example, the polypeptide may be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or a mature polypeptide thereof, provided that the polypeptide has nuclease activity at high solution ionic strength and the polypeptide comprises one or more of the following mutations: S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0040] The present disclosure also includes a reaction mixture comprising (a) one or more polypeptides disclosed herein, (b) one or more nucleic acid molecules, and (c) an aqueous solution in which the polypeptide hydrolyzes the one or more nucleic acid molecules. In some embodiments, the one or more nucleic acid molecules comprise single-stranded DNA molecules, double-stranded DNA molecules, single-stranded RNA molecules, double-stranded RNA molecules, or any combination thereof. In some embodiments, the one or more nucleic acid molecules are derived from a host cell for protein production. In some embodiments, the polypeptide is expressed in a host cell selected from the group consisting of bacterial cells, mammalian cells, fungal cells, yeast cells, and insect cells. In some embodiments, the temperature of the reaction mixture is about 10°C to about 70°C. In some embodiments, the temperature of the reaction mixture is about 30°C to about 60°C. In some embodiments, the reaction mixture is at a pH of about 4 to about 11. In some embodiments, the reaction mixture is at a pH of about 7 to about 9. In some embodiments, the aqueous solution is a detergent composition, a detergent additive, a food, a food supplement, a feed supplement, a feed, a pharmaceutical composition, a fermentation product, a fermentation intermediate, a fermentation downstream reaction mixture, a product from a protein production process, an intermediate from a protein production process, a protein purification solution, or a combination thereof. For example, the polypeptide may be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or a mature polypeptide thereof, provided that the polypeptide has nuclease activity at high solution ionic strength, and the polypeptide comprises one or more of the following mutations: S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0041] The present disclosure includes methods for degrading DNA or RNA in a protein production mixture. In some embodiments, the methods include culturing host cells containing a nucleic acid encoding a protein of interest and expressing one or more polypeptides disclosed herein under conditions that allow the polypeptide to degrade the DNA or RNA. In some embodiments, expression of the polypeptide may be delayed until after production of the protein of interest. In some embodiments, expression of the polypeptide is not delayed. Expression of the polypeptide may begin before, after, or simultaneously with the initiation of expression of the protein of interest. For example, the host cell may be a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, a plant cell, or an insect cell. In some embodiments, the polypeptide is expressed from an expression vector present in the host cell, or the polypeptide is encoded by a nucleic acid sequence in a chromosome of the host cell. In some embodiments, the polypeptide is expressed by a cell that does not express the protein of interest. Expression of the one or more proteins of interest and polypeptides may be inducible or non-inducible. [Brief explanation of the drawings]
[0042] [Figure 1] Electrostatic potential surfaces of wild-type NucA from Serratia marcescens (left) and an engineered NucA with more positively charged residues (HighSalt NucA) (right). White represents uncharged surfaces, and darker shaded areas represent charged surfaces. Compared to wild-type NucA, the engineered HighSalt NucA has more positively charged surface area. [Figure 2] Digestion of λ DNA by wild-type NucA and engineered NucA (HighSalt NucA) from Serratia marcescens at different NaCl concentrations. When the salt concentration increases to 300 mM, the enzymatic activity of wild-type NucA is significantly inhibited. However, the engineered HighSalt NucA can efficiently digest λ DNA at 500 mM NaCl. [Figure 3]Digestion of plasmid DNA by wild-type NucA and engineered NucA (HighSalt NucA) from Serratia marcescens at different NaCl concentrations. When the salt concentration increases to 300 mM, the enzymatic activity of wild-type NucA is significantly inhibited. However, the engineered HighSalt NucA can efficiently digest plasmid DNA at 500 mM NaCl. [Figure 4] Digestion of total RNA extracted from CHO cells by wild-type NucA and engineered NucA (HighSalt NucA) from Serratia marcescens at different NaCl concentrations. When the salt concentration increases to 300 mM, the enzymatic activity of wild-type NucA is significantly inhibited. However, the engineered HighSalt NucA can efficiently digest RNA at 500 mM NaCl. [Figure 5] Digestion of single-stranded DNA (ssDNA) by wild-type NucA and engineered NucA (HighSalt NucA) from Serratia marcescens at different NaCl concentrations. When the salt concentration increases to 300 mM, the enzymatic activity of wild-type NucA is greatly inhibited. However, the engineered HighSalt NucA can efficiently digest ssDNA at 500 mM NaCl. [Figure 6] Inhibitory effect of monovalent salts (left: NaCl, right: KCl) on HighSalt NucA and wild-type nuclease A. [Figure 7] Inhibitory effect of divalent salts (left: MgCl2, right: MnCl2) on HighSalt NucA and wild-type nuclease A. [Figure 8] Inhibitory effect of (NH4)2SO4 on HighSalt NucA and wild-type nuclease A. [Figure 9] Inhibitory effect of Na2HPO4 on HighSalt NucA and wild-type nuclease A. [Figure 10] Salt tolerance of wild-type NucA, HighSalt NucA, and 19 mutants with single mutations in the N79 residue. [Figure 11] Salt tolerance of wild-type NucA, HighSalt NucA, and 19 mutants with single mutations in the A95 residue. [Figure 12] Salt tolerance of wild-type NucA, HighSalt NucA, and 19 mutants with single mutations in the A102 residue. [Figure 13] Salt tolerance of wild-type NucA, HighSalt NucA, and 19 mutants with single mutations in the D149 residue. [Figure 14] Salt tolerance of wild-type NucA, HighSalt NucA, and mutants with combined mutations of positively charged residues. DETAILED DESCRIPTION OF THE INVENTION
[0043] All patents, applications, published applications and other publications mentioned herein are incorporated by reference for the material referenced and in their entirety. If a term or phrase used herein is contrary to or otherwise inconsistent with a definition set forth in a patent, application, published application or other publication incorporated by reference herein, the usage herein shall take precedence over the definition incorporated by reference herein.
[0044] definition Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. In the event that there are multiple definitions for terms herein, those in this section prevail unless stated otherwise.
[0045] As used herein, the singular forms "a," "one," and "said" include plural references unless expressly or by context dictates otherwise. For example, "a" dimer includes one or more dimers unless expressly or by context dictates otherwise.
[0046] The term "amplification" ("polymerase extension reaction") means increasing the number of copies of a polynucleotide.
[0047] As used herein, "sequence identity" or "identity" in the context of two protein sequences (or nucleotide sequences) includes reference to residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window.
[0048] Sequence identity is usually provided as "sequence identity %" or "identity %". To determine the percent identity between two amino acid sequences, the first step is to generate a pairwise sequence alignment between the two sequences, where the two sequences are aligned over their entire length (i.e., pairwise global alignment). The alignment is generated using a program that implements the Needleman and Wunsch algorithm (J. Mol. Biol. (1979) 48, pp. 443-453), and such programs are known to those skilled in the art, such as "NEEDLE". For the purposes of this specification, the preferred alignment is the one that allows the highest sequence identity to be determined.
[0049] After aligning the two sequences, the second step is to determine an identity value from the resulting alignment. For purposes herein, percent identity is calculated as follows: % identity = (identical residues / length of the aligned region representing each sequence herein over its entire length) * 100.
[0050] Thus, the sequence identity associated with a comparison of two amino acid sequences according to this embodiment is calculated by dividing the number of identical residues by the length of the alignment region representing each sequence herein over its entire length, and multiplying this value by 100 to arrive at the "% identity."
[0051] To calculate the percent identity of two DNA sequences, the same rules apply as for calculating the percent identity of two amino acid sequences, with some specifications.
[0052] For DNA sequences encoding proteins, pairwise alignment should be performed over the entire length of the coding region from the start codon to the stop codon, excluding introns. Introns present in the other sequence being compared to the sequences herein may also be removed for pairwise alignment. Percent identity is calculated as follows: % identity = (identical residues / length of the aligned region showing the coding region from the start codon to the stop codon of the sequences herein, excluding introns) * 100.
[0053] Sequences having regions identical or similar to the sequences herein and which can be compared to the sequences herein to determine percent identity can be readily identified by a variety of methods known to those of skill in the art, including, for example, using publicly available computer methods and programs such as BLAST, available at NCBI and elsewhere.
[0054] Variants of parent enzyme molecules may have an amino acid sequence that is at least n% identical to the amino acid sequence of the respective parent enzyme that has enzymatic activity, compared to the full-length polypeptide sequence, where n is an integer between 50 and 100, preferably 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99. Preferably, variant enzymes that are n% identical to the parent enzyme have enzymatic activity.
[0055] Enzyme variants can be defined by their sequence similarity compared to the parent enzyme. Sequence similarity is usually provided as "sequence similarity %" or "similarity %." To calculate sequence similarity, in the first step, a sequence alignment must be generated as described above. In the second step, the sequence similarity percentage must be calculated, taking into account that a defined set of amino acids share similar properties, such as their size, hydrophobicity, charge, or other characteristics. Here, the replacement of an amino acid with a similar amino acid is called a "conservative mutation." Enzyme variants containing conservative mutations have minimal impact on protein folding, resulting in the substantial maintenance of some enzymatic properties compared to those of the parent enzyme.
[0056] Conservative amino acid substitutions can occur throughout the entire polypeptide sequence of a functional protein, such as an enzyme. In one embodiment, such mutations do not involve functional domains of the enzyme. In one embodiment, conservative mutations do not involve the catalytic center of the enzyme.
[0057] For example, amino acid A is similar to amino acid S; amino acid D is similar to amino acids E and N; amino acid E is similar to amino acids D, K, and Q; amino acid F is similar to amino acids W and Y; amino acid H is similar to amino acids N and Y; amino acid I is similar to amino acids L, M, and V; amino acid K is similar to amino acids E, Q, and R; amino acid L is similar to amino acids I, M, and V; amino acid M is similar to amino acids I, L, and V; amino acid N is similar to amino acids D, H, and S; amino acid Q is similar to amino acids E, K, and R; amino acid R is similar to amino acids K and Q; amino acid S is similar to amino acids A, N, and T; amino acid T is similar to amino acid S; amino acid V is similar to amino acids I, L, and M; amino acid W is similar to amino acids F and Y; and amino acid Y is similar to amino acids F, H, and W.
[0058] In particular, variant enzymes containing conservative mutations that are at least m% similar to their respective parent sequences compared to the full-length polypeptide sequence, where m is an integer between 50 and 100, preferably 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99, are expected to have essentially unchanged enzymatic properties. Preferably, variant enzymes that are m% similar compared to the parent enzyme have enzymatic activity.
[0059] Homology refers to a high degree of similarity between genes, polypeptides, and polynucleotides in terms of position, structure, function, or characteristics, but does not necessarily mean that the sequence identity is high.
[0060] As used herein, "substantially complementary or substantially matched" means that two nucleic acid sequences have at least about 90% sequence identity. Preferably, two nucleic acid sequences have at least, or at least about, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity. Alternatively, "substantially complementary or substantially matched" means that two nucleic acid sequences are capable of hybridizing under high stringency conditions.
[0061] As defined herein, the term "hybridization" refers to the process by which substantially complementary nucleotide sequences anneal to each other. The hybridization process may be performed entirely in solution, i.e., under conditions in which both complementary nucleic acids are in solution. The hybridization process may also be performed using one of the complementary nucleic acids immobilized on a matrix, such as magnetic beads, Sepharose beads, or other resins. The hybridization process may also be performed using one of the complementary nucleic acids immobilized on a solid support, such as a nitrocellulose or nylon membrane, or immobilized on, for example, a siliceous glass support, for example, by photolithography (the latter known as a nucleic acid array or microarray, or nucleic acid chip). To allow hybridization, nucleic acid molecules are typically thermally or chemically denatured to melt the duplex into two single strands and / or remove hairpins or other secondary structures from single-stranded nucleic acids. Hybridization as used herein means that hybridization must occur over the entire length of the sequences of the present invention. As defined herein, full-length hybridization means that when the sequences herein are fragmented into fragments of 300 to 500 bases, each fragment hybridizes.
[0062] The term "stringency" refers to the conditions under which hybridization occurs. Hybridization stringency is influenced by conditions such as temperature, salt concentration, ionic strength, and hybridization buffer composition. Generally, low stringency conditions are selected to be approximately 30°C lower than the thermal melting point (Tm) of a specific sequence at a defined ionic strength and pH. Moderate stringency conditions are those in which the temperature is 20°C lower than the Tm, and high stringency conditions are those in which the temperature is 10°C lower than the Tm. High stringency hybridization conditions are typically used to isolate hybridizing sequences with high sequence identity to the target nucleic acid sequence. However, due to the degeneracy of the genetic code, nucleic acids may encode substantially identical polypeptides despite mismatches in sequence. Therefore, moderate stringency hybridization conditions may be required to identify such nucleic acid molecules. "Tm" is the temperature at which 50% of a target sequence hybridizes to a perfectly matched probe at a defined ionic strength and pH. The Tm depends on the solution conditions, the base composition, and the length of the probe. For example, longer sequences hybridize specifically at higher temperatures. The maximum rate of hybridization is achieved at approximately 16°C–32°C below the Tm. The presence of monovalent cations in the hybridization solution reduces electrostatic repulsion between the two nucleic acid strands, promoting hybrid formation; this effect is observed at sodium concentrations up to 0.4M (at higher concentrations, the effect is negligible). Formamide reduces the melting temperature of DNA-DNA and DNA-RNA duplexes by 0.6–0.7°C per 1% formamide; adding 50% formamide reduces the rate of hybridization but allows hybridization at 30–45°C. Base pair mismatches decrease the rate of hybridization and reduce the thermal stability of the duplex. On average, for large probes, the Tm decreases by approximately 1°C per 1% base mismatch. Tm can be calculated according to routine knowledge in the art.In addition to hybridization conditions, the specificity of hybridization generally also depends on the post-hybridization washes. Those skilled in the art are aware of the various parameters that can be altered during washing to maintain or change stringency conditions.
[0063] For example, typical high stringency hybridization conditions for DNA hybrids longer than 50 nucleotides include hybridization in 1xSSC at 65°C, or in 1xSSC and 50% formamide at 42°C, followed by a wash in 0.3xSSC at 65°C.
[0064] For purposes of defining stringency levels, reference may be made to Sambrook et al. (2001) Molecular Cloning: a laboratory manual, 3rd Edition, Cold Spring Harbor Laboratory Press, CSH, New York, or Current Protocols in Molecular Biology, John Wiley & Sons, NY (1989 and annual updates).
[0065] As used herein, a "primer" refers to a nucleic acid molecule that can anneal to a template nucleic acid and function as a starting point for DNA amplification. A primer may be fully or partially complementary to a specific region of a template polynucleotide, for example, 20 nucleotides upstream or downstream from a target codon. Non-complementary nucleotides are defined herein as mismatches. Mismatches may be located within the primer or at both ends of the primer. Preferably, one nucleotide mismatch, more preferably two, and even more preferably three or more consecutive or non-consecutive nucleotide mismatches are located within the primer. A primer may have, for example, 5 to 200 nucleotides, preferably 20 to 80 nucleotides, and more preferably 43 to 65 nucleotides. More preferably, the primer has 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, or 190 nucleotides. As defined herein, a "forward primer" is a primer that is complementary to the minus strand of a template polynucleotide. As defined herein, a "reverse primer" is a primer that is complementary to the plus strand of a template polynucleotide. Preferably, the forward primer and the reverse primer do not contain overlapping nucleotide sequences. As defined herein, "does not contain overlapping nucleotide sequences" means that the forward primer and reverse primer do not anneal to regions of the minus and plus strands, respectively, of a template polynucleotide where the plus and minus strands are complementary to each other. With respect to primers that anneal to the same strand of a template polynucleotide, "does not contain overlapping nucleotide sequences" means that they do not contain sequences that are complementary to the same region of the same strand of the template polynucleotide. As used herein, a "primer set" refers to a combination of a "forward primer" and a corresponding "reverse primer."
[0066] As used herein, the plus strand corresponds to the sense strand and is sometimes called the coding strand or non-template strand. It is the strand that has the same sequence as the mRNA (except that T's replace U's). The other strand is called the template, minus, or antisense strand and is complementary to the mRNA.
[0067] As described herein, "codon optimization" refers to a design process in which codons are changed to codons known to maximize protein expression efficiency. In some alternatives, codon optimization for expression in cells is described, although codon optimization can be performed using algorithms known to those skilled in the art to create synthetic gene transcripts optimized for high mRNA and protein yields in a host cell of interest, such as, for example, a bacterial, fungal, insect, or mammalian cell (including a human cell). Codons may be optimized for protein expression in, for example, bacterial cells, mammalian cells, yeast cells, insect cells, plant cells, etc. Programs containing algorithms for codon optimization in human cells are readily available. In some embodiments, a gene is codon-optimized for expression in a bacterial, yeast, fungal, or insect cell.
[0068] The term "heterologous" (or foreign or recombinant) polypeptide is defined herein as: (a) a polypeptide that is not native to the host cell; the protein sequence of such a heterologous polypeptide is a synthetic, non-naturally occurring, "artificial" protein sequence; (b) a polypeptide that is native to the host cell but has undergone structural modifications, such as deletions, substitutions, and / or insertions, to alter the native polypeptide; or c) a polypeptide that is native to the host cell but whose expression has been quantitatively altered as a result of manipulation of the host cell's DNA by recombinant DNA techniques, such as stronger promoters, or whose expression is derived from a different genomic location than in the native host cell.
[0069] Explanations b) and c) above refer to sequences that are native but not naturally expressed by the cells used to produce them. Thus, the produced polypeptide is more accurately defined as a "recombinantly expressed endogenous polypeptide," which does not contradict the above definition and reflects the special situation of a method for producing a polypeptide molecule, rather than a synthetic or engineered protein sequence.
[0070] Similarly, the term "heterologous" (or foreign or recombinant) polynucleotide refers to: (a) a polynucleotide that is not native to a host cell; (b) a polynucleotide that is native to a host cell that has undergone structural modifications, such as deletions, substitutions, and / or insertions, to alter the native polynucleotide; (c) a polynucleotide that is native to a host cell whose expression has been quantitatively altered as a result of manipulation of the polynucleotide's regulatory elements by recombinant DNA techniques, such as a stronger promoter; or (d) a polynucleotide that is native to a host cell that has been genetically engineered by recombinant DNA techniques so that it is not integrated into its natural genetic environment.
[0071] With respect to two or more polynucleotide sequences or two or more amino acid sequences, the term "heterologous" is used to characterize the two or more polynucleotide sequences or two or more amino acid sequences as not occurring in nature in that particular combination with each other.
[0072] As used herein, "transgenic," "transgenic," or "recombinant" refers to, for example, a nucleic acid sequence, an expression cassette, a genetic construct, or a vector containing a nucleic acid sequence, or an organism transformed with a nucleic acid sequence, expression cassette, or vector, all of which constructs have been synthetically produced by recombinant or genetic engineering methods, provided that (a) the nucleic acid sequence containing the desired genetic information to be expressed, or (b) a genetic control sequence (e.g., a promoter) operably linked to the nucleic acid sequence containing the desired genetic information, or (c) neither (a) nor (b) is located in their natural genetic environment or has been modified by recombinant methods. Natural genetic environment is understood to mean the natural genomic or chromosomal locus in the original organism. A naturally occurring expression cassette (e.g., the naturally occurring combination of a nucleic acid sequence's native promoter and the corresponding nucleic acid sequence encoding a polypeptide) becomes a transgenic expression cassette when the expression cassette is modified by human intervention, e.g., mutagenic treatment. Furthermore, a naturally occurring expression cassette becomes a recombinant expression cassette when the expression cassette is separated from its natural genetic environment and then reintroduced into a genetic environment that is not its natural genetic environment.
[0073] "Synthetic" or "artificial" compounds are produced by chemical or enzymatic synthesis and include, but are not limited to, mutant nucleic acids made with optimal codon usage for a host organism, such as a yeast cell host or other expression host of choice, or mutant protein sequences with amino acid modifications, such as substitutions, compared to a parent protein sequence, e.g., to optimize the properties of the polypeptide.
[0074] A "reference sequence" is a defined sequence used as a basis for sequence comparison; it may be a subset of a larger sequence, for example, a segment of a full-length cDNA or gene sequence set forth in a sequence listing, or it may include the entire cDNA or gene sequence. Generally, the length of a reference sequence is at least 20 nucleotides, frequently at least 25 nucleotides, and often at least 50 nucleotides.
[0075] The terms "fragment," "derivative," and "analog" when referring to a reference polypeptide include polypeptides that retain at least one biological function or activity that is at least substantially the same as the reference polypeptide.
[0076] The term "functional fragment" refers to a nucleic acid sequence or amino acid sequence that contains only a portion of a full-length nucleic acid sequence or full-length amino acid sequence, respectively, but still has the same or similar activity and / or function. In one embodiment, the fragment contains at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the original sequence. In one embodiment, the functional fragment contains contiguous nucleic acids or amino acids compared to the original nucleic acid or original amino acid sequence, respectively.
[0077] The term "gene" means a segment of DNA involved in producing a polypeptide chain, including regions preceding and following the coding region (leader and trailer), and any intervening sequences (introns) between individual coding segments (exons).
[0078] As used herein, the term "isolated" means that a material is removed from its original environment (e.g., the natural environment if it occurs in nature). For example, a polynucleotide or enzyme naturally occurring in a living animal is not isolated, but the same polynucleotide or enzyme becomes isolated when separated from some or all of the coexisting materials in the natural system. Such a polynucleotide may be part of a vector and / or such a polynucleotide or enzyme may be part of a composition, but such a vector or composition is still isolated in that it is not part of its natural environment. As a further example, an isolated nucleic acid, e.g., a DNA or RNA molecule, is one that is not immediately contiguous with the 5' and 3' flanking sequences that are normally immediately contiguous with it when present in the naturally occurring genome of the organism from which it is derived. Such polynucleotides may be part of a vector, may be integrated into the genome of a cell of an unrelated genetic background (or, if integrated into the genome of a cell of a substantially similar genetic background, may be integrated into a site different from the site where it naturally occurs), or may be produced by PCR amplification or restriction enzyme digestion, or may be an RNA molecule produced by in vitro transcription, and / or such polynucleotides, polypeptides, or enzymes may be part of a composition, but such vector or composition is nevertheless isolated in that it is not part of its natural environment.
[0079] The term "isolated" refers to DNA incorporated into a vector, such as a plasmid or viral vector; nucleic acid incorporated into the genome of a heterologous cell (or into the genome of a homologous cell, but at a site where it does not naturally occur); and nucleic acid that exists as a separate molecule, e.g., a DNA fragment produced by PCR amplification or restriction enzyme digestion, or an RNA molecule produced by in vitro transcription.
[0080] As used herein, the term "purified" does not require absolute purity; rather, it is intended as a relative definition. Individual nucleic acids obtained from a library have traditionally been purified to electrophoretic homogeneity. For example, purified nucleic acids of the present disclosure may be at least 10- to 10-fold purified from the remainder of the genomic DNA in the organism. However, the term "purified" also encompasses nucleic acids that have been purified by at least one order of magnitude, typically two or three orders of magnitude, and more typically four or five orders of magnitude, from the remainder of the genomic DNA or from other sequences in the library or other environment. "Purified" means that the material is in a relatively pure state, e.g., at least about 90% pure, at least about 95% pure, or at least about 98% or 99% pure. Preferably, "purified" means that the material is in a 100% pure state.
[0081] The term "operably linked" means that the components described are in a relationship permitting them to function in their intended manner. For example, a regulatory sequence operably linked to a coding sequence is ligated in such a way that expression of the coding sequence is achieved under conditions compatible with the control sequences. As used herein, a promoter sequence is "operably linked" to a coding sequence if it is capable of transcribing the coding sequence into mRNA by RNA polymerase that initiates transcription at the promoter.
[0082] The term "mutation" is defined as an alteration in the genetic code of a nucleic acid sequence or an alteration in the sequence of a peptide. Such a mutation may be a point mutation, such as a transition or translocation. A mutation may be a change in one or more nucleotides or in the encoded amino acid sequence. A mutation may be a deletion, insertion, or duplication.
[0083] The terms "polynucleotide," "nucleic acid sequence," "nucleotide sequence," "nucleic acid," and "nucleic acid molecule" are used interchangeably herein and refer to a polymeric, unbranched form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides, or a combination of both.
[0084] The term "nucleic acid sequence" or "DNA coding sequence" or "nucleotide sequence" encoding a particular protein or polypeptide refers to a DNA sequence that is transcribed and translated into a protein or polypeptide when placed under the control of appropriate regulatory sequences.
[0085] The terms "nucleic acid encoding a protein or peptide" or "DNA encoding a protein or peptide" or "polynucleotide encoding a protein or peptide" and other synonyms encompass polynucleotides that contain only the coding sequence for a protein or peptide, as well as polynucleotides that contain additional coding and / or non-coding sequences.
[0086] The terms "regulatory element," "control sequence," and "promoter" are all used interchangeably herein and are interpreted in a broad context to refer to a regulatory nucleic acid sequence capable of effectively effecting the expression of an associated sequence. As used herein, a "regulatory element" or "regulatory nucleotide sequence" may refer to a fragment of nucleic acid that drives the expression of a nucleic acid sequence upon transformation into a host cell or organelle. Regulatory nucleotide sequences may include any nucleotide sequence that has a function or purpose, both individually and in a particular arrangement or group of other elements or sequences within that arrangement. Examples of regulatory nucleotide sequences include, but are not limited to, transcription control elements such as promoters, enhancers, and termination elements. Regulatory nucleotide sequences may be native (i.e., from the same gene) or foreign (i.e., from a different gene) to the nucleotide sequence to be expressed.
[0087] The term "promoter" generally refers to a nucleic acid control sequence located upstream of the transcription start site of a gene and involved in the recognition and binding of RNA polymerase and other proteins, thereby directing transcription of an operably linked nucleic acid. As used herein, "promoter" may further include any nucleic acid sequence capable of driving transcription of a coding sequence. In particular, as used herein, the term "promoter" may refer to a polynucleotide sequence generally described as the 5' regulatory region of a gene, located proximal to the start codon. Transcription of one or more coding sequences is initiated at the promoter region. The term promoter may also include a fragment of a promoter functional to initiate transcription of a gene. A promoter may also be referred to as a "transcription start site" (TSS).
[0088] The foregoing term encompasses additional transcriptional regulatory sequences derived from classical eukaryotic genomic genes (including the TATA box, with or without a CCAAT box sequence, required for accurate transcription initiation), and additional regulatory elements (i.e., upstream activating sequences, enhancers and silencers) that modify gene expression in response to developmental and / or external stimuli or in a tissue-specific manner.
[0089] For example, enhancers, as known in the art and used herein, are typically short segments of DNA (e.g., 50-1500 bp) that may be bound by proteins, such as transcription factors, to increase the likelihood that transcription of a coding sequence will occur.
[0090] Additional elements may be "transcription termination elements," which comprise segments of nucleic acid sequences that mark the ends of genes and mediate transcription termination by providing a signal within the mRNA that initiates release of the mRNA from the transcription complex. Transcription termination in prokaryotes and eukaryotes is conventional knowledge in the art.
[0091] "Oligonucleotide" (or synonymously, "oligo") refers to either a single stranded polydeoxynucleotide or two complementary polydeoxynucleotide strands, which may be chemically synthesized. Such synthetic oligonucleotides may or may not have a 5' phosphate.
[0092] The starting nucleic acid (also defined as a "template polynucleotide") can be any purified source of nucleic acid. Thus, DNA or RNA, including messenger RNA, can be used in this process, and the DNA or RNA can be single-stranded or preferably double-stranded. Furthermore, DNA-RNA hybrids containing one of each strand can also be used. The nucleic acid sequence can have a variety of lengths depending on the size of the nucleic acid sequence to be mutated. Preferably, the specific nucleic acid sequence is 50 to 50,000 base pairs, more preferably 50 to 11,000 base pairs.
[0093] All methods and materials similar or equivalent to those described herein can be used in the practice or testing of the methods and compositions disclosed herein, and suitable methods and materials are described herein. All publications, patent applications, patents, and other materials mentioned herein are incorporated by reference in their entirety. Furthermore, the materials, methods, and examples are illustrative only and, unless otherwise specified, are not intended to be limiting.
[0094] Nuclease Polypeptides This specification discloses polypeptides with nuclease activity and salt tolerance. To overcome the drawbacks of nucleases under high salt concentrations, the present inventors engineered Serratia marcescens nuclease by introducing positively charged amino acid residues into the nuclease to reshape its electrostatic surface. After screening, the nuclease mutants were found to be able to tolerate at least 200 mM NaCl without a significant decrease in enzymatic activity. This mutant can broaden its application range to higher salt solutions.
[0095] Serratia marcescens nuclease A is an extracellular enzyme widely used for nucleic acid removal and has been used in many applications. However, this enzyme is salt-sensitive, and its enzymatic activity is significantly inhibited at solution ionic strengths above 200 mM. Analysis of the crystal structure of nuclease A revealed that the interaction between nuclease A and nucleic acids is primarily dependent on electrostatic interactions. At high salt concentrations, these interactions are lost, resulting in the inhibition of nuclease A enzymatic activity under high-salt conditions. Therefore, enhancing the nucleic acid-protein interaction is an effective way to improve the salt tolerance of nuclease A. Based on this principle, we engineered Serratia marcescens nuclease A by introducing more positively charged residues into the predicted nucleic acid-binding surface, including residues S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156. Mutation of these residues to specific amino acids increases the positive charge on the nucleic acid-binding surface, promoting nucleic acid binding. It should be noted that improved nucleic acid binding may also affect the dissociation of cleaved nucleic acids from the protein. Furthermore, excessive nucleic acid binding may inhibit enzymatic activity. Therefore, the balance between nucleic acid affinity and enzymatic activity must be impaired. We designed six mutants by combining mutations in the above 11 residues and expressed and purified all of them. We compared the electrostatic potential surfaces of wild-type and mutant nucleases A and found that the mutant nucleases A indeed have an expanded positively charged surface (Figure 1). Rational engineering of Serratia marcescens nuclease A by introducing positively charged residues into specific regions of its surface significantly improved the enzyme's tolerance to high-salt conditions.
[0096] In some embodiments, the N79 mutation is to a nonpolar amino acid such as alanine, valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine, or to an uncharged polar amino acid such as serine, threonine, glutamine, tyrosine, or cysteine, or to a positively charged polar amino acid such as histidine, lysine, or arginine. In certain embodiments, the N79 mutation is to Ala, Cys, Gly, Gln, Ser, Thr, Val, Trp, Tyr, Lys, or Arg.
[0097] In some embodiments, the A95 mutation is to a nonpolar amino acid such as valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine, or to an uncharged polar amino acid such as serine, threonine, glutamine, asparagine, tyrosine, or cysteine, or to a positively charged polar amino acid such as histidine, lysine, or arginine. In certain embodiments, the A95 mutation is to Phe, Asn, Pro, Gln, Ser, Thr, Val, Tyr, Lys, or Arg.
[0098] In some embodiments, the mutation at A102 is to a nonpolar amino acid such as valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine, or to an uncharged polar amino acid such as serine, threonine, glutamine, asparagine, tyrosine, or cysteine, or to a positively charged polar amino acid such as histidine, lysine, or arginine. In certain embodiments, the mutation at A102 is to Thr, Val, Trp, Tyr, Lys, or Arg.
[0099] In some embodiments, the mutation at D149 is to a nonpolar amino acid such as alanine, valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine, or to an uncharged polar amino acid such as serine, threonine, glutamine, asparagine, tyrosine, or cysteine, or to a positively charged polar amino acid such as histidine, lysine, or arginine. In certain embodiments, the mutation at D149 is to Ala, Phe, His, Asn, Gln, Ser, Thr, Val, Trp, Tyr, Lys, or Arg.
[0100] In some embodiments, the mutations at S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156 are to a positively charged polar amino acid, preferably histidine, lysine, or arginine, more preferably lysine or arginine.
[0101] Serratia nuclease A is a polypeptide that has the amino acid sequence of SEQ ID NO: 1 and exhibits nuclease activity. The first 21 amino acids at the N-terminus of SEQ ID NO: 1 are a signal sequence, and the remaining amino acids form the mature polypeptide.
[0102] In some embodiments, a polypeptide disclosed herein is an isolated, synthetic, or recombinant polypeptide comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or more sequence identity to SEQ ID NO:1 or the mature polypeptide thereof, provided that the polypeptide has nuclease activity. The polypeptide may, for example, have 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% sequence identity to SEQ ID NO:1 or the mature polypeptide thereof, or a range between any two of these values.
[0103] In some embodiments, a polypeptide comprises (a) a mutation at one, two, three, four, five, six, seven, or more or all of S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, or D156 of SEQ ID NO:1. The mutations include amino acid modifications, substitutions, or deletions. In some embodiments, a polypeptide comprising an amino acid sequence having at least 70% sequence identity to the sequence of (a) and having mutations at the corresponding sites described in (a) retains nuclease activity. The amino acid mutations are described herein with respect to the corresponding amino acid positions in SEQ ID NO:1. For example, an amino acid substitution from S to K at position 74 of SEQ ID NO:1 is described herein as S74K.
[0104] In some embodiments, the polypeptide has mutations at at least four sites selected from the group consisting of S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156 of SEQ ID NO: 1 or the mature polypeptide thereof. In some embodiments, the polypeptide has a mutation at D149 of SEQ ID NO: 1 and at least three sites selected from the group consisting of S74, N79, A95, T98, N101, A102, S137, D138, Q141, and D156 of SEQ ID NO: 1. In certain embodiments, the mutations at each site are to lysine or arginine.
[0105] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1, or the mature polypeptide thereof, and has mutations at at least three sites selected from the group consisting of N79, A95, A102, and D149, and optionally has a mutation at at least one site selected from the group consisting of S74, T98, N101, S137, D138, Q141, and D156. In certain embodiments, the mutation at each site is to lysine or arginine.
[0106] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1, or the mature polypeptide thereof, and has mutations at A95, A102, and D149, and optionally a mutation at at least one site selected from the group consisting of S74, N79, T98, N101, S137, D138, Q141, and D156. In certain embodiments, the mutation at each site is to lysine or arginine.
[0107] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1, or the mature polypeptide thereof, and has mutations at N79, A102, and D149, and optionally at least one mutation selected from the group consisting of S74, A95, T98, N101, S137, D138, Q141, and D156. In certain embodiments, the mutation at each site is to lysine or arginine.
[0108] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1, or the mature polypeptide thereof, and has mutations at N79, A95, and D149, and optionally at least one mutation selected from the group consisting of S74, T98, N101, A102, S137, D138, Q141, and D156. In certain embodiments, the mutation at each site is to lysine or arginine.
[0109] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1, or the mature polypeptide thereof, and has mutations at N79, A95, and A102, and optionally a mutation at at least one site selected from the group consisting of S74, T98, N101, A102, S137, D138, Q141, D149, and D156. In certain embodiments, the mutation at each site is to lysine or arginine.
[0110] In some embodiments, the polypeptide has mutations at N79, A95, and A102 of SEQ ID NO: 1, and optionally has a mutation at at least one site selected from the group consisting of S74, T98, N101, S137, D138, Q141, and D156 of SEQ ID NO: 1. In some embodiments, the polypeptide has mutations at N79, A95, A102, and D149, and optionally has mutations at one or two sites selected from the group consisting of Q141 and N101. In certain embodiments, the mutation at each site is to lysine or arginine.
[0111] In some embodiments, the polypeptide has mutations at A95, N101, Q141, and D149 of SEQ ID NO: 1, and optionally mutations at at least three sites selected from the group consisting of S74, N79, T98, A102, S137, D138, and D156 of SEQ ID NO: 1. In some embodiments, the polypeptide has mutations at A95, N101, Q141, and D149, and optionally mutations at S74, N79, and T98. In certain embodiments, the mutations at each site are to lysine or arginine.
[0112] In some embodiments, the polypeptide may be a polypeptide comprising an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 1 or a mature polypeptide thereof, provided that the polypeptide has nuclease activity at high solution ionic strength, and the polypeptide is selected from the group consisting of: (1) N79K or N79R; (2) A95K or A95R; (3) (3) A102K or A102R, (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R; for example, one or more mutations of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0113] In some embodiments, the polypeptide has at least four mutations in SEQ ID NO:1 selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R (e.g., the group consisting of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R). In some embodiments, the polypeptide has a D149K or D149R mutation in SEQ ID NO:1 and at least three (or at least four, at least five, at least six, at least seven or more) mutations selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K and (11) D156R (e.g., the group consisting of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K and D156R).
[0114] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has at least three mutations selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, and (4) D149K or D149R, and optionally has mutations selected from the group consisting of S74K, T98K, N101K, S137K, D138K, Q141K, and D156R.
[0115] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has three mutations: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, and optionally a mutation selected from the group consisting of (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R.
[0116] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has three mutations: (1) N79K or N79R, (2) A95K or A95R, (3) D149K or D149R, and optionally a mutation selected from the group consisting of (4) A102K or A102R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R.
[0117] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has three mutations: (1) N79K or N79R, (2) A102K or A102R, (3) D149K or D149R, and optionally a mutation selected from the group consisting of (4) A95K or A95R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R.
[0118] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has three mutations: (1) A95K or A95R, (2) A102K or A102R, (3) D149K or D149R, and optionally a mutation selected from the group consisting of (4) N79K or N79R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R.
[0119] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has three mutations: (1) A95K or A95R, (2) A102K or A102R, (3) D149K or D149R, and optionally a mutation selected from the group consisting of (4) N79K or N79R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R.
[0120] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has three mutations: (i) A95K, A102K, and D149K, or (ii) A95R, A102R, and D149R, and optionally has a mutation selected from the group consisting of: (1) N79K or N79R, (2) S74K, (3) T98K, (4) N101K, (5) S137K, (6) D138K, (7) Q141K, and (8) D156R.
[0121] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has three mutations: (i) N79K, A102K, and D149K, or (ii) N79R, A102R, and D149R, and optionally has a mutation selected from the group consisting of: (1) A95K or A95R, (2) S74K, (3) T98K, (4) N101K, (5) S137K, (6) D138K, (7) Q141K, and (8) D156R.
[0122] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has three mutations: (i) N79K, A95K, and D149K, or (ii) N79R, A95R, and D149R, and optionally has a mutation selected from the group consisting of: (1) A102K or A102R, (2) S74K, (3) T98K, (4) N101K, (5) S137K, (6) D138K, (7) Q141K, and (8) D156R.
[0123] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has three mutations: (i) N79K, A95K, and A102K, or (ii) N79R, A95R, and A102R, and optionally has a mutation selected from the group consisting of: (1) D149K or D149R, (2) S74K, (3) T98K, (4) N101K, (5) S137K, (6) D138K, (7) Q141K, and (8) D156R.
[0124] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO:1 or the mature polypeptide thereof, and has three mutations: (i) N79K, A95K, A102K, and D149K, or (ii) N79R, A95R, A102R, and D149R, and optionally has a mutation selected from the group consisting of: (1) S74K, (2) T98K, (3) N101K, (4) S137K, (5) D138K, (6) Q141K, and (7) D156R.
[0125] In some embodiments, the polypeptide has (i) N79K, A95K, A102K, and D149K mutations, or (ii) N79R, A95R, A102R, and D149R mutations, and optionally has mutations selected from the group consisting of S74K, T98K, N101K, S137K, D138K, Q141K, and D156R. In some embodiments, the polypeptide has (i) N79K, A95K, A102K, and D149K mutations, or (ii) N79R, A95R, A102R, and D149R mutations, and optionally has mutations selected from the group consisting of Q141K and N101K.
[0126] In some embodiments, the polypeptide has the following mutations in SEQ ID NO: 1: A95K, N101K, Q141K, D149K, and optionally has mutations selected from the group consisting of S74K, N79K, T98K, A102K, S137K, D138K, and D156R. In some embodiments, the polypeptide has the following mutations in SEQ ID NO: 1: A95K, N101K, Q141K, D149K, and optionally has mutations S74K, N79K, and T98K. In some embodiments, the polypeptide has the following mutations in SEQ ID NO: 1: S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0127] In some embodiments, the polypeptide comprises a mutation or combination of mutations selected from the group consisting of: (a) N79K, A95K, A102K, and D149K; (b) N79K, A95K, A102K, Q141K, and D149K; (c) N79K, A95K, N101K, A102K, Q141K, and D149K; (d) A95K, N101K, Q141K, and D149K; (e) S74K, N79K, A95K, T98K, N101K, Q141K, and D149K; and (f) S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0128] In some embodiments, the polypeptide is any one of the nuclease variants disclosed herein, provided that the polypeptide has nuclease activity.
[0129] In some embodiments, the nuclease polypeptide is thermostable. For example, the nuclease polypeptide is more thermostable than a nuclease having the sequence of SEQ ID NO: 1. In some embodiments, the nuclease activity of the polypeptide is at least 1%, 2%, 3%, 4%, 5%, 7%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% or more greater than the activity of a nuclease having the sequence of SEQ ID NO: 1 at a given temperature, e.g., between 30°C and 60°C. In some embodiments, the nuclease activity of the polypeptide is at least 1%, 2%, 3%, 4%, 5%, 7%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100%, or a range between two of these numbers, greater than the activity of a nuclease having the sequence of SEQ ID NO: 1 at a given temperature, e.g., between 30° C. and 60° C. In some embodiments, the nuclease activity of the polypeptide is at least 1%, 2%, 3%, 4%, 5%, 7%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% or more greater than the activity of a nuclease having the sequence of SEQ ID NO: 1 at a given pH, e.g., between pH 4 and pH 11.
[0130] In some embodiments, the nuclease polypeptide is salt-tolerant. In some embodiments, the mutant polypeptide has at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the nuclease activity at a solution ionic strength of greater than 200 mM compared to a nuclease having the sequence of SEQ ID NO: 1, where the ion is a monovalent or divalent ion. In some embodiments, the solution ionic strength is greater than 300 mM, e.g., greater than 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1 M, 2 M, 3 M, 4 M, 5 M, or more.
[0131] The optimal ionic strength of a polypeptide can be different (e.g., higher or lower) than the ionic strength of a nuclease having the sequence of SEQ ID NO: 1 or its parent nuclease. For example, the optimal ionic strength of a polypeptide can be 10 mM, 20 mM, 30 mM, 40 mM, 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM, 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM higher than the optimal ionic strength of a nuclease having the sequence of SEQ ID NO: 1, or a range between any two of these values. In some embodiments, the optimal ionic strength of the polypeptide is or is about 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM, 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, or a range between any two of these values. In some embodiments, the optimal temperature of the polypeptide is between 100 mM and 600 mM.
[0132] In fermentation processes, DNA from the production host can complicate many aspects of the protein recovery process. Currently, DNase is added to remove DNA from the final product, often requiring the addition of expensive substances from other sources. In some embodiments, the nucleases disclosed herein can be expressed in the same production host as the product of interest, eliminating the need for external DNase addition and reducing overall processing costs to the final product.
[0133] The nuclease polypeptides disclosed herein may have one or more signal sequences. In some embodiments, at least one of the one or more signal sequences is heterologous to the nuclease polypeptide in which it is contained. In some embodiments, the nuclease polypeptides disclosed herein do not contain a signal sequence.
[0134] Also disclosed herein are antibodies or binding fragments thereof (e.g., isolated or purified antibodies or binding fragments thereof) that specifically bind to an isolated, synthetic, or recombinant polypeptide comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or more sequence identity to SEQ ID NO:1, or the mature polypeptide thereof, provided that the polypeptide has nuclease activity. In some embodiments, the polypeptide comprises one or more of the following mutations: S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0135] In some embodiments, the polypeptide comprises: (1) N79K, A95K, A102K, and D149K; (2) N79K, A95K, A102K, Q141K, and D149K; (3) N79K, A95K, N101K, A102K, Q141K, and D149K; (4) A95K, N101K, Q141K, and D149K; (5) S74K, N79K, A95K, T98K, N101K, Q141K, and D149K; (6) S74K, N79K, A95K, T98K, N101K, A102K, S137K, D149K; (13) N79R, A102R, D149R; (14) N79R, A95R, A102R, D149R; (15) N79R, A95R, A102R, D149R; (16) N79R, A95R, A102R, D149R; (17) N79R, A95R, A102R, D149R; (18) N79R, A95R, D149R; (19) N79R, A95R, D149R; (20) N79R, A95R, D149R; (21) N79R, A95R, D149R; (22) N79R, A95R, D149R; (23) N79R, A102R, D149R; (24) N79R, A95R, D149R; (25) N79R, A95R, D149R.
[0136] Variants of the nucleases disclosed herein may contain one or more substitutions, deletions, and insertions at one or more amino acid positions of the nuclease. In some embodiments, the number of amino acid substitutions, deletions, and / or insertions introduced into a parent nuclease (e.g., a nuclease having the sequence of SEQ ID NO: 1) is 30 or less, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, or 29. Amino acid changes can be minor, i.e., conservative amino acid substitutions or insertions that do not significantly affect protein folding and / or activity; small deletions (e.g., 1-20 amino acids); small amino- or carboxyl-terminal extensions, such as an amino-terminal methionine residue; small linker peptides of up to 20-25 residues; or small extensions that facilitate purification by altering net charge or another function, such as polyhistidine tracts, antigenic epitopes, or binding domains. In some embodiments, amino acid changes to a nuclease can alter one or more physical or chemical properties of the parent nuclease.
[0137] Examples of conservative substitutions are within the group of basic amino acids (arginine, lysine, histidine), acidic amino acids (glutamic acid, aspartic acid), polar amino acids (glutamine, asparagine), hydrophobic amino acids (leucine, isoleucine, valine), aromatic amino acids (phenylalanine, tryptophan, tyrosine), and small amino acids (glycine, alanine, serine, threonine, methionine). Amino acid substitutions that generally do not alter a specific activity are known in the art and are described, for example, in H. Neurath and R.L. Hill, 1979, In, The Proteins, Academic Press, New York. Non-limiting exemplary amino acid substitutions include Ala to Ser, Val to Ile, Asp to Glu, Thr to Ser, Ala to Gly, Ala to Thr, Ser to Asn, Ala to Val, Ser to Gly, Tyr to Phe, Ala to Pro, Lys to Arg, Asp to Asn, Leu to Ile, Leu to Val, Ala to Glu, and Asp to Gly.
[0138] Amino acid substitutions, deletions, and / or insertions can be made and tested using methods known in the art for protein / DNA engineering, including, but not limited to, mutagenesis, recombination, and / or shuffling, followed by associated screening procedures. In some embodiments, mutagenesis / shuffling methods can be combined with high-throughput automated screening methods to detect the activity of cloned mutagenized polypeptides expressed by host cells. Mutagenized DNA molecules encoding active polypeptides can be recovered from host cells and rapidly sequenced using standard methods in the art. These methods allow the rapid determination of the importance of individual amino acid residues in a polypeptide.
[0139] In some embodiments, a nuclease polypeptide can tolerate higher or lower ionic strengths than other nucleases, such as a nuclease having the sequence of SEQ ID NO: 1. For example, a nuclease polypeptide can retain nuclease activity (e.g., at least 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 98% of its nuclease activity) at or about 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM, 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, or a range between any two of these values.
[0140] The nuclease polypeptides disclosed herein may have the same or different substrate specificity compared to a nuclease having the sequence of SEQ ID NO:1 or its respective parent nuclease. For example, a nuclease polypeptide may have substantially the same substrate specificity compared to a nuclease having the sequence of SEQ ID NO:1 or its respective parent nuclease. In some embodiments, a nuclease polypeptide has a substrate specificity of about 50%, 60%, 70%, 80%, 90%, 95%, 98%, or a range between any two of these values compared to a nuclease having the sequence of SEQ ID NO:1 or its respective parent nuclease. Also disclosed herein are compositions or kits comprising one or more polypeptides disclosed herein.
[0141] Also provided herein are immobilized nuclease polypeptides, wherein the immobilized polypeptide comprises one of the nuclease polypeptides disclosed herein. In some embodiments, the polypeptide may be immobilized on a cell, a metal, a resin, a polymer, a ceramic, a glass, a microelectrode, a graphite particle, a bead, a gel, a plate, an array, a capillary tube, or a combination thereof.
[0142] nucleic acid As used herein, the terms "nucleic acid" or "nucleotide" or "polynucleotide" or "nucleic acid sequence" are interchangeable and may be in the form of DNA or RNA. DNA includes cDNA, genomic DNA, or artificially synthesized DNA. DNA may be single-stranded or double-stranded. DNA may be a coding strand or a non-coding strand. As used herein, the term "variant" when referring to a nucleic acid may refer to a naturally occurring allelic variant or a non-naturally occurring variant. These nucleotide variants include degenerate variants, substitution variants, deletion variants, and insertion variants. As known in the art, an allelic variant is a form of nucleic acid substitution, which may involve the substitution, deletion, or insertion of one or more nucleotides, but does not substantially alter the function of the protein it encodes. The nucleic acids of the present invention may comprise a nucleotide sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or 100% sequence identity to a nucleic acid sequence. The present invention also relates to nucleic acid fragments that hybridize to the above sequences. As used herein, the length of a "nucleic acid fragment" is at least 15 nucleotides, preferably at least 30 nucleotides, more preferably at least 50 nucleotides, and even more preferably at least 100 nucleotides or more. Nucleic acid fragments can be used in nucleic acid amplification techniques (e.g., PCR).
[0143] The full-length coding sequence or fragments of the polypeptides of the present invention can be obtained by PCR amplification, artificial synthesis, or recombinant procedures. Mutations in the polypeptides can be introduced during PCR, synthesis, or recombinant procedures.
[0144] For PCR amplification, primers can be designed according to the nucleotide sequences disclosed herein, and the target sequence can be amplified using a commercially available cDNA library or a cDNA library prepared by conventional methods known to those skilled in the art as a template. If the nucleotide sequence is longer than 2500 bp, it is preferable to perform two to six rounds of PCR amplification, followed by splicing the individually amplified fragments together in the correct order. The PCR amplification procedures and systems described herein are not particularly limited, and conventional PCR amplification procedures and systems in the art can be utilized.
[0145] Furthermore, for particularly short fragments, the desired sequence can be artificially synthesized. For example, if the nucleotide sequence of the optical probe is less than 2500 bp, it can be synthesized by artificial synthesis. The artificial synthesis method can be any conventional DNA synthesis method known in the art. Generally, many small fragments are first synthesized and then ligated to obtain a long sequence. It is also possible to obtain a DNA sequence encoding the protein of the present invention entirely by chemical synthesis. This DNA sequence can then be introduced into various existing DNA molecules known in the art, such as vectors, or into cells.
[0146] Recombination can also be used to obtain relevant sequences in bulk, typically by cloning them into a vector, reintroducing them into cells, and then isolating and purifying the desired polypeptide or protein from grown host cells using conventional methods.
[0147] Production of nuclease polypeptides and variants thereof Provided herein are methods for modifying and generating variants of the nuclease polynucleotides disclosed herein. In some embodiments, synthetic or recombinant nucleic acids encoding one or more polypeptides disclosed herein, and vectors (e.g., expression vectors) containing the nucleic acids, are provided. Non-limiting examples of methods include synthetic ligation reassembly, random mutagenesis, targeted mutagenesis, optimized directed evolution systems, and / or saturation mutagenesis, such as gene site saturation mutagenesis (GSSM), and any combination thereof. The term "mutant" refers to a polynucleotide or polypeptide according to the present invention that has been modified in one or more base pairs, codons, introns, exons, or amino acid residues (respectively), yet retains the biological activity of a nuclease. Mutants can be generated by methods such as error-prone PCR, shuffling, site-directed mutagenesis, assembly PCR, sexual PCR mutagenesis, in vivo mutagenesis (phage-assisted continuous evolution, in vivo continuous evolution), cassette mutagenesis, recursive ensemble mutagenesis, exponential ensemble mutagenesis, site-directed mutagenesis, gene reassembly, gene site saturation mutagenesis, synthetic ligation reassembly, recombination, recursive sequence recombination, phosphothioate-modified DNA mutagenesis, uracil-containing template mutagenesis, gapped duplex mutagenesis, point mismatch repair mutagenesis, repair-deficient host strain mutagenesis, chemical mutagenesis, radiation-induced mutagenesis, deletion mutagenesis, restriction-selection mutagenesis, restriction-purification mutagenesis, artificial gene synthesis, ensemble mutagenesis, chimeric nucleic acid multimer generation, and / or combinations of these and other methods.
[0148] Cloning vehicles containing expression cassettes (e.g., vectors) can be utilized herein to express one or more nuclease polypeptides disclosed herein. As used herein, the term "vector" encompasses all types of cloning vehicles, including, but not limited to, plasmids, phagemids, viral vectors (e.g., phages), bacteriophages, baculoviruses, cosmids, fosmids, artificial chromosomes, and other vectors specific to a particular host of interest. Low-copy number and high-copy number vectors are also included. The exogenous polynucleotide sequence typically contains a coding sequence, sometimes referred to herein as a "gene of interest." The gene of interest may contain introns and exons depending on the type of origin or destination in the host cell. The cloning vehicle may be a viral vector, a plasmid, a phage, a phagemid, a cosmid, a fosmid, a bacteriophage, an artificial chromosome, or a combination thereof. The viral vector may include an adenoviral vector, a retroviral vector, or an adeno-associated viral vector. Cloning vehicles may include bacterial artificial chromosomes (BACs), plasmids, bacteriophage P1-derived vectors (PACs), yeast artificial chromosomes (YACs), and mammalian artificial chromosomes (MACs). In some embodiments, the polynucleotide sequence encoding one or more nuclease polypeptides is integrated into the chromosome of the host cell in which the polynucleotide sequence resides, such that the polynucleotide becomes part of the chromosome of the host cell. For example, the host cell may be a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, an insect cell, or a plant cell. In some embodiments, the polynucleotide sequence encoding one or more nuclease polypeptides is not located in the chromosome of the host cell.
[0149] Also provided herein are transformed host cells comprising a nucleic acid or expression cassette (e.g., a vector) or cloning vehicle comprising a nucleic acid sequence encoding one or more nuclease polypeptides disclosed herein. Some embodiments provide methods for producing recombinant polypeptides having nuclease activity, comprising expressing a polynucleotide encoding one or more nuclease polypeptides disclosed herein under conditions that allow for expression of at least one of the one or more nuclease polypeptides, thereby producing a recombinant polypeptide having nuclease activity. In some embodiments, the polynucleotide encoding one or more nuclease polypeptides disclosed herein is operably linked to a promoter. In some embodiments, the polynucleotide is present in an expression vector. In some embodiments, the polynucleotide is present in a host cell in a manner that allows for expression of the polypeptide. In some embodiments, the polynucleotide is present in a chromosome of the host cell in a manner that allows for expression of the polypeptide. In some embodiments, the transformed host cell is a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, an insect cell, or a plant cell. In some embodiments, the transformed host cell is a cell derived from Escherichia coli.
[0150] Non-limiting examples of expression vectors include virus particles, baculoviruses, phages, plasmids, phagemids, cosmids, fosmids, bacterial artificial chromosomes, viral DNA (e.g., vaccinia virus, adenovirus, fowlpox virus, pseudorabies, and SV40 derivatives), Pi-based artificial chromosomes, yeast plasmids, yeast artificial chromosomes, and other vectors specific to particular hosts of interest (e.g., Bacillus, Aspergillus, and yeast). The nuclease-encoding DNA disclosed herein may be included in any of a variety of expression vectors for expressing nuclease polypeptides. Such vectors include chromosomal, non-chromosomal, and synthetic DNA sequences. Many suitable vectors are known to those skilled in the art and are commercially available, for example, pET-28a (Novagen). Depending on the intended use, low- or high-copy-number vectors can be used.
[0151] Codon optimization can be applied to achieve high levels of protein expression in host cells. In some embodiments, codons in a nucleic acid encoding one or more nuclease polypeptides disclosed herein may be optimized to increase or decrease expression in a host cell. For example, one or more non-preferred or less preferred codons in a nucleic acid encoding a nuclease polypeptide can be replaced with one or more "preferred codons" that encode the same amino acid for a host cell of interest. As used herein, a "preferred codon" is a codon that is over-represented in the coding sequence of genes in a host cell, and a "non-preferred or less preferred codon" is a codon that is under-represented in the coding sequence of genes in a host cell. As an illustrative example, a codon-optimized encoding nucleic acid sequence of SEQ ID NO:1 is disclosed herein as SEQ ID NO:2.
[0152] Host cells for expressing the nucleic acids, expression cassettes, and vectors of the present disclosure can be eukaryotic or prokaryotic, including bacteria, yeast, fungi, plant cells, insect cells, and mammalian cells, and methods are provided for optimizing codon usage in all of these cells, codon-modified nucleic acids, and polypeptides produced by the codon-modified nucleic acids. Exemplary host cells include Gram-negative and Gram-positive bacteria. Exemplary host cells also include eukaryotic organisms such as various yeasts, mammalian cells, and insect cells. In some embodiments, the host cell is a cell derived from an organism selected from the group consisting of Pichia pastoris, Bacillus subtilis, Pseudomonas fluorescens, Myceliopthora thermophile fungus, Trichoderma reesei, Escherichia coli, Bacillus licheniformis, Aspergillus niger, Schizosaccharomyces pombe, and Sacaramyces cerevisiae. Nucleic acids encoding the nuclease polypeptides disclosed herein may be located in the genome of the host cell, e.g., may be part of a chromosome of the host cell. In some embodiments, the nucleic acid encoding the nuclease polypeptide is located on an expression vector separate from the genome of the host cell. In some embodiments, methods of producing a nuclease polypeptide disclosed herein may include expressing a nucleic acid encoding a nuclease polypeptide under conditions that allow expression of the nuclease polypeptide, thereby producing the nuclease polypeptide.In some embodiments, the nucleic acid encoding the nuclease polypeptide is operably linked to an inducible promoter, e.g., a promoter that is induced by a change in temperature and / or pH, and / or by the presence, absence, or change in amount / concentration of a compound (e.g., IPTG, arabinose, tetracycline, steroids, and metals).
[0153] The present specification also includes nucleic acids and polypeptides optimized for expression in these organisms and species.
[0154] Compositions and Uses Methods for degrading polynucleotides using one or more nuclease polypeptides disclosed herein are provided. In some embodiments, the methods include contacting one or more polynucleotide molecules with one or more nuclease polypeptides disclosed herein, thereby degrading the polynucleotide molecules. The polynucleotide molecules may comprise DNA (e.g., single-stranded or double-stranded DNA), RNA (e.g., single-stranded or double-stranded DNA), or any combination thereof. The contacting may be performed at various pH values, such as pH 4, pH 5, pH 6, pH 7, pH 8, pH 9, pH 10, pH 11, or a range between any two of these values. The contacting may also be performed at various temperatures, such as 10°C, 15°C, 20°C, 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, or a range between any two of these values. The one or more polynucleotide molecules may be included in a variety of objects, such as a laundry composition (e.g., textiles, clothing, cotton, fabric, or any combination thereof), or an aqueous solution (e.g., a reaction mixture).
[0155] Provided herein are compositions and kits comprising one or more nuclease polypeptides disclosed herein. The compositions may be enzyme compositions, detergent compositions, detergent additives, foods, food supplements, feed supplements, feeds, pharmaceutical compositions, fermentation products, fermentation intermediates, fermentation downstream reaction mixtures, or combinations thereof. In some embodiments, the reaction mixture is for protein expression or purification. The enzyme composition may, for example, comprise one or more nuclease polypeptides disclosed herein and a storage buffer. In some embodiments, the storage buffer comprises a pH buffer system (e.g., Tris-HCl) that provides a pH of about 6.0-9.0, e.g., pH 6.0-7.0 or about pH 6.5. The storage buffer may also include a stabilizer, such as glycerol, at a concentration of at least about 5%, 10%, 20%, 30%, 40%, 50%, or more.
[0156] Also disclosed herein is the use of a polypeptide, nucleic acid, nucleic acid construct, or host cell disclosed herein in the manufacture of a pharmaceutical composition for preventing or treating a disease or disorder, such as wounds, dental plaque, dental caries, periodontitis, native valve endocarditis, chronic bacterial prostatitis, otitis media, infections associated with medical devices such as artificial heart valves, artificial pacemakers, contact lenses, artificial joints, sutures, catheters, and arteriovenous shunts; infections associated with mucosal lesions such as wounds, lacerations, sores, and ulcers; infections of the oral cavity, oropharynx, nasopharynx, and laryngopharynx; infections of the outer ear; eye infections; infections of the stomach, small intestine, and large intestine; infections of the urethra and vagina; skin infections; intranasal infections, such as sinus infections; or combinations thereof. In some embodiments, a method of using the composition comprises contacting the composition with a wound, laceration, sore, mucosal lesion, or any combination thereof. In some embodiments, a method of using the composition comprises contacting the composition with a medical device. In some embodiments, a method of using the composition comprises contacting the composition with the skin, outer ear, eye, or a combination thereof. The composition can be formulated in a variety of forms, such as a tablet, gel, pill, implant, liquid, spray, film, micelle, powder, food, feed pellet, encapsulated form, or a combination thereof.
[0157] In some embodiments, a composition comprising one or more nuclease polypeptides disclosed herein is used to contact and wash a surface containing a DNA substrate.
[0158] In some embodiments, the nuclease polypeptides disclosed herein may be used alone or in combination with one or more additional enzymes in an application or use. Some embodiments provide compositions comprising the nucleases and variants thereof disclosed herein. In some embodiments, the compositions may further comprise one or more additional enzymes, or one or more additional components, or any combination thereof. The one or more additional enzymes may include, but are not limited to, one or more of nucleases, proteases, lipases, cutinases, amylases, carbohydrases, cellulases, pectinases, mannanases, arabinases, galactanases, xylanases, oxidases (e.g., laccases and peroxidases), and deoxyribonucleases (DNases).
[0159] Also provided are pharmaceutically acceptable prodrugs of pharmaceutical compositions and methods of treatment using such pharmaceutically acceptable prodrugs. A "prodrug" refers to a precursor of a specified compound that, after administration to a subject, generates the compound in vivo through a chemical or physiological process, such as solvolysis or enzymatic cleavage, or under physiological conditions (e.g., a prodrug returned to physiological pH is converted to a drug). A "pharmaceutically acceptable prodrug" is a prodrug that is nontoxic, biologically acceptable, and biologically suitable for administration to a subject. Exemplary procedures for the selection and preparation of suitable prodrug derivatives are described, for example, in Bundgaard, Design of Prodrugs (Elsevier Press, 1985).
[0160] Also provided are pharmaceutically active metabolites of the pharmaceutical compositions and the use of such metabolites in the methods herein. "Pharmaceutical active metabolites" refer to pharmacologically active products resulting from the metabolism of a compound or its salt in vivo. Prodrugs and active metabolites of a compound can be measured using routine techniques known or available in the art. See, for example, Bundgaard, Design of Prodrugs (Elsevier Press, 1985).
[0161] Any suitable formulation of the compounds described herein may be prepared. See generally Remington's Pharmaceutical Sciences (2000) Hoover, JE, editor, 20th edition. Formulations are selected to accommodate the appropriate route of administration. Some routes of administration include oral, parenteral, inhalation, topical, rectal, nasal, buccal, vaginal, via an implanted reservoir, or other drug administration methods. If the compound is sufficiently basic or acidic to form stable, non-toxic acid or base salts, administering the compound as a salt may be appropriate. Examples of pharmaceutically acceptable salts include organic acid addition salts formed with acids that form physiologically acceptable anions, such as tosylate, methanesulfonate, acetate, citrate, malonate, tartrate, succinate, benzoate, ascorbate, α-ketoglutarate, and α-glycerophosphate. Suitable inorganic salts may also be formed, including hydrochlorides, sulfates, nitrates, bicarbonates, and carbonates. Pharmaceutically acceptable salts are obtained using standard procedures well known in the art, for example, by combining a sufficiently basic compound, such as an amine, with a suitable acid to yield a physiologically acceptable anion. Alkali metal (sodium, potassium, lithium, etc.) and alkaline earth metal (calcium, etc.) salts of carboxylic acids are also made.
[0162] method Disclosed herein includes methods for degrading polynucleotides, comprising contacting a polynucleotide molecule with one or more polypeptides disclosed herein, thereby degrading the polynucleotide molecule.
[0163] The present disclosure also includes methods for degrading DNA or RNA during cell lysis. In some embodiments, the methods include lysing a host cell of interest and adding a polypeptide disclosed herein under conditions that allow the polypeptide to degrade the DNA or RNA.
[0164] The present disclosure also includes methods for degrading DNA or RNA during viral vector production. In some embodiments, the method includes culturing a host cell containing a viral vector of interest and expressing or adding one or more polypeptides disclosed herein under conditions that allow the one or more polypeptides to degrade DNA or RNA other than the viral vector.
[0165] The present disclosure also includes methods for degrading polynucleotides during expression or production of a protein of interest. In some embodiments, the methods include culturing host cells containing a nucleic acid encoding the protein of interest and expressing one or more nuclease polypeptides disclosed herein under conditions that allow degradation of the polynucleotide by at least one of the one or more nuclease polypeptides. The polynucleotide molecules may comprise DNA, RNA, or any combination thereof. In some embodiments, at least one of the one or more nuclease polypeptides is not expressed by cells that express the protein of interest. In some embodiments, one or more nuclease polypeptides is expressed by cells that do not express the protein of interest.
[0166] In these methods, one or more nuclease polypeptides may be expressed, for example, from one or more expression vectors present in the host cell or from a nucleic acid sequence in the host cell chromosome. Expression of the protein of interest and / or expression of the nuclease polypeptide may be inducible. For example, the coding sequence for the protein of interest, the coding sequence for the nuclease polypeptide, or both, may be operably linked to an inducible promoter. The inducible promoter may be induced, for example, by the presence, absence, and / or change in amount of one or more chemical or biological compounds, changes in pH, temperature, osmolality, ionic strength / concentration, or a combination thereof. Some embodiments provide host cells comprising a nucleic acid encoding a protein of interest and a nucleic acid encoding one or more nuclease polypeptides disclosed herein. [Example]
[0167] Aspects of the above-described embodiments are described in further detail with reference to the following experimental examples. These examples are provided for illustrative purposes and are not intended to be limiting unless otherwise specified. Therefore, the present invention should not be construed as being limited to the following embodiments in any way, but rather should be construed to include all variations that become evident as a result of the teachings provided herein. The methods and reagents used in the examples are conventional in the art unless otherwise specified.
[0168] Materials and Methods Protein Engineering Based on the crystal structure of Serratia marcescens nuclease A (PDB accession number: 1G8T), we performed electrostatic resurfacing by site-specifically introducing positively charged residues onto the protein surface to improve the salt tolerance of NucA. Eleven residues, corresponding to S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156 in wild-type nuclease A, were mutated to positively charged lysine or arginine residues. To investigate the combinatorial effect of positively charged residues at multiple sites, 15 other mutants (B0 to B14) were designed. To investigate the contribution and role of different residues to high-salt tolerance, other mutants (B15 to B90) with single-site saturation mutations were designed. The nucleotide sequences of wild-type NucA, B0, B1, B2, B3, B4, and B5 are shown below: Wild type NucA (SEQ ID NO:1): JPEG2026501926000001.jpg150154 JPEG2026501926000002.jpg125154
[0169] Plasmid preparation, protein expression and purification The genes encoding the nuclease A mutants were synthesized with codon optimization (WuXi Biologics) and cloned into pET-28a(+) (Novagen). The plasmids were transformed into E. coli BL21(DE3), and the overexpressed proteins contained a C-terminal hexa-histidine tag and were purified using Ni-NTA resin (Cytiva). Positive clones were inoculated into 200 ml of Luria-Bertani medium and grown at 37°C until the OD600 reached 0.6–0.8. Protein expression was induced by the addition of 0.4 mM isopropylthio-β-galactoside, and the cells were then transferred to 18°C and grown for an additional 20 hours. The cell culture was harvested by centrifugation at 3500g for 20 min at 4°C, and the cell pellet was resuspended in lysis buffer (20 mM Tris, 300 mM NaCl, 10 mM imidazole, 5% v / v glycerol, pH 8.0). Cell lysis was then performed by sonication, and the cell lysate was centrifuged at 16000g for 20 min at 4°C. The supernatant was loaded onto a Ni-NTA column and washed twice with wash buffer (20 mM Tris, 300 mM NaCl, 40 mM imidazole, 5% v / v glycerol, pH 8.0). The target protein was then eluted with elution buffer (20 mM Tris, 300 mM NaCl, 250 mM imidazole, 5% v / v glycerol, pH 8.0). The eluate was then further purified by gel filtration using HiLoad 16 / 600 Superdex 75 pg (Cytiva), and the protein was dialyzed into storage buffer (20 mM Tris-Cl, pH 8.0, 20 mM NaCl, 2 mM MgCl), concentrated to approximately 1 mg / ml, and supplemented with 50% glycerol for storage.
[0170] Enzyme specificity assay To evaluate the enzyme specificity of wild-type nuclease A and its mutants, 4 U of the enzyme was mixed with 1 μg of different substrates (including CHO cell total RNA, plasmid DNA-pET28(a+), single-stranded DNA, and λ DNA) in buffers containing different salt concentrations. The reactions were incubated at 37°C for 30 minutes and then stopped by adding 10 mM EDTA. The samples were loaded onto 1% or 2% agroin gels and stained with GelRed (Sigma-Aldrich) to analyze the digestion efficiency.
[0171] Enzyme residual activity assay Salmon sperm DNA was dissolved in reaction buffer (50 mM Tris, 2 mM MgCl2, pH 8.0) to a concentration of 1 mg / ml. The enzymatic activity of wild-type nuclease A and its mutants was measured using 2 ng of each enzyme and 8 μl of salmon sperm DNA (1 mg / ml) in different buffers with different salt concentrations and pHs. After incubation at 37°C for 30 minutes, the reaction was stopped by adding 10 mM EDTA, and the A260 absorbance of each reaction was collected using a NanoDrop (Thermo Fisher). Relative enzyme activity was calculated using the following formula: Relative enzyme activity = (A260 サンプル -A260 酵素無し ) / (A260 塩無し -A260 酵素無し ), A260 サンプル represents the A260 absorption of each sample, and A260 酵素無し represents the A260 absorbance of the control reaction without enzyme, and A260 塩無し represents the A260 absorbance of the reaction in reaction buffer (50 mM Tris, 2 mM MgCl2, pH 8.0).
[0172] Example 1 To investigate whether HighSalt NucA is a nonspecific nuclease similar to wild-type Nuclease A, different forms of nucleic acid, including supercoiled plasmid DNA, single-stranded DNA, linear double-stranded DNA (λ phage genomic DNA), and total RNA, were digested with HighSalt NucA and wild-type Nuclease A in buffers containing different NaCl concentrations. With or without further explanation, mutant B0 corresponds to the term "HighSalt NucA" in this and the following examples.
[0173] HighSalt NucA can digest linear double-stranded DNA under high salt concentrations λ phage genomic DNA (New England Biolabs, catalog number N3011S) was incubated with HighSalt NucA and wild-type Nuclease A. The results showed that digestion by wild-type Nuclease A was significantly inhibited in buffers containing more than 200 mM NaCl (Figure 2). However, HighSalt NucA could effectively digest λ DNA in buffers containing 500 mM NaCl (Figure 2).
[0174] HighSalt NucA can digest supercoiled plasmid DNA under high salt concentrations Plasmid pET28a(+) (Novagen) was incubated with HighSalt NucA and wild-type Nuclease A, respectively. The results are similar to those of the λ DNA digestion assay. Digestion by wild-type Nuclease A was significantly inhibited in buffers containing more than 200 mM NaCl, whereas HighSalt NucA can effectively digest plasmid DNA at salt concentrations up to 500 mM (Figure 3).
[0175] HighSalt NucA can digest RNA under high salt conditions Total RNA from CHO cells (WuXi Biologics Co, Ltd) extracted using the RNeasy Plus Mini Kit (QIAGEN, Cat. No. 74134) was incubated with HighSalt NucA and wild-type nuclease A, respectively. The most abundant RNA in total RNA is ribosomal RNA, which contains double-stranded and single-stranded portions in different regions. Digestion results also showed that wild-type nuclease A was significantly inhibited at salt concentrations above 200 mM (Figure 4). HighSalt NucA could effectively digest RNA at salt concentrations up to 500 mM (Figure 4).
[0176] HighSalt NucA can digest single-stranded DNA under high salt concentrations Single-stranded DNA (120-130 nt) was synthesized and incubated with HighSalt NucA and wild-type Nuclease A. The results showed that digestion by wild-type Nuclease A was significantly inhibited in buffers containing more than 200 mM NaCl (Figure 5). However, HighSalt NucA could also effectively digest single-stranded DNA in buffers containing 500 mM NaCl (Figure 5).
[0177] conclusion Wild-type Serratia marcescens nuclease A is a nonspecific nuclease that digests various forms of nucleic acids. Enzyme specificity assays revealed that the engineered HighSalt NucA maintains nonspecific catalytic activity toward various forms of nucleic acids. Furthermore, HighSalt NucA exhibits significantly higher salt tolerance than wild-type nuclease A and remains effective even under high salt concentrations of 500 mM NaCl.
[0178] Example 2 To examine the inhibitory effects of various salts on wild-type nuclease A and HighSalt NucA, salmon sperm DNA was incubated with wild-type nuclease A and HighSalt NucA in buffers containing different concentrations of monovalent salts (NaCl, KCl), divalent salts (MgCl2, MnCl2, (NH4)2SO4), and trivalent salt (Na2HPO4), and residual enzyme activity assays were performed.
[0179] Inhibitory effect of monovalent salts on HighSalt NucA Wild-type nuclease A is sensitive to both NaCl and KCl concentrations, and monovalent salt concentrations above 300 mM dramatically reduce the activity of wild-type nuclease A to below 60% (Figure 6). However, engineered HighSalt NucA maintains more than 60% of its enzymatic activity even at a salt concentration of 500 mM (Figure 6). Thus, engineered HighSalt NucA has higher tolerance to monovalent salts.
[0180] Inhibitory effect of divalent salts on HighSalt NucA Because divalent salts provide stronger ionic strength in solution, we also investigated the effects of divalent ions on Nuclease A and HighSalt NucA. Similar to the performance of HighSalt NucA in monovalent salt buffers, the results showed that HighSalt NucA maintained 60% of its enzymatic activity at a MgCl concentration of 400 mM. HighSalt NucA is more tolerant of MgCl than wild-type Nuclease A (Figure 7). Furthermore, HighSalt NucA is somewhat more tolerant of (NH)SO than wild-type Nuclease A (Figure 8). However, the inhibitory effect of MnCl appears to be greater for HighSalt NucA than for wild-type Nuclease A (Figure 7).
[0181] Inhibitory effect of trivalent salts on HighSalt NucA Due to the chelating effect, metal-dependent nucleases react with PO4 3- Both HighSalt NucA and Nuclease A are sensitive to PO4 3-Although they perform similarly in buffer, the enzymatic activity of both enzymes is higher at PO4 >100 mM. 3- is significantly inhibited by (Figure 9).
[0182] conclusion Residual enzyme activity assays showed that HighSalt NucA was more tolerant to high concentrations of NaCl, KCl, and (NH4)2SO4 than wild-type nuclease A. Engineering of nuclease A by introducing positively charged residues on specific surfaces indeed improved its high-salt tolerance.
[0183] Example 3 To investigate the contribution of the single residue mutations at N79, A95, A102, and D149 to salt tolerance, we mutated each residue to one of 19 other amino acids and analyzed the activity of each mutant under different salt concentrations.
[0184] Mutations in N79 and their contribution to salt tolerance Residue N79 was mutated to one of 19 other amino acids. The mutants were purified and their salt tolerance was analyzed. In addition to mutations at the positively charged residues K and R, many other mutations, including A, C, G, Q, S, T, V, W, and Y, showed better salt tolerance than the wild-type enzyme (Figure 10 and Table 1). [Table 1]
[0185] Mutations in A95 and their contribution to salt tolerance Residue A95 was mutated to each of the other 19 amino acids. The mutants were purified and their salt tolerance was analyzed. In addition to mutations of the positively charged residues K and R, other mutations, such as F, N, P, Q, S, T, V, and Y, showed positive effects on improving the salt tolerance of the mutants (Figure 11 and Table 2). [Table 2]
[0186] Mutations in A102 and their contribution to salt tolerance Residue A102 was mutated to one of 19 other amino acids. The mutants were purified and their salt tolerance was analyzed. Among these, mutants containing K, R, T, V, W, and Y had better salt tolerance than the wild-type enzyme (Figure 12 and Table 3). [Table 3]
[0187] Mutations in D149 and their contribution to salt tolerance Residue D149 was mutated to each of the other 19 amino acids. The mutants were purified and their salt tolerance was analyzed. Mutation to most of the other residues improved the salt tolerance of the enzyme by reducing the acidic surface at this site, but promising residues at this site could also be replaced with A, F, H, K, N, Q, R, S, T, V, W, and Y (Figure 13 and Table 4). [Table 4]
[0188] Combinatorial effect of positively charged residues on salt tolerance Based on these results, single-residue mutations, especially those at K and R, improve the enzyme's salt tolerance, but the effect is limited. Therefore, we hypothesized that salt tolerance could be improved by combining mutations. N79, A95, A102, and D149 were selected, and combined mutations of these residues resulted in significantly improved salt tolerance compared with the wild-type enzyme and enzymes with single mutations (Figure 14 and Table 5). Furthermore, these results support another hypothesis: that the improvement of enzyme activity in high-salt buffers depends on the balance between binding and substrate release. [Table 5]
[0189] conclusion In addition to mutations at the positively charged residues K and R, other residues also contribute to the improvement of salt tolerance. However, because single-residue mutations are insufficient for optimal performance, multiple mutations, especially multiple K / R mutations, significantly improve enzyme activity under high-salt conditions.
Claims
1. A polypeptide derived from Serratia marcescens nuclease A and having high salt tolerance. However, the polypeptide contains one or more mutations such that there is more positively charged surface area in the three-dimensional structure of the polypeptide. Preferably, the polypeptide has at least 60% of the nuclease activity at a solution ionic strength of greater than 200 mM compared to a nuclease having the sequence of SEQ ID NO:
1.
2. (a) an amino acid sequence having at least 70% identity to SEQ ID NO: 1, or the mature polypeptide thereof, and having mutations at one, two, three, four, five, six, seven, or more or all of S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, D156; or (b) A polypeptide comprising an amino acid sequence having at least 70% sequence identity to the sequence of (a), having a mutation at the site described in (a), and retaining nuclease activity. Preferably, the mutations at N79, A95, A102, and D149 are each independently to a nonpolar amino acid such as alanine, valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine, or to an uncharged polar amino acid such as serine, threonine, glutamine, tyrosine, or cysteine, or to a positively charged polar amino acid such as histidine, lysine, or arginine. More preferably, the mutations at S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156 are to a positively charged amino acid residue, preferably to histidine, lysine, or arginine.
3. The polypeptide of claim 2, comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or its mature polypeptide, and (1) having a mutation at A95 and having mutations at at least three sites selected from the group consisting of S74, N79, T98, N101, A102, S137, D138, Q141, D149, and D156; or (2) having a mutation at D149 and having mutations at at least three sites selected from the group consisting of S74, N79, A95, T98, N101, A102, S137, D138, Q141, and D156. Preferably, the polypeptide has mutations at A95 and D149, and at least two sites selected from the group consisting of S74, N79, T98, N101, A102, S137, D138, Q141, and D156. More preferably, the polypeptide has mutations at N79, A95, A102 and D149, and optionally has a mutation at at least one site selected from the group consisting of S74, T98, N101, S137, D138, Q141 and D156. More preferably, the polypeptide has mutations at N79, A95, A102 and D149, and optionally has mutations at one or two sites selected from the group consisting of Q141 and N101. More preferably, the polypeptide has mutations at A95, N101, Q141 and D149, and optionally mutations at at least three sites selected from the group consisting of S74, N79, T98, A102, S137, D138 and D156. More preferably, the polypeptide has mutations at A95, N101, Q141, D149, and optionally mutations at S74, N79 and T98.
4. 3. The polypeptide of claim 2, wherein the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof, and has mutations at at least three sites selected from the group consisting of N79, A95, A102, and D149, and optionally has a mutation at at least one site selected from the group consisting of S74, T98, N101, S137, D138, Q141, and D156. Preferably, the polypeptide has mutations at A95, A102, and D149, and optionally has mutations at at least one site selected from the group consisting of S74, N79, T98, N101, S137, D138, Q141, and D156; or the polypeptide has mutations at N79, A102, and D149, and optionally has mutations at at least one site selected from the group consisting of S74, A95, T98, N101, S137, D138, Q141, and D156; or Alternatively, the polypeptide has mutations at N79, A95, and D149, and optionally has a mutation at at least one site selected from the group consisting of S74, T98, N101, A102, S137, D138, Q141, and D156; or, the polypeptide has mutations at N79, A95, and A102, and optionally has a mutation at at least one site selected from the group consisting of S74, T98, N101, A102, S137, D138, Q141, D149, and D156.
5. 3. The polypeptide of claim 2, wherein the polypeptide comprises an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or a mature polypeptide thereof, wherein the polypeptide has nuclease activity at high solution ionic strength, and the polypeptide comprises one or more mutations of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R. Preferably, the polypeptide contains (a) an A95K or A95R mutation and at least three selected from the group consisting of (1) N79K or N79R, (2) A102K or A102R, (3) D149K or D149R, (4) S74K, (5) T98K, (6) N101K, (7) S137K, (8) D138K, (9) Q141K, and (10) D156R. or (b) a D149K or D149R mutation and at least three mutations selected from the group consisting of (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) S74K, (5) T98K, (6) N101K, (7) S137K, (8) D138K, (9) Q141K, and (10) D156R. More preferably, the polypeptide has at least three mutations selected from the group consisting of (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, and (4) D149K or D149R, and optionally has a mutation selected from the group consisting of S74K, T98K, N101K, S137K, D138K, Q141K, and D156R. More preferably, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and has the mutations A95K, N101K, Q141K, D149K, and optionally has a mutation selected from the group consisting of S74K, N79K, T98K, A102K, S137K, D138K, and D156R. More preferably, the polypeptide has the mutations A95K, N101K, Q141K, D149K, and optionally the mutations S74K, N79K, and T98K.
4. A composition or kit comprising one or more polypeptides according to any one of claims 1 to 3. Preferably, the composition is a reaction mixture, a detergent composition, a detergent additive, a food, a food supplement, a feed supplement, a feed, a pharmaceutical composition, a fermentation product, a fermentation intermediate, a fermentation downstream reaction mixture, or a combination thereof.
5. (1) a sequence encoding one or more polypeptides according to any one of claims 1 to 3, or a complementary sequence thereof; or (2) a sequence having at least 50%, 60%, 70%, 80%, or 90% identity to (1); A nucleic acid comprising:
6. A nucleic acid construct comprising the polynucleotide sequence of the nucleic acid of claim 5. Preferably, the nucleic acid construct is a cloning vector, an expression vector, or a recombinant vector. More preferably, the expression vector comprises a viral vector, a plasmid, a phage, a phagemid, a cosmid, a fosmid, a bacteriophage, an artificial chromosome, or a combination thereof.
7. A reaction mixture comprising: (a) one or more polypeptides according to any one of claims 1 to 3; (b) one or more nucleic acid molecules; and (c) an aqueous solution in which the polypeptides hydrolyze the one or more nucleic acid molecules.
8. A cell comprising one or more polypeptides according to any one of claims 1 to 3, one or more nucleic acids encoding said polypeptides, one or more nucleic acid constructs comprising the polynucleotide sequences of said nucleic acids, or a combination thereof.
9. A method for producing a polypeptide having nuclease activity at high solution ionic strength, comprising expressing a nucleic acid encoding one of the polypeptides described in any one of claims 1 to 3 under conditions that allow expression of the polypeptide, thereby producing a recombinant polypeptide having nuclease activity, wherein the nucleic acid is operably linked to a promoter. Preferably, the method comprises the steps of: culturing the cells of claim 7 under conditions that allow the cells to express the polypeptide.
10. A method for degrading polynucleotides, comprising contacting a polynucleotide molecule with one or more polypeptides according to any one of claims 1 to 3, thereby degrading the polynucleotide molecule. Preferably, the contacting is carried out at a pH of 4 to 11. Preferably, the contacting is carried out at a temperature of from 10°C to 70°C.
11. 10. A method for degrading DNA or RNA during cell lysis, comprising lysing a host cell of interest and contacting the lysate with one or more polypeptides of any one of claims 1 to 3 under conditions that allow degradation of the DNA or RNA by the polypeptides. Preferably, the host cell is selected from the group consisting of a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, a plant cell, or an insect cell.
12. A method for degrading DNA or RNA during viral vector production, comprising culturing host cells containing the viral vector of interest, and expressing or adding one or more polypeptides according to any one of claims 1 to 3 under conditions that allow the one or more polypeptides to degrade DNA or RNA other than the viral vector. Preferably, the host cell is selected from the group consisting of a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, a plant cell, or an insect cell.
13. A method for degrading DNA or RNA during protein production, comprising culturing a host cell containing a nucleic acid encoding a protein of interest, and expressing or adding one or more polypeptides of the present disclosure under conditions that allow for degradation of the DNA or RNA by the one or more polypeptides. Preferably, the host cell is selected from the group consisting of a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, a plant cell, or an insect cell.
14. A method for degrading DNA or RNA in a protein production mixture, comprising culturing a host cell containing nucleic acid encoding a protein of interest and expressing one or more polypeptides according to any one of claims 1 to 3 under conditions that allow the polypeptides to degrade the DNA or RNA. Preferably, the host cell is selected from the group consisting of a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, a plant cell, or an insect cell.