Engineered nuclease with high salt tolerance
By introducing specific amino acid mutations into Serratia marcescens nuclease A, its enzyme activity in high-salt environments was increased, solving the problem of enzyme activity inhibition under high-salt conditions and achieving effective application under high-salt conditions.
Patent Information
- Application Number
- CN202380094583.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-20
- Filing Date
- 2023-12-20
- Publication Date
- 2025-10-10
Smart Images

Figure BDA0005557794380000391 
Figure BDA0005557794380000401 
Figure BDA0005557794380000402
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of protein engineering and recombinant expression technology, and more particularly to modifying the biochemical properties of a nuclease derived from Serratia marcescens by combining amino acid residue mutations in the primary sequence of the protein to enhance the enzyme activity of the mutant under high salt concentration. Background Art
[0002] Nucleases are enzymes that cleave phosphodiester bonds in nucleic acid molecules by endo- or exo-cleavage. Serratia marcescens nuclease A (Uniprot No.: P13717) has been shown to be a non-specific nuclease that catalyzes the hydrolysis of single-stranded / double-stranded DNA and RNA by cleaving phosphodiester bonds. Crystal structures show that this enzyme is a magnesium ion (Mg)-dependent 2+ ) is a dimeric nuclease whose activity depends on a key water cluster. It has been widely used for nucleic acid removal in protein production and pharmaceutical virus production. It is commercialized by EMD Millipore under the name of .
[0003] Despite widespread use in industry and research, Serratia marcescens nuclease A suffers from a significant drawback: its susceptibility to inhibition by high salt concentrations. This unfavorable high-salt nature of the enzyme limits its applications and necessitates additional buffer exchange steps in the production of certain pharmaceutical products. Enzyme activity is significantly inhibited when the ionic strength of the solution exceeds 200 mM. However, recent studies have shown that high-salt buffers are more beneficial for virus production. Therefore, a high-salt-tolerant Serratia marcescens nuclease A would be more practical for many applications under high-salt conditions. Summary of the Invention
[0004] Disclosed herein are polypeptides having nuclease activity (hereinafter referred to as "nucleases" or "nuclease polypeptides"), polynucleotides comprising sequences encoding these polypeptides, and methods for preparing and using these polypeptides and polynucleotides. Also provided herein are compositions and kits comprising one or more nuclease polypeptides disclosed herein, one or more nuclease-encoding polynucleotides, and any combination thereof.
[0005] Disclosed herein is a synthetic or recombinant polypeptide derived from Serratia marcescens nuclease A having high salt tolerance. In some embodiments, the polypeptide comprises one or more mutations such that the polypeptide has more positively charged surface area in its three-dimensional structure.
[0006] In some embodiments, the polypeptide has at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of the nuclease activity compared to the nuclease having the sequence of SEQ ID NO: 1 or the mature polypeptide thereof, in a solution having an ionic strength greater than 200 mM. In some embodiments, the ion is a monovalent ion or a divalent ion. In some embodiments, the solution ionic strength is greater than 300 mM. In some embodiments, the solution ionic strength is greater than 400 mM. In some embodiments, the solution ionic strength is greater than 500 mM.
[0007] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 1 or its mature polypeptide. In some embodiments, the polypeptide has nuclease activity, wherein the polypeptide comprises a mutation at one, two, three, four, five, six, seven, or more of S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, or D156, or mutations at all of the positions. Mutations include amino acid modifications, substitutions, or deletions.
[0008] In some embodiments, the polypeptide:
[0009] (a) comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or the mature polypeptide thereof, and having a mutation at one, two, three, four, five, six, seven or more positions, or all positions, of S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, or D156; or
[0010] (b) comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% sequence identity to the sequence described in (a), and has a mutation at the position described in (a) and retains nuclease activity.
[0011] In some embodiments, N79 is mutated to a nonpolar amino acid, such as alanine, valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine; or to an uncharged polar amino acid, such as serine, threonine, glutamine, tyrosine, or cysteine; or to a positively charged polar amino acid, such as histidine, lysine, or arginine.
[0012] In some embodiments, A95 is mutated to a non-polar amino acid, such as valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine; or to an uncharged polar amino acid, such as serine, threonine, glutamine, asparagine, tyrosine, or cysteine; or to a positively charged polar amino acid, such as histidine, lysine, or arginine.
[0013] In some embodiments, A102 is mutated to a non-polar amino acid, such as valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine; or to an uncharged polar amino acid, such as serine, threonine, glutamine, asparagine, tyrosine, or cysteine; or to a positively charged polar amino acid, such as histidine, lysine, or arginine.
[0014] In some embodiments, D149 is mutated to a non-polar amino acid, such as alanine, valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine; or to an uncharged polar amino acid, such as serine, threonine, glutamine, asparagine, tyrosine, or cysteine; or to a positively charged polar amino acid, such as histidine, lysine, or arginine.
[0015] In some embodiments, the mutation at the S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, D156 positions is a mutation to a positively charged polar amino acid, preferably histidine, lysine, or arginine, more preferably lysine or arginine.
[0016] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof, and has a mutation at at least 4 positions selected from S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, D156.
[0017] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof, and has a mutation at the A95 or D149 position, and has a mutation at at least 3 positions selected from S74, N79, A95, T98, N101, A102, S137, D138, Q141, D156. Preferably, the polypeptide has a mutation at the A95 and D149 positions, and has a mutation at at least 2 positions selected from S74, N79, T98, N101, A102, S137, D138, Q141, D156.
[0018] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or its mature polypeptide, and has mutations in at least three of the positions selected from N79, A95, A102 and D149, and optionally has a mutation in at least one of the positions selected from S74, T98, N101, S137, D138, Q141, D156.
[0019] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or its mature polypeptide, and (1) has a mutation at positions A95, A102 and D149, and optionally has a mutation at at least one position selected from the group consisting of S74, N79, T98, N101, S137, D138, Q141, and D156; (2) has a mutation at positions N79, A102 and D149, and optionally has a mutation at a position selected from the group consisting of S74, A95, T98, N101, S137, D138, Q141, and D156. at least one site has a mutation; (3) has mutations at N79, A95 and D149, and optionally has a mutation at at least one site selected from S74, T98, N101, A102, S137, D138, Q141, D156; or (4) has mutations at N79, A95 and A102, and optionally has a mutation at at least one site selected from S74, T98, N101, A102, S137, D138, Q141, D149, D156.
[0020] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or its mature polypeptide, and has mutations at N79, A95, A102, and D149, and optionally has mutations at at least one of the positions selected from S74, T98, N101, S137, D138, Q141, and D156. In some embodiments, the polypeptide has mutations at N79, A95, A102, and D149, and optionally has mutations at one or two of the positions selected from Q141 and N101.
[0021] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or its mature polypeptide, and has mutations at A95, N101, Q141, D149, and optionally at least three of the positions selected from S74, N79, T98, A102, S137, D138, and D156. In some embodiments, the polypeptide has mutations at A95, N101, Q141, and D149, and optionally at S74, N79, and T98.
[0022] In some embodiments, the polypeptide may be a polypeptide comprising NO:1 or its mature polypeptide having an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 95% or at least 99% sequence identity, wherein the polypeptide has nuclease activity under high solution ionic strength, and the polypeptide comprises one or more mutations of (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R; preferably one or more mutations of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K and D156R.
[0023] In some embodiments, the polypeptide:
[0024] (i) comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and having a mutation selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R, preferably having a mutation selected from the group consisting of: S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, D156R; or
[0025] (ii) comprising an amino acid sequence having at least 70%, 80%, 85%, 90%, 95% or 99% sequence identity to the sequence described in (i), and having the mutation described in (i) and retaining nuclease activity.
[0026] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and has at least four mutations selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R, preferably a mutation selected from the group consisting of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0027] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and has: (a) an A95K or A95R mutation, and at least three (or at least four, at least five, at least six, at least seven or more) mutations selected from the group consisting of: (1) N79K or N79R, (2) A102K or A102R, (3) D149K or D149R, (4) S74K, (5) T98K, (6) N101K, (7) S137K, (8) D138K, (9) Q141K, and (10) D156R, preferably selected from the group consisting of S74K, N79K, T98K, N101K, A102K, S137K, D138K, Q141K , D149K and D156R mutations; or (b) D149K or D149R mutations, and at least 3 (or at least 4, at least 5, at least 6, at least 7 or more) mutations selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) S74K, (5) T98K, (6) N101K, (7) S137K, (8) D138K, (9) Q141K, and (10) D156R, preferably mutations selected from the group consisting of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K and D156R. More preferably, the polypeptide has (i) an A95K or A95R mutation and (ii) a D149K or D149R mutation, and at least 2 (or at least 3, at least 4, at least 5, at least 6 or more) mutations selected from the group consisting of: (1) N79K or N79R, (2) A102K or A102R, (3) S74K, (4) T98K, (5) N101K, (6) S137K, (7) D138K, (8) Q141K, and (9) D156R, preferably a mutation selected from the group consisting of S74K, N79K, T98K, N101K, A102K, S137K, D138K, Q141K and D156R.
[0028] In some embodiments, the polypeptide comprises an amino acid sequence that is at least 70% identical to SEQ ID NO: 1 or the mature polypeptide thereof, and has at least three mutations selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, and optionally has a mutation selected from the group consisting of S74K, T98K, N101K, S137K, D138K, Q141K, and D156R.
[0029] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or its mature polypeptide, and has: (i) N79K, A95K, A102K, D149K mutations, or (ii) N79R, A95R, A102R and D149R mutations, and optionally has a mutation selected from S74K, T98K, N101K, S137K, D138K, Q141K and D156R. In some embodiments, the polypeptide has (i) N79K, A95K, A102K and D149K mutations, or (ii) N79R, A95R, A102R and D149R mutations, and optionally has a mutation selected from Q141K and N101K.
[0030] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or its mature polypeptide, and having A95K, N101K, Q141K, D149K mutations, and optionally having a mutation selected from S74K, N79K, T98K, A102K, S137K, D138K and D156R. In some embodiments, the polypeptide has A95K, N101K, Q141K, D149K mutations, and optionally having S74K, N79K and T98K mutations.
[0031] In some embodiments, the polypeptide comprises an amino acid sequence that is at least 70% identical to SEQ ID NO: 1 or the mature polypeptide thereof, and has S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K and D156R mutations.
[0032] In some embodiments, the polypeptide has an optimum temperature between 30°C and 60°C. In some embodiments, the polypeptide has an optimum pH between pH 4 and pH 11. In some embodiments, the polypeptide does not comprise a signal sequence. In some embodiments, the polypeptide further comprises a signal sequence. In some embodiments, the signal sequence is a heterologous sequence or a native signal sequence.
[0033] Also disclosed herein are compositions or kits comprising one or more polypeptides described herein. In some embodiments, the composition is a reaction mixture, a detergent composition, a detergent additive, a food, a food supplement, a feed supplement, a feed, a pharmaceutical composition, a fermentation product, a fermentation intermediate, a fermentation downstream reaction mixture, or a combination thereof. In some embodiments, a reaction mixture comprising one or more of the polypeptides is used for expression or purification of a protein or viral vector. For example, the polypeptide can be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or its mature polypeptide, wherein the polypeptide has nuclease activity under high solution ionic strength, and the polypeptide comprises one or more mutations in S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0034] Also disclosed herein is a nucleic acid comprising: (1) a sequence encoding any one of the polypeptides described herein or its complementary sequence, or (2) a sequence having at least 50%, 60%, 70%, 80% or 90% identity to (1).
[0035] Also disclosed herein are nucleic acid constructs comprising the nucleic acid sequences described herein. In one or more embodiments, the nucleic acid construct is a cloning vector, an expression vector, or a recombinant vector. In one or more embodiments, the nucleic acid sequence is operably linked to an expression control sequence. In some embodiments, the expression vector comprises a viral vector, a plasmid, a phage, a phagemid, a cosmid, a fosmid, a phage, an artificial chromosome, or a combination thereof.
[0036] Disclosed herein is a recombinant cell comprising one or more polypeptides disclosed herein, one or more nucleic acids encoding any of the polypeptides disclosed herein, one or more nucleic acid constructs comprising the nucleic acid polynucleotide sequences, or a combination thereof. In some embodiments, the nucleic acid is part of a chromosome of a recombinant cell. In some embodiments, the recombinant cell is a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, an insect cell, or a plant cell. For example, the polypeptide can be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or its mature polypeptide, wherein the polypeptide has nuclease activity under high solution ionic strength, and wherein the polypeptide comprises one or more mutations in S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0037] Also disclosed herein is a method for producing a polypeptide having nuclease activity under high solution ionic strength. In some embodiments, the method comprises: expressing a nucleic acid encoding any of the polypeptides disclosed herein under conditions that permit expression of the polypeptide, thereby producing a recombinant polypeptide having nuclease activity, wherein the nucleic acid is operably linked to a promoter. In some embodiments, the nucleic acid is present in an expression vector. In some embodiments, the nucleic acid is present in a host cell to permit expression of the polypeptide. In some embodiments, the nucleic acid is present in a chromosome of the host cell. In some embodiments, the host cell is a cell of an organism selected from the group consisting of Pichia pastoris, Bacillus subtilis, Pseudomonas fluorescens, Myceliopthora thermophile fungus, Tricodermea reesei, Escherichia coli, Bacillus licheniformis, Aspergillus niger, Schizosaccharomyces pombe, and Saccharamyces cerevisiae. In some embodiments, the nucleic acid is expressed via an in vitro expression system. The polypeptide may be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or a mature polypeptide thereof, wherein the polypeptide has nuclease activity under high solution ionic strength, and wherein the polypeptide comprises one or more mutations of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K and D156R.
[0038] Also disclosed herein is a method for degrading a polynucleotide, comprising contacting a polynucleotide molecule with one or more polypeptides disclosed herein, thereby degrading the polynucleotide molecule. For example, the polypeptide can be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or its mature polypeptide, wherein the polypeptide has nuclease activity under high solution ionic strength, and wherein the polypeptide comprises one or more mutations in S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K and D156R. In some embodiments, the polynucleotide molecule is a DNA molecule or an RNA molecule. In some embodiments, the contacting is carried out under conditions of pH 4 to pH 11. In some embodiments, the temperature of the reaction mixture is from about 10°C to about 70°C. In some embodiments, the contacting is carried out under conditions of 30°C to 60°C.
[0039] Also disclosed herein is a method for degrading DNA or RNA during cell lysis. In some embodiments, the method comprises: lysing a desired host cell; and adding a polypeptide disclosed herein under conditions that allow the polypeptide to degrade DNA or RNA. The host cell can be, for example, a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, a plant cell, or an insect cell.
[0040] Also disclosed herein is a method for degrading DNA or RNA during viral vector production. In some embodiments, the method comprises: culturing a host cell, wherein the host cell comprises a viral vector of interest; and expressing or adding one or more polypeptides disclosed herein under conditions that allow one or more polypeptides to degrade DNA or RNA other than the viral vector. In some embodiments, the host cell is a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, a plant cell, or an insect cell. In some embodiments, the polypeptide is expressed by an expression vector present in the host cell, or the polypeptide is encoded by a nucleic acid sequence in the host cell chromosome. In some embodiments, the virus is an adeno-associated virus (AAV) or a lentivirus.
[0041] In some embodiments, the polypeptide can be expressed by cells that do not express the target viral vector. In some embodiments, the expression of the target viral vector and the polypeptide can be inducible or non-inducible. In some embodiments, one or more polypeptides disclosed herein can be added externally. For example, the polypeptide can be a polypeptide comprising an amino acid sequence having at least 70% sequence identity with SEQ ID NO: 1 or its mature polypeptide, wherein the polypeptide has nuclease activity under high solution ionic strength and comprises one or more mutations in S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K and D156R.
[0042] Also disclosed herein is a method for degrading DNA or RNA during protein production. In some embodiments, the method comprises: culturing a host cell containing a nucleic acid encoding a protein of interest; and expressing or adding one or more polypeptides disclosed herein under conditions that allow for degradation of the DNA or RNA. In some embodiments, the host cell is a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, a plant cell, or an insect cell. In some embodiments, the polypeptide is expressed by an expression vector in the host cell or is encoded by a nucleic acid sequence in the host cell chromosome.
[0043] In some embodiments, the polypeptide can be expressed by cells that do not express the protein of interest. In some embodiments, the expression of one or more of the protein of interest and the polypeptide can be inducible or non-inducible. In some embodiments, one or more polypeptides disclosed herein can be added externally. For example, the polypeptide can be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or its mature polypeptide, wherein the polypeptide has nuclease activity under high solution ionic strength and comprises one or more mutations in S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K and D156R.
[0044] Also disclosed herein is a reaction mixture comprising: (a) one or more polypeptides disclosed herein, (b) one or more nucleic acid molecules, and (c) an aqueous solution that allows the polypeptide to hydrolyze the one or more nucleic acid molecules. In some embodiments, the nucleic acid molecules include single-stranded DNA, double-stranded DNA, single-stranded RNA, double-stranded RNA, or any combination thereof. In some embodiments, the nucleic acid molecules are from host cells used for protein production. In some embodiments, the polypeptides are expressed in host cells selected from bacterial cells, mammalian cells, fungal cells, yeast cells, and insect cells. In some embodiments, the temperature of the reaction mixture is about 10°C to 70°C, preferably 30°C to 60°C; the pH is about 4 to 11, preferably 7 to 9. In some embodiments, the aqueous solution is a detergent composition, a detergent additive, a food, a food supplement, a feed supplement, a feed, a pharmaceutical composition, a fermentation product, a fermentation intermediate, a fermentation downstream reaction mixture, a product of a protein production process, an intermediate of a protein production process, a protein purification solution, or a combination thereof. For example, the polypeptide can be a polypeptide comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or a mature polypeptide thereof, wherein the polypeptide has nuclease activity under high solution ionic strength and comprises one or more mutations of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K and D156R.
[0045] Also disclosed herein is a method for degrading DNA or RNA in a protein production mixture. In some embodiments, the method comprises: culturing a host cell comprising a nucleic acid encoding a target protein; and expressing one or more of the polypeptides disclosed herein under conditions that allow the polypeptides to degrade DNA or RNA. In some embodiments, expression of the polypeptides may be delayed until after the target protein is produced. In one or more embodiments, expression of the polypeptides is not delayed. Expression of the polypeptides may be initiated before, after, or simultaneously with the start of expression of the target protein. The host cell may be a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, a plant cell, or an insect cell. In some embodiments, the polypeptides are expressed by an expression vector in the host cell, or are encoded by a nucleic acid sequence in the host cell chromosome. In some embodiments, the polypeptides are expressed by cells that do not express the target protein. Expression of the target protein and the polypeptides may be inducible or non-inducible. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1Electrostatic potential surface plots of wild-type NucA from Serratia marcescens (left) and an engineered NucA with more positively charged residues (HighSalt NucA, right). White represents uncharged surface, and dark shaded areas represent charged surface. HighSalt NucA has a greater positively charged surface area than wild-type NucA.
[0047] Figure 2 Digestion of λ-DNA by wild-type NucA and engineered NucA (HighSalt NucA) from Serratia marcescens at different NaCl concentrations. Wild-type NucA activity was significantly inhibited when the salt concentration increased to 300 mM. HighSalt NucA still efficiently digested λ-DNA at 500 mM NaCl.
[0048] Figure 3 Digestion of plasmid DNA by wild-type NucA and engineered NucA (HighSalt NucA) from Serratia marcescens at different NaCl concentrations. Wild-type NucA's enzymatic activity was significantly inhibited at 300 mM NaCl. HighSalt NucA still efficiently digested plasmid DNA at 500 mM NaCl.
[0049] Figure 4 Digestion of total RNA extracted from CHO cells by wild-type NucA and engineered NucA (HighSalt NucA) from Serratia marcescens at different NaCl concentrations. Wild-type NucA's enzymatic activity was significantly inhibited at 300 mM NaCl, while HighSalt NucA still efficiently digested RNA at 500 mM NaCl.
[0050] Figure 5 Digestion of single-stranded DNA (ssDNA) by wild-type NucA and engineered NucA (HighSalt NucA) from Serratia marcescens at different NaCl concentrations. Wild-type NucA activity was significantly inhibited at 300 mM NaCl. HighSalt NucA still efficiently digested ssDNA at 500 mM NaCl.
[0051] Figure 6 .Inhibitory effects of monovalent salts (left: NaCl; right: KCl) on HighSalt NucA and wild-type nuclease A.
[0052] Figure 7 .Inhibitory effect of divalent salts (left: MgCl2; right: MnCl2) on HighSalt NucA and wild-type nuclease A.
[0053] Figure 8 .Inhibitory effect of (NH4)2SO4 on HighSalt NucA and wild-type nuclease A.
[0054] Figure 9 .Inhibitory effect of Na2HPO4 on HighSalt NucA and wild-type nuclease A.
[0055] Figure 10 Comparison of salt tolerance among wild-type NucA, HighSalt NucA, and 19 mutants with single mutation at N79 residue.
[0056] Figure 11 Comparison of salt tolerance among wild-type NucA, HighSalt NucA, and 19 mutants with single mutations at residue A95.
[0057] Figure 12 Comparison of salt tolerance among wild-type NucA, HighSalt NucA, and 19 mutants with single mutations at residue A102.
[0058] Figure 13 Comparison of salt tolerance among wild-type NucA, HighSalt NucA, and 19 mutants with single mutation at residue D149.
[0059] Figure 14 .Comparison of salt tolerance among wild-type NucA, HighSalt NucA, and mutants with combined positively charged residues. DETAILED DESCRIPTION
[0060] All patents, patent applications, published applications, and other publications cited herein are hereby incorporated by reference herein for their entirety and for all purposes. If a term or phrase used herein conflicts or is inconsistent with a definition of a term or phrase in the patents, patent applications, published applications, and other publications incorporated herein by reference, the usage in this document takes precedence over the definitions incorporated by reference.
[0061] definition
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. If there are multiple definitions of a term herein, the definition in this section shall prevail unless otherwise stated.
[0063] As used herein, the singular forms "a," "an," and "the" include plural references unless the context or clear statement indicates otherwise. For example, a dimer includes one or more dimers unless the context or clear statement indicates otherwise.
[0064] The term "amplification" ("polymerase extension reaction") refers to an increase in the number of copies of a polynucleotide.
[0065] As used herein, "sequence identity" or "identity" in the context of two protein sequences (or nucleotide sequences) refers to the residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window.
[0066] Sequence identity is usually expressed as "% sequence identity" or "% identity". To determine the percent identity between two amino acid sequences, the first step is to generate a pairwise sequence alignment between the two sequences, wherein the two sequences are aligned over their entire length (i.e., a pairwise global alignment). This alignment is generated by a program that implements the algorithm of Needleman and Wunsch (J. Molecular Biology (1979) 48, pp. 443-453), such programs are within the routine skill of those skilled in the art, for example, the "NEEDLE" program. For the purposes of this specification, the preferred alignment is the alignment from which the highest sequence identity can be determined.
[0067] After the two sequences are aligned, the second step is to determine the identity value from the generated alignment. For the purposes of this specification, percent identity is calculated using the following formula: % identity = (number of identical residues / length of the aligned region showing the entire length of the corresponding sequences described in this specification) × 100.
[0068] Therefore, according to this embodiment, sequence identity involving the comparison of two amino acid sequences is calculated as follows: the number of identical residues is divided by the length of the comparison region showing the entire length of the corresponding sequences as described in this specification, and this value is multiplied by 100 to obtain "% identity".
[0069] The calculation method for calculating the percentage identity between two DNA sequences is the same as that for calculating the percentage identity between two amino acid sequences, except for certain special instructions.
[0070] For protein-encoding DNA sequences, pairwise alignments should be performed over the entire coding region from the start codon to the stop codon, excluding introns. In the other sequence being compared to the sequence described herein, if introns are present, they can also be removed before pairwise alignment. Percent identity is then calculated using the following formula: % identity = (number of identical residues / length of the aligned region showing the entire coding region from the start codon to the stop codon of the sequence described herein (excluding introns)) × 100.
[0071] Sequences that have identical or similar regions to the sequences described in this specification and are to be compared to determine % identity with the sequences described in this specification can be easily identified by various methods well known to those skilled in the art, for example, using publicly available computer methods and programs (such as BLAST provided by the National Center for Biotechnology Information (NCBI)).
[0072] The variant of the parent enzyme molecule may have an amino acid sequence that is at least n% identical to the corresponding parent enzyme amino acid sequence, wherein the parent enzyme has enzymatic activity, and n is an integer between 50 and 100, preferably 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99 (based on full-length polypeptide sequence comparison). Preferably, the variant enzyme having n% identity compared to the parent enzyme has enzymatic activity.
[0073] Enzyme variants can be defined by their sequence similarity to the parent enzyme. Sequence similarity is usually expressed as "% sequence similarity" or "% similarity". To calculate sequence similarity, the first step is to generate a sequence alignment as described above; the second step is to calculate the percentage similarity, where the percentage sequence similarity takes into account that a particular group of amino acids has similar properties (e.g., based on size, hydrophobicity, charge or other characteristics). Herein, a mutation that replaces an amino acid with a similar amino acid is referred to as a "conservative mutation". Enzyme variants containing conservative mutations appear to have minimal effect on protein folding, and therefore some of their enzymatic properties can be substantially maintained compared to the enzymatic properties of the parent enzyme.
[0074] Conservative amino acid substitutions can occur within the full length of the polypeptide sequence of a functional protein (e.g., an enzyme). In one embodiment, such mutations do not involve the functional domains of the enzyme. In one embodiment, conservative mutations do not involve the catalytic center of the enzyme.
[0075] For example, amino acid A is similar to amino acid S; amino acid D is similar to amino acids E and N; amino acid E is similar to amino acids D, K, and Q; amino acid F is similar to amino acids W and Y; amino acid H is similar to amino acids N and Y; amino acid I is similar to amino acids L, M, and V; amino acid K is similar to amino acids E, Q, and R; amino acid L is similar to amino acids I, M, and V; amino acid M is similar to amino acids I, L, and V; amino acid N is similar to amino acids D, H, and S; amino acid Q is similar to amino acids E, K, and R; amino acid R is similar to amino acids K and Q; amino acid S is similar to amino acids A, N, and T; amino acid T is similar to amino acid S; amino acid V is similar to amino acids I, L, and M; amino acid W is similar to amino acids F and Y; and amino acid Y is similar to amino acids F, H, and W.
[0076] In particular, variant enzymes comprising conservative mutations have at least m% similarity to the corresponding parent sequence (m is an integer between 50 and 100, preferably 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99, based on full-length polypeptide sequence comparison), and it is expected that their enzymatic properties are substantially unchanged. Preferably, variant enzymes having m% similarity to the parent enzyme have enzymatic activity.
[0077] "Homologous" refers to genes, polypeptides or polynucleotides that have a high degree of similarity in position, structure, function or characteristics, but their sequence identity is not necessarily highly identical.
[0078] As used herein, "substantially complementary or substantially matching" means that two nucleic acid sequences have at least about 90% sequence identity. Preferably, the two nucleic acid sequences have at least or at least about 95%, 96%, 97%, 98%, 99% or 100% sequence identity. Alternatively, "substantially complementary or substantially matching" means that the two nucleic acid sequences can hybridize under high stringency conditions.
[0079] The term "hybridization" as defined herein refers to the process by which substantially complementary nucleotide sequences anneal to each other. The hybridization process can occur entirely in solution, i.e., both complementary nucleic acids are in solution. The hybridization process can also occur when one of the complementary nucleic acids is fixed to a matrix (e.g., magnetic beads, agarose beads, or any other resin). In addition, the hybridization process can also occur when one of the complementary nucleic acids is fixed to a solid support (e.g., a nitrocellulose membrane or a nylon membrane) or is fixed to a silica glass support (the latter is referred to as a nucleic acid array, microarray, or nucleic acid chip) by, for example, photolithography. In order for hybridization to occur, the nucleic acid molecules are generally subjected to thermal or chemical denaturation to melt the double strands into two single strands and / or eliminate hairpin structures or other secondary structures of the single-stranded nucleic acids. Hybridization according to this specification refers to hybridization that must cover the entire length of the sequence of the present invention. As defined herein, such full-length hybridization means that when the sequences described herein are cut into fragments of 300-500 bases, each fragment will hybridize.
[0080] The term "stringency" refers to the conditions under which hybridization occurs. Hybridization stringency is affected by conditions such as temperature, salt concentration, ionic strength, and hybridization buffer composition. Typically, low stringency conditions are selected to be approximately 30°C below the thermal melting point (Tm) at a specific ionic strength and pH. Medium stringency conditions are temperatures 20°C below the Tm, and high stringency conditions are temperatures 10°C below the Tm. High stringency hybridization conditions are typically used to isolate hybridizing sequences that have a high degree of sequence identity to the target nucleic acid sequence. However, due to the degeneracy of the genetic code, nucleic acids may have sequence deviations but still encode essentially identical polypeptides. Therefore, moderate stringency hybridization conditions may sometimes be required to identify such nucleic acid molecules. "Tm" refers to the temperature at which 50% of the target sequence hybridizes to a perfectly matched probe under specific ionic strength and pH conditions. The Tm depends on the solution conditions, the base composition of the probe, and the length. For example, longer sequences hybridize specifically at higher temperatures. The maximum rate of hybridization occurs in the range of approximately 16°C to 32°C below the Tm. The presence of monovalent cations in the hybridization solution can reduce the electrostatic repulsion between the two nucleic acid chains, thereby promoting hybrid formation; this effect is significant at sodium concentrations up to 0.4M (the effect is negligible at higher concentrations). Formamide can reduce the melting temperature of DNA-DNA and DNA-RNA duplexes by 0.6 to 0.7°C per percentage of formamide. Adding 50% formamide can allow hybridization to be carried out at 30 to 45°C, although the hybridization rate will be reduced. Base pair mismatches will reduce the hybridization rate and the thermal stability of the duplex. On average, for large probes, each 1% base mismatch will reduce Tm by about 1°C. Tm can be calculated based on conventional knowledge in the art. In addition to hybridization conditions, hybridization specificity generally depends on the operation of post-hybridization washing. Those skilled in the art are aware of the various parameters that can be changed during the washing process, which will maintain or change the stringency conditions.
[0081] For example, for DNA hybrids greater than 50 nucleotides in length, typical high stringency hybridization conditions include hybridization at 65°C in 1X SSC, or hybridization at 42°C in 1X SSC and 50% formamide, followed by a wash in 0.3X SSC at 65°C.
[0082] For definition of stringency levels, reference may be made to Sambrook et al. (2001), Molecular Cloning: A Laboratory Manual, 3rd ed. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY) or Molecular Biology: A Laboratory Manual (John Wiley & Sons, New York, 1989 and annual updates).
[0083] As used herein, "primer" refers to a nucleic acid molecule that can anneal to a template nucleic acid and serve as a starting point for DNA amplification. A primer can be fully or partially complementary to a specific region of a template polynucleotide, for example 20 nucleotides upstream or downstream of the target codon. Non-complementary nucleotides are defined herein as mismatches. Mismatches can be located inside a primer or at either end of the primer. Preferably, a single nucleotide mismatch, more preferably two, and more preferably three or more continuous or non-continuous nucleotide mismatches are located within the primer. A primer can have, for example, 5 to 200 nucleotides, preferably 20 to 80 nucleotides, more preferably 43 to 65 nucleotides. More preferably, the primer has 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185 or 190 nucleotides. "Forward primer" as defined herein is a primer that is complementary to the negative strand of the template polynucleotide. "Reverse primer" as defined herein is a primer that is complementary to the positive strand of the template polynucleotide. Preferably, the forward primer and the reverse primer do not comprise overlapping nucleotide sequences. "Does not comprise overlapping nucleotide sequences" as defined herein refers to that the forward primer and the reverse primer will not anneal to the region where the positive and negative strands are complementary to each other in the negative and positive strands of the template polynucleotide, respectively. For primers that anneal to the same strand of a template polynucleotide, "does not contain overlapping nucleotide sequences" means that the primers do not contain a sequence that is complementary to the same region of the same strand of the template polynucleotide. As used herein, a "primer set" refers to a combination of a "forward primer" and a corresponding "reverse primer."
[0084] As used herein, "positive strand" is equivalent to "sense strand" and may also be referred to as "coding strand" or "non-template strand." This strand has the same sequence as the mRNA (except it contains T instead of U). The other strand, called the "template strand," the "minus strand" or "antisense strand," is complementary to the mRNA.
[0085] As described herein, "codon optimization" refers to a design process in which codons are changed into codons known to improve maximum protein expression efficiency. In some alternatives, codon optimization for expression in cells has been described, wherein codon optimization can be performed by algorithms known to those skilled in the art to create synthetic genetic transcripts optimized for high mRNA and protein yield in target host cells (such as bacteria, fungi, insects or mammalian cells including human cells). For example, codons can be optimized for protein expression in bacterial cells, mammalian cells, yeast cells, insect cells or plant cells. Programs comprising codon optimization algorithms for human cells are readily available. In some embodiments, the gene is codon optimized to be expressed in bacteria, yeast, fungi or insect cells.
[0086] The term "heterologous" (or exogenous, foreign, or recombinant) polypeptide is defined herein as: (a) a polypeptide that does not occur naturally in the host cell. The protein sequence of such a heterologous polypeptide is a synthetic, non-naturally occurring "artificial" protein sequence; (b) a polypeptide native to the host cell in which structural modifications (e.g., deletions, substitutions, and / or insertions) have been made to alter the native polypeptide; or (c) a polypeptide native to the host cell that is expressed in altered amounts or from a genomic location that is different from the native host cell due to manipulation of the host cell's DNA through recombinant DNA technology (e.g., a stronger promoter).
[0087] Items (b) and (c) above refer to sequences that are in their native form but are not naturally expressed by the cells used to produce the polypeptide. Therefore, the polypeptide produced is more accurately defined as a "recombinantly expressed endogenous polypeptide," which does not contradict the above definition but rather reflects the specific situation in which it is not the protein sequence that is synthesized or manipulated, but rather the method by which the polypeptide molecule is produced.
[0088] Similarly, the term "heterologous" (or exogenous, foreign or recombinant) polynucleotide refers to: (a) a polynucleotide that is not native to the host cell; (b) a polynucleotide that is native to the host cell in which structural modifications (e.g., deletions, substitutions and / or insertions) have been made to alter the native polynucleotide; (c) a polynucleotide that is native to the host cell and whose expression levels are altered due to manipulation of the polynucleotide's regulatory elements (e.g., a stronger promoter) through recombinant DNA techniques; or (d) a polynucleotide that is native to the host cell but has not been integrated into its natural genetic environment due to genetic manipulation using recombinant DNA techniques.
[0089] With respect to two or more polynucleotide sequences or two or more amino acid sequences, the term "heterologous" is used to characterize that the two or more polynucleotide sequences or two or more amino acid sequences do not occur in a specific combination in nature.
[0090] As used herein, "transgenic," "transgenic," or "recombinant" refers, for example, to a nucleic acid sequence, an expression cassette, a genetic construct or vector comprising the nucleic acid sequence, or an organism transformed with the nucleic acid sequence, expression cassette, or vector, to all constructs produced synthetically by recombinant or gene technology methods, wherein (a) the nucleic acid sequence comprising the desired genetic information to be expressed, or (b) a genetic control sequence (such as a promoter) operably linked to the nucleic acid sequence comprising the desired genetic information, or (c) both (a) and (b) are not located in their natural genetic environment or have been modified by recombinant methods. The natural genetic environment is understood to be the natural genome or chromosomal site in the original organism. A naturally occurring expression cassette (e.g., the natural combination of a natural promoter of a nucleic acid sequence and a corresponding nucleic acid sequence encoding a polypeptide) becomes a transgenic expression cassette after modification by human intervention (e.g., mutagenesis). In addition, a naturally occurring expression cassette becomes a recombinant expression cassette when it is separated from its natural genetic environment and subsequently reintroduced into a non-natural genetic environment.
[0091] "Synthetic" or "artificial" compounds are produced by in vitro chemical or enzymatic synthesis. They include, but are not limited to, nucleic acid variants prepared using optimal codon usage for a host organism (e.g., a yeast cell host or other selected expression host), or protein variants that have amino acid modifications (e.g., substitutions) compared to the parent protein sequence, e.g., to optimize the properties of the polypeptide.
[0092] A "reference sequence" is a defined sequence used as a basis for sequence comparison. A reference sequence may be a subset of a larger sequence, such as a segment of a full-length cDNA or gene sequence listed in a sequence listing, or it may comprise the entire cDNA or gene sequence. Typically, a reference sequence is at least 20 nucleotides in length, usually at least 25 nucleotides, and often at least 50 nucleotides in length.
[0093] The terms "fragment," "derivative," and "analog," when referring to a reference polypeptide, include polypeptides that retain at least one biological function or activity that is at least substantially the same as the reference polypeptide.
[0094] The term "functional fragment" refers to any nucleic acid or amino acid sequence that comprises only a portion of a full-length nucleic acid or full-length amino acid sequence, respectively, but still has the same or similar activity and / or function. In one embodiment, the fragment comprises at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the original sequence. In one embodiment, the functional fragment comprises contiguous nucleic acids or amino acids, respectively, compared to the original nucleic acid or amino acid sequence.
[0095] The term "gene" refers to a segment of DNA involved in producing a polypeptide chain; it includes regions preceding and following the coding region (leader and trailer sequences), as well as optional intervening sequences (introns) between individual coding segments (exons).
[0096] As used herein, the term "isolated" refers to a substance that has been removed from its original environment (e.g., if naturally occurring, its natural environment). For example, a naturally occurring polynucleotide or enzyme present in a living animal is not isolated, but the same polynucleotide or enzyme is isolated when separated from some or all of the coexisting materials in the natural system. Such a polynucleotide can be part of a vector and / or such a polynucleotide or enzyme can be part of a composition and still be isolated because the vector or composition is not part of its natural environment. As another example, an isolated nucleic acid (e.g., a DNA or RNA molecule) is a nucleic acid molecule that is not directly adjacent to the 5' and 3' flanking sequences that would normally be directly adjacent to it when present in the native genome of the source organism. Such a polynucleotide can be part of a vector, integrated into the genome of a cell with an unrelated genetic background (or integrated into the genome of a cell with a substantially similar genetic background but at a different location from the naturally occurring site), produced by PCR amplification or restriction enzyme digestion, or as an RNA molecule produced by in vitro transcription, and / or such a polynucleotide, polypeptide, or enzyme can be part of a composition and still be isolated because such a vector or composition is not part of its natural environment.
[0097] The term "isolated" refers to DNA incorporated into a vector, such as a plasmid or viral vector; nucleic acid integrated into the genome of a heterologous cell (or into the genome of a homologous cell but at a non-natural site); and nucleic acid existing as an independent molecule, such as a DNA fragment generated by PCR amplification or restriction enzyme digestion or an RNA molecule generated by in vitro transcription.
[0098] The term "purified" as used herein does not require absolute purity, but is intended as a relative definition. Individual nucleic acids obtained from a library are typically purified to electrophoretic homogeneity. For example, a purified nucleic acid of the present disclosure may be purified to at least 10% purity from the remaining genomic DNA of an organism. 4 -10 6 times. However, the term "purified" also includes nucleic acids that have been purified from the rest of the genomic DNA or other sequences in a library or other environment by at least one order of magnitude (usually two or three orders of magnitude, more usually four or five orders of magnitude). "Purified" means that the material is in a relatively pure state, for example, at least about 90% pure, at least about 95% pure, or at least about 98% or 99% pure. Preferably, "purified" means that the material is in a 100% pure state.
[0099] The term "operably linked" refers to a relationship between the described components that allows them to function in the intended manner. For example, a regulatory sequence operably linked to a coding sequence is linked in such a way that expression of the coding sequence is achieved under conditions compatible with the control sequences. As used herein, a promoter sequence is "operably linked to" a coding sequence if RNA polymerase initiating transcription at the promoter is capable of transcribing the coding sequence into mRNA.
[0100] The term "mutation" is defined as a change in the genetic code of a nucleic acid sequence or an alteration in a peptide sequence. Such mutations can be point mutations, such as transitions or transversions. A mutation can be a change in one or more nucleotides or the encoded amino acid sequence. A mutation can be a deletion, insertion, or duplication.
[0101] The terms "polynucleotide," "nucleic acid sequence," "nucleotide sequence," "nucleic acid," and "nucleic acid molecule" are used interchangeably herein to refer to a polymeric, unbranched form of nucleotides of any length, which may be ribonucleotides, deoxyribonucleotides, or a combination of both.
[0102] The term "nucleic acid sequence encoding a specific protein or polypeptide" or "DNA coding sequence of..." or "nucleotide sequence encoding a specific protein or polypeptide" refers to a DNA sequence that is transcribed and translated into a protein or polypeptide when placed under the control of appropriate regulatory sequences.
[0103] The terms "nucleic acid encoding a protein or peptide" or "DNA encoding a protein or peptide" or "polynucleotide encoding a protein or peptide" and other synonymous terms include polynucleotides that contain only protein or peptide coding sequences, as well as polynucleotides that contain additional coding and / or non-coding sequences.
[0104] The terms "regulatory element," "control sequence," and "promoter" are used interchangeably herein and should be broadly understood to refer to regulatory nucleic acid sequences that can affect the expression of sequences associated therewith. "Regulatory element" or "regulatory nucleotide sequence" herein may refer to a nucleic acid fragment that drives the expression of a nucleic acid sequence after transformation into a host cell or organelle. A regulatory nucleotide sequence may include any nucleotide sequence that has a function or use alone in a specific arrangement or in a grouping of other elements or sequences in the arrangement. Examples of regulatory nucleotide sequences include, but are not limited to, transcriptional control elements such as promoters, enhancers, and termination elements. The regulatory nucleotide sequence may be the natural sequence (i.e., from the same gene) or an exogenous sequence (i.e., from a different gene) of the nucleotide sequence to be expressed.
[0105] The term "promoter" generally refers to a nucleic acid control sequence located upstream of the start point of gene transcription, which participates in the recognition and binding of RNA polymerase and other proteins, thereby guiding the transcription of operably linked nucleic acids. "Promoter" herein may further include any nucleic acid sequence capable of driving transcription of a coding sequence. In particular, the term "promoter" used herein may refer to a polynucleotide sequence generally described as a gene 5' regulatory region, located near the start codon. Transcription of one or more coding sequences is initiated in the promoter region. The term may also include promoter fragments that play a role in initiating gene transcription. A promoter may also be referred to as a "transcription start site" (TSS).
[0106] The above terms encompass transcriptional regulatory sequences derived from classic eukaryotic genomic genes (including the TATA box, with or without a CCAAT box sequence, required for accurate transcription initiation), as well as other regulatory elements (i.e., upstream activating sequences, enhancers, and silencers) that alter gene expression in response to developmental and / or external stimuli or in a tissue-specific manner.
[0107] For example, enhancers, as known in the art and used herein, are generally short DNA segments (eg, 50-1500 bp) to which proteins (eg, transcription factors) bind to increase the likelihood that a coding sequence will be transcribed.
[0108] Other elements may be "transcription termination elements," which comprise a nucleic acid sequence segment that marks the end of a gene and mediates transcription termination by providing a signal within the mRNA that initiates release of the mRNA from the transcription complex. Transcription termination in prokaryotes and eukaryotes is common knowledge in the art.
[0109] "Oligonucleotide" (or synonymously "oligo") refers to a single-stranded polydeoxynucleotide or two complementary polydeoxynucleotide chains that can be chemically synthesized. Such synthetic oligonucleotides may or may not have a 5' phosphate group.
[0110] Any source of nucleic acid in purified form can be used as the starting nucleic acid (also defined as a "template polynucleotide"). Thus, the method can be used with DNA or RNA (including messenger RNA), which can be single-stranded, preferably double-stranded. In addition, DNA-RNA hybrids containing one strand of each strand can be used. The length of the nucleic acid sequence can vary depending on the size of the nucleic acid sequence to be mutated. Preferably, the specific nucleic acid sequence is 50 to 50,000 base pairs, more preferably 50-11,000 base pairs.
[0111] All methods and materials similar or equivalent to those described herein can be used to practice or test the methods and compositions disclosed herein, and suitable methods and materials are described herein. All publications, patent applications, patents, and other references mentioned herein are incorporated herein by reference in their entirety. In addition, unless otherwise indicated, the materials, methods, and examples are for illustrative purposes only and are not intended to be limiting.
[0112] Nuclease polypeptide
[0113] Disclosed herein are salt-tolerant polypeptides with nuclease activity. To overcome the limitations of nucleases in high salt concentrations, the inventors modified a nuclease from Serratia marcescens by introducing positive amino acid residues to reshape its electrostatic surface. Screening revealed a nuclease mutant capable of tolerating at least 200 mM NaCl without a significant decrease in enzyme activity. This mutant may extend its application to solutions with even higher salt concentrations.
[0114] Serratia marcescens nuclease A is an extracellular enzyme that has been widely used for nucleic acid removal in a variety of application scenarios. However, the enzyme is sensitive to salt, and its enzymatic activity is significantly inhibited by solutions with ionic strengths >200mM. By analyzing the crystal structure of nuclease A, it was found that its interaction with nucleic acids basically depends on electrostatic interactions. High salt concentrations destroy these interactions, resulting in the inhibition of the enzymatic activity of nuclease A under high salt conditions. Therefore, enhancing nucleic acid-protein interactions is an effective way to improve the salt tolerance of nuclease A. Based on this principle, the inventors modified Serratia marcescens nuclease A by introducing more positively charged residues (including residues S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156) into the predicted nucleic acid binding surface. Mutating these residues to specific amino acids will make the nucleic acid binding surface more positively charged, thereby enhancing binding to nucleic acids. It should be noted that the improved nucleic acid binding ability may also affect the dissociation of nucleic acid and protein after cleavage. If the nucleic acid binding is too strong, it may also inhibit the enzyme activity. Therefore, a balance must be achieved between nucleic acid affinity and enzyme activity. The inventors designed 6 mutants by combining the mutations of the above 11 residues and expressed and purified all mutants. By comparing the electrostatic potential surface of wild-type nuclease A and mutant nuclease A, it was found that mutant nuclease A does have an enlarged positively charged surface ( Figure 1 Rational engineering of Serratia marcescens nuclease A by introducing positively charged residues into specific surface regions significantly enhanced the enzyme's tolerance to high salt conditions.
[0115] In some embodiments, N79 is mutated to a non-polar amino acid (such as alanine, valine, isoleucine, proline, phenylalanine, tryptophan, methionine or glycine), or to an uncharged polar amino acid (such as serine, threonine, glutamine, tyrosine or cysteine), or to a positively charged polar amino acid (such as histidine, lysine or arginine). In certain embodiments, N79 is mutated to alanine (Ala), cysteine (Cys), glycine (Gly), glutamine (Gln), serine (Ser), threonine (Thr), valine (Val), tryptophan (Trp), tyrosine (Tyr), lysine (Lys) or arginine (Arg).
[0116] In some embodiments, A95 is mutated to a non-polar amino acid (such as valine, isoleucine, proline, phenylalanine, tryptophan, methionine or glycine), or to an uncharged polar amino acid (such as serine, threonine, glutamine, asparagine, tyrosine or cysteine), or to a positively charged polar amino acid (such as histidine, lysine or arginine). In certain embodiments, A95 is mutated to phenylalanine (Phe), asparagine (Asn), proline (Pro), glutamine (Gln), serine (Ser), threonine (Thr), valine (Val), tyrosine (Tyr), lysine (Lys) or arginine (Arg).
[0117] In some embodiments, A102 is mutated to a non-polar amino acid (such as valine, isoleucine, proline, phenylalanine, tryptophan, methionine or glycine), or to an uncharged polar amino acid (such as serine, threonine, glutamine, asparagine, tyrosine or cysteine), or to a positively charged polar amino acid (such as histidine, lysine or arginine). In certain embodiments, A102 is mutated to threonine (Thr), valine (Val), tryptophan (Trp), tyrosine (Tyr), lysine (Lys) or arginine (Arg).
[0118] In some embodiments, D149 is mutated to a non-polar amino acid (such as alanine, valine, isoleucine, proline, phenylalanine, tryptophan, methionine or glycine), or to an uncharged polar amino acid (such as serine, threonine, glutamine, asparagine, tyrosine or cysteine), or to a positively charged polar amino acid (such as histidine, lysine or arginine). In certain embodiments, D149 is mutated to alanine (Ala), phenylalanine (Phe), histidine (His), asparagine (Asn), glutamine (Gln), serine (Ser), threonine (Thr), valine (Val), tryptophan (Trp), tyrosine (Tyr), lysine (Lys) or arginine (Arg).
[0119] In some embodiments, S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156 are mutated to positively charged polar amino acids, preferably to histidine, lysine, or arginine, and more preferably to lysine or arginine.
[0120] Serratia marcescens nuclease A is a polypeptide having the amino acid sequence of SEQ ID NO: 1 and exhibiting nuclease activity. The first 21 amino acids at the N-terminus of SEQ ID NO: 1 are a signal sequence, and the remaining amino acids constitute the mature polypeptide.
[0121] In some embodiments, the polypeptides disclosed herein are isolated, synthetic, or recombinant polypeptides comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or more sequence identity to SEQ ID NO: 1 or its mature polypeptide, and the polypeptides have nuclease activity. For example, the polypeptides may have 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% sequence identity to SEQ ID NO: 1 or its mature polypeptide, or within a range between any two of these values.
[0122] In some embodiments, the polypeptide: (a) comprises a mutation at one, two, three, four, five, six, seven or more or all of S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, D156 of SEQ ID NO: 1, wherein the mutation comprises a modification, substitution or deletion of an amino acid. In some embodiments, a polypeptide comprising an amino acid sequence having at least 70% sequence identity to the sequence of (a) and having a mutation at the corresponding position described in (a) retains nuclease activity. The amino acid mutations described herein are relative to the corresponding amino acid positions in SEQ ID NO: 1. For example, the mutation of amino acid position 74 of SEQ ID NO: 1 from S to K is described herein as S74K.
[0123] In some embodiments, the polypeptide has mutations at at least four sites selected from S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156 of SEQ ID NO: 1 or its mature polypeptide. In some embodiments, the polypeptide has mutations at D149 and at least three sites selected from S74, N79, A95, T98, N101, A102, S137, D138, Q141, and D156 of SEQ ID NO: 1. In certain embodiments, each site is mutated to lysine or arginine.
[0124] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or its mature polypeptide, and has mutations at at least three positions selected from N79, A95, A102, and D149, and optionally has a mutation at at least one position selected from S74, T98, N101, S137, D138, Q141, and D156. In certain embodiments, each position is mutated to lysine or arginine.
[0125] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or its mature polypeptide, and has mutations at positions A95, A102, and D149, and optionally at least one mutation selected from the group consisting of S74, N79, T98, N101, S137, D138, Q141, and D156. In certain embodiments, each position is mutated to lysine or arginine.
[0126] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof and has mutations at positions N79, A95, and D149 and optionally at least one position selected from S74, T98, N101, A102, S137, D138, Q141, D156. In certain embodiments, each position is mutated to lysine or arginine.
[0127] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof and has mutations at positions N79, A95, and D149 and optionally at least one position selected from S74, T98, N101, A102, S137, D138, Q141, D156. In certain embodiments, each position is mutated to lysine or arginine.
[0128] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof and has mutations at positions N79, A95, and D149 and optionally at least one position selected from S74, T98, N101, A102, S137, D138, Q141, D156. In certain embodiments, each position is mutated to lysine or arginine.
[0129] In some embodiments, the polypeptide has mutations at positions N79, A95, A102, and D149 of SEQ ID NO: 1 and optionally at least one position selected from S74, T98, N101, S137, D138, Q141, D156. In some embodiments, the polypeptide has mutations at positions N79, A95, A102, and D149 and optionally at one or two positions selected from Q141 and N101. In certain embodiments, each position is mutated to lysine or arginine.
[0130] In some embodiments, the polypeptide has mutations at positions A95, N101, Q141, D149 of SEQ ID NO: 1 and optionally at least three positions selected from S74, N79, T98, A102, S137, D138, D156. In some embodiments, the polypeptide has mutations at positions A95, N101, Q141, D149 and optionally at positions S74, N79, and T98. In certain embodiments, each position is mutated to lysine or arginine.
[0131] In some embodiments, the polypeptide may be a polypeptide comprising an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 1 or its mature polypeptide, wherein the polypeptide has nuclease activity under high solution ionic strength, and wherein the polypeptide comprises (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D14 9R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K and (11) D156R; for example, one or more mutations of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K and D156R.
[0132] In some embodiments, the polypeptide in SEQ ID NO: 1 has at least 4 mutations selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R (e.g., selected from the group consisting of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R). In some embodiments, the polypeptide in SEQ ID NO: 1 has a D149K or D149R mutation and at least 3 (or at least 4, at least 5, at least 6, at least 7 or more) mutations selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K and (11) D156R (e.g., the group consisting of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K and D156R).
[0133] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof, and has at least three mutations selected from (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, and optionally has a mutation selected from S74K, T98K, N101K, S137K, D138K, Q141K, and D156R.
[0134] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof, and has three mutations of (1) N79K or N79R, (2) A95K or A95R, and (3) A102K or A102R, and optionally has a mutation selected from (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R.
[0135] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof, and has three mutations of (1) N79K or N79R, (2) A95K or A95R, and (3) D149K or D149R, and optionally has a mutation selected from (4) A102K or A102R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R.
[0136] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof, and has three mutations of (1) N79K or N79R, (2) A102K or A102R, and (3) D149K or D149R, and optionally has a mutation selected from (4) A95K or A95R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R.
[0137] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and having three mutations of (1) A95K or A95R, (2) A102K or A102R and (3) D149K or D149R, and optionally having a mutation selected from (4) N79K or N79R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K and (11) D156R.
[0138] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and having three mutations of (i) A95K, A102K and D149K or (ii) A95R, A102R and D149R, and optionally having a mutation selected from (1) N79K or N79R, (2) S74K, (3) T98K, (4) N101K, (5) S137K, (6) D138K, (7) Q141K and (8) D156R.
[0139] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and having three mutations of (i) N79K, A102K and D149K or (ii) N79R, A102R and D149R, and optionally having a mutation selected from (1) A95K or A95R, (2) S74K, (3) T98K, (4) N101K, (5) S137K, (6) D138K, (7) Q141K and (8) D156R.
[0140] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and having three mutations of (i) N79K, A95K and D149K or (ii) N79R, A95R and D149R, and optionally having a mutation selected from (1) A102K or A102R, (2) S74K, (3) T98K, (4) N101K, (5) S137K, (6) D138K, (7) Q141K and (8) D156R.
[0141] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof, and has three mutations of (i) N79K, A95K, and A102K or (ii) N79R, A95R, and A102R, and optionally has a mutation selected from (1) D149K or D149R, (2) S74K, (3) T98K, (4) N101K, (5) S137K, (6) D138K, (7) Q141K, and (8) D156R.
[0142] In some embodiments, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or a mature polypeptide thereof, and has four mutations of (i) N79K, A95K, A102K, and D149K or (ii) N79R, A95R, A102R, and D149R, and optionally has a mutation selected from (1) S74K, (2) T98K, (3) N101K, (4) S137K, (5) D138K, (6) Q141K, and (7) D156R.
[0143] In some embodiments, the polypeptide has mutations of (i) N79K, A95K, A102K, D149K or (ii) N79R, A95R, A102R, and D149R, and optionally has a mutation selected from S74K, T98K, N101K, S137K, D138K, Q141K, and D156R. In some embodiments, the polypeptide has mutations of (i) N79K, A95K, A102K, and D149K or (ii) N79R, A95R, A102R, and D149R, and optionally has a mutation selected from Q141K and N101K.
[0144] In some embodiments, the polypeptide has mutations of A95K, N101K, Q141K, D149K in SEQ ID NO: 1, and optionally has a mutation selected from S74K, N79K, T98K, A102K, S137K, D138K, and D156R. In some embodiments, the polypeptide has mutations of A95K, N101K, Q141K, D149K, and optionally has mutations of S74K, N79K, and T98K. In some embodiments, the polypeptide has mutations of S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R in SEQ ID NO: 1.
[0145] In some embodiments, the polypeptide comprises a mutation or combination of mutations selected from the group consisting of: (a) N79K, A95K, A102K, and D149K; (b) N79K, A95K, A102K, Q141K, and D149K; (c) N79K, A95K, N101K, A102K, Q141K, and D149K; (d) A95K, N101K, Q141K, and D149K; (e) S74K, N79K, A95K, T98K, N101K, Q141K, and D149K; (f) S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0146] In some embodiments, the polypeptide is any one of the nuclease variants disclosed herein, wherein the polypeptide has nuclease activity.
[0147] In some embodiments, the nuclease polypeptide has thermotolerance. For example, the nuclease polypeptide can be more thermotolerant than the nuclease with SEQ ID NO:1 sequence. In some embodiments, at a given temperature (for example, at a temperature between 30°C and 60°C), the nuclease activity of the polypeptide is higher than the nuclease with SEQ ID NO:1 sequence by at least 1%, 2%, 3%, 4%, 5%, 7%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% or more. In some embodiments, at a given temperature, for example, at a temperature between 30°C and 60°C, the nuclease activity of the polypeptide is higher than the nuclease activity with SEQ ID NO:1 sequence by 1%, 2%, 3%, 4%, 5%, 7%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 75%, 100% or a value between these values. In some embodiments, at a given pH, for example, at pH 4 to pH 11, the nuclease activity of the polypeptide is at least 1%, 2%, 3%, 4%, 5%, 7%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% or more higher than the nuclease activity of the nuclease having the sequence of SEQ ID NO: 1.
[0148] In some embodiments, the nuclease polypeptide has salt tolerance. In some embodiments, compared with the nuclease with SEQ ID NO: 1 sequence, the mutant polypeptide has at least 60%, at least 70%, at least 80%, at least 90% or at least 95% nuclease activity at a solution ionic strength greater than 200mM. Ions herein are monovalent ions or divalent ions. In some embodiments, the solution ionic strength is greater than 300mM, for example, greater than 400mM, 500mM, 600mM, 700mM, 800mM, 900mM, 1mM, 2mM, 3mM, 4mM, 5mM or higher ionic strength.
[0149] The optimal ionic strength of the polypeptide can be different (e.g., higher or lower) than the optimal ionic strength of the nuclease having the sequence of SEQ ID NO: 1 or its parent nuclease. For example, the optimal ionic strength of the polypeptide can be 10 mM, 20 mM, 30 mM, 40 mM, 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM, 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, or a range between any two of these values, which is higher than the optimal ionic strength of the nuclease having the sequence of SEQ ID NO: 1. In some embodiments, the optimal ionic strength of the polypeptide is or is about 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM, 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, or a range between any two of these values. In some embodiments, the optimal temperature of the polypeptide is between 100 mM and 600 mM.
[0150] During fermentation, DNA from the production host can complicate multiple steps in the protein recovery process. Currently, removing DNA from the final product typically requires the addition of expensive materials from other sources, such as DNase. In some embodiments, the nucleases disclosed herein can be expressed in the same production host as the target product, eliminating the need for the addition of external DNase and reducing the overall processing cost of the final product.
[0151] The nuclease polypeptides disclosed herein may have one or more signal sequences. In some embodiments, at least one of the one or more signal sequences is heterologous to the nuclease polypeptide it comprises. In some embodiments, the nuclease polypeptides disclosed herein do not comprise any signal sequence.
[0152] Also disclosed herein are antibodies or binding fragments thereof (e.g., isolated or purified antibodies or binding fragments thereof) that specifically bind to an isolated, synthetic, or recombinant polypeptide comprising an amino acid sequence having at least 80%, 85%, 90%, 95%, 98%, 99% or greater sequence identity to SEQ ID NO: 1 or its mature polypeptide, and wherein the polypeptide has nuclease activity. In some embodiments, the polypeptide comprises one or more mutations selected from S74K, N79K, A95K, T98K, N101K, A102K, S137K, D138K, Q141K, D149K, and D156R.
[0153] In some embodiments, the polypeptide comprises a combination of mutations selected from the group consisting of: (1) N79K, A95K, A102K, and D149K; (2) N79K, A95K, A102K, Q141K, and D149K; (3) N79K, A95K, N101K, A102K, Q141K, and D149K; (4) A95K, N101K, Q141K, and D149K; (5) S74K, N79K, A95K, T98K, N101K, Q141K, and D149K; (6) S74K, N79K, A95K, T98K, N102K, A102K, Q141K, and D149K K, S137K, D138K, Q141K, D149K and D156R; (7) N79R, A95R, A102R, D149R; (8) A95K, A102K, D149K; (9) N79K, A102K, D149K; (10) N79K, A95 K, D149K; (11) N79K, A95K, A102K; (12) A95R, A102R, D149R; (13) N79R, A102R, D149R; (14) N79R, A95R, D149K; (15) N79R, A95R, A102R.
[0154] Nuclease variants disclosed herein may comprise one or more substitutions, deletions, and insertions at one or more amino acid positions of the nuclease. In some embodiments, the number of amino acid substitutions, deletions, and / or insertions introduced into a parent nuclease (e.g., a nuclease having a sequence of SEQ ID NO: 1) is no more than 30, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, or 29. The amino acid changes may be minor, i.e., conservative amino acid substitutions or insertions that do not significantly affect the folding and / or activity of the protein; small deletions (e.g., 1-20 amino acids); small amino or carboxyl terminal extensions, such as an amino terminal methionine residue; small connecting peptides of up to 20-25 residues in length; or small extensions that facilitate purification by altering the net charge or other functions (e.g., a polyhistidine sequence, an antigenic epitope, or a binding domain). In some embodiments, amino acid changes to a nuclease can alter one or more physicochemical properties of the parent nuclease.
[0155] Examples of conservative substitutions include replacements within the basic amino acid group (arginine, lysine, and histidine), the acidic amino acid group (glutamic acid and aspartic acid), the polar amino acid group (glutamine and asparagine), the hydrophobic amino acid group (leucine, isoleucine, and valine), the aromatic amino acid group (phenylalanine, tryptophan, and tyrosine), and the small amino acid group (glycine, alanine, serine, threonine, and methionine). Amino acid substitutions that generally do not alter specific activity are known in the art, such as those described by H. Neurath and RL Hill in 1979 in The Proteins (Academic Press, New York). Non-limiting exemplary amino acid substitutions include alanine to serine, valine to isoleucine, aspartic acid to glutamic acid, threonine to serine, alanine to glycine, alanine to threonine, serine to asparagine, alanine to valine, serine to glycine, tyrosine to phenylalanine, alanine to proline, lysine to arginine, aspartic acid to asparagine, leucine to isoleucine, leucine to valine, alanine to glutamic acid, and aspartic acid to glycine.
[0156] Amino acid replacement, deletion and / or insertion can be performed and tested using protein / DNA engineering methods known in the art, including but not limited to mutagenesis, recombination and / or shuffling, followed by relevant screening procedures. In some embodiments, mutagenesis / shuffling methods can be combined with high-throughput automated screening methods to detect the activity of cloned mutagenized polypeptides expressed by host cells. Mutagenized DNA molecules encoding active polypeptides can be recovered from host cells and rapidly sequenced using standard methods in the art. These methods can quickly determine the importance of individual amino acid residues in a polypeptide.
[0157] In some embodiments, the nuclease polypeptides of the present invention can tolerate higher or lower ionic strengths compared to other nucleases (e.g., nucleases having SEQ ID NO: 1). For example, these nuclease polypeptides can retain nuclease activity (e.g., retain at least 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 98% of their nuclease activity) at 50 mM, 60 mM, 70 mM, 80 mM, 90 mM, 100 mM, 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, or any two of these values.
[0158] The nuclease polypeptides disclosed herein may have the same or different substrate specificities compared to the nuclease having the sequence of SEQ ID NO: 1 or its respective parent nuclease. For example, these nuclease polypeptides may have substantially the same substrate specificity compared to the nuclease having the sequence of SEQ ID NO: 1 or its respective parent nuclease. In some embodiments, the substrate specificity of these nuclease polypeptides is about 50%, 60%, 70%, 80%, 90%, 95%, 98%, or a range between any two of the above values compared to the nuclease having the sequence of SEQ ID NO: 1 or its respective parent nuclease. Also disclosed herein is a composition comprising one or more nuclease polypeptides disclosed herein.
[0159] Also provided herein are immobilized nuclease polypeptides, wherein the immobilized polypeptide comprises one of the nuclease polypeptides disclosed herein. In some embodiments, the polypeptide can be immobilized on a cell, a metal, a resin, a polymer, a ceramic, a glass, a microelectrode, a graphite particle, a bead, a gel, a plate, an array, a capillary, or a combination thereof.
[0160] Nucleic Acids
[0161] As used herein, the terms "nucleic acid," "nucleotide," "polynucleotide," or "nucleic acid sequence" are used interchangeably and can be in the form of DNA or RNA. DNA includes cDNA, genomic DNA, or synthetic DNA. DNA can be single-stranded or double-stranded. DNA can be a coding strand or a non-coding strand. When referring to nucleic acids, the term "variant" as used herein can be a naturally occurring allelic variant or a non-naturally occurring variant. These nucleotide variants include degenerate variants, substitution variants, deletion variants, and insertion variants. As known in the art, an allelic variant is a substituted form of a nucleic acid that can be a substitution, deletion, or insertion of one or more nucleotides without substantially altering the function of the protein it encodes. The nucleic acids of the present invention can comprise a nucleotide sequence that has at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or 100% sequence identity to the nucleic acid sequence. The present invention also relates to nucleic acid fragments that hybridize to the above sequences. As used herein, a "nucleic acid fragment" is at least 15 nucleotides in length, preferably at least 30 nucleotides, more preferably at least 50 nucleotides, and even more preferably at least 100 nucleotides or more. The nucleic acid fragment can be used in nucleic acid amplification techniques (such as PCR).
[0162] The full-length coding sequence or fragments of the polypeptide of the present invention can be obtained by PCR amplification, artificial synthesis or recombination. Mutations in the polypeptide can be introduced during PCR, synthesis or recombination.
[0163] For PCR amplification, primers can be designed based on the nucleotide sequences disclosed herein, and the sequence of interest can be amplified using a commercially available cDNA library or a cDNA library prepared using conventional methods known to those skilled in the art as a template. When the nucleotide sequence is greater than 2500 bp, two to six rounds of PCR amplification are preferably performed, followed by splicing the individually amplified fragments together in the correct sequence. The PCR amplification procedures and systems described herein are not particularly limited, and conventional PCR amplification procedures and systems in the art can be used.
[0164] In addition, there are also methods for artificially synthesizing sequences of interest, particularly for short fragments. For example, when the nucleotide sequence of the optical probe is less than 2500bp, it can be synthesized by an artificial synthesis method. The artificial synthesis method can be any conventional DNA synthesis method in the art. Generally speaking, many small fragments are first synthesized and then connected to obtain a longer sequence. The DNA sequence encoding the protein of the present invention can also be obtained completely by chemical synthesis. The DNA sequence can then be introduced into a variety of existing DNA molecules known in the art (such as vectors) and introduced into cells.
[0165] Recombination can also be used to obtain related sequences in bulk, typically by cloning into a vector, transferring it into cells, and then isolating and purifying the polypeptide or protein of interest from the proliferating host cells by conventional methods.
[0166] Preparation of nuclease polypeptides and their variants
[0167] Provided herein are methods for modifying and preparing nucleic acid enzyme polynucleotide variants disclosed herein. Some embodiments provide synthetic or recombinant nucleic acids encoding one or more polypeptides disclosed herein, and vectors (e.g., expression vectors) comprising the nucleic acid. Non-limiting examples of the method include synthetic ligation reassembly, random mutagenesis, directed mutagenesis, optimized directed evolution systems and / or saturation mutagenesis (e.g., gene site saturation mutagenesis, GSSM), and any combination thereof. The term "variant" refers to a polynucleotide or polypeptide that is modified (respectively) at one or more base pairs, codons, introns, exons, or amino acid residues according to the description but still retains the biological activity of the nucleic acid enzyme. Variants can be prepared by methods such as error-prone PCR, shuffling, site-directed mutagenesis, assembly PCR, sexual PCR mutagenesis, in vivo mutagenesis (phage-assisted continuous evolution, in vivo continuous evolution), cassette mutagenesis, recursive ensemble mutagenesis, exponential ensemble mutagenesis, site-specific mutagenesis, gene reassembly, gene site-saturation mutagenesis, synthetic ligation reassembly, recombination, recursive sequence recombination, phosphorothioate-modified DNA mutagenesis, uracil-containing template mutagenesis, gapped duplex mutagenesis, point mismatch repair mutagenesis, repair-deficient host strain mutagenesis, chemical mutagenesis, radiation mutagenesis, deletion mutagenesis, restriction-selection mutagenesis, restriction-purification mutagenesis, artificial gene synthesis, ensemble mutagenesis, creation of chimeric nucleic acid multimers, and / or a combination of these methods with other methods.
[0168] A cloning vector, such as a vector comprising an expression cassette, can be used to express one or more nuclease polypeptides disclosed herein. The term "vector" as used herein encompasses any type of cloning vector, including, but not limited to, a plasmid, a phagemid, a viral vector (such as a bacteriophage), a phage, a baculovirus, a cosmid, a fosmid, an artificial chromosome, or any other vector for a particular host cell. Low copy number or high copy number vectors are also included. The exogenous polynucleotide sequence typically comprises a coding sequence, which can be referred to herein as a "gene of interest." The gene of interest can comprise introns and exons, depending on the source or use of the host cell. The cloning vector can be a viral vector, a plasmid, a phage, a phagemid, a cosmid, a fosmid, a phage, an artificial chromosome, or a combination thereof. The viral vector can comprise an adenoviral vector, a retroviral vector, or an adeno-associated viral vector. The cloning vector can comprise a bacterial artificial chromosome (BAC), a plasmid, a bacteriophage PI -derived vector (PAC), a yeast artificial chromosome (YAC), or a mammalian artificial chromosome (MAC). In some embodiments, the polynucleotide sequence encoding one or more nuclease polypeptides is integrated into the chromosome of a host cell, such that the polynucleotide becomes part of the host cell chromosome. The host cell can be, for example, a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, an insect cell, or a plant cell. In some embodiments, the polynucleotide sequence encoding one or more nuclease polypeptides is not located in the chromosome of a host cell.
[0169] Also provided herein are transformed host cells comprising a nucleic acid or expression cassette (e.g., a vector) or cloning vector comprising a nucleic acid sequence encoding one or more nuclease polypeptides disclosed herein.
[0170] Some embodiments provide a method of producing a recombinant polypeptide having nuclease activity, wherein the method comprises expressing a polynucleotide encoding one or more nuclease polypeptides disclosed herein under conditions that allow expression of at least one of the one or more nuclease polypeptides, thereby producing a recombinant polypeptide having nuclease activity. In some embodiments, the polynucleotide encoding one or more nuclease polypeptides disclosed herein is operably linked to a promoter. In some embodiments, the polypeptide is present in an expression vector. In some embodiments, the polynucleotide is present in a host cell to allow expression of the polypeptide. In some embodiments, the polynucleotide is present in the chromosome of a host cell to allow expression of the polypeptide. In some embodiments, the transformed host cell is a bacterial cell, a mammalian cell, a fungal cell, a yeast cell, an insect cell, or a plant cell. In some embodiments, the transformed host cell is a cell from E. coli.
[0171] Non-limiting examples of expression vectors include viral particles, baculovirus, phage, plasmid, phagemid, clay, F clay, bacterial artificial chromosome, viral DNA (e.g., derivatives of vaccinia virus, adenovirus, fowlpox virus, pseudorabies virus and SV40), artificial chromosomes based on Pi, yeast plasmids, yeast artificial chromosomes, and any other vectors for specific hosts (e.g., Bacillus, Aspergillus and yeast). The nuclease encoding DNA disclosed herein can be contained in any of a variety of expression vectors for expressing the nuclease polypeptide. Such vectors include chromosomal, non-chromosomal and synthetic DNA sequences. Many suitable vectors are known to those skilled in the art and are commercially available, such as pET-28a (Novagen). Depending on the desired use, low copy number or high copy number vectors can be used.
[0172] Codon optimization can be used to achieve high levels of protein expression in host cells. In some embodiments, the codons in the nucleic acids encoding one or more nuclease polypeptides disclosed herein can be optimized to increase or decrease their expression in host cells. For example, one or more non-preferred or less preferred codons in the nucleic acids encoding nuclease polypeptides can be replaced with one or more "preferred codons" encoding the same amino acids of the target host cell. "Preferred codons" used herein refer to codons that appear excessively in the coding sequence of a host cell gene, and "non-preferred or less preferred codons" refer to codons that are not expressed enough in the coding sequence of a host cell gene. For example, the codon-optimized encoding nucleic acid sequence of SEQ ID NO:1 is disclosed herein as SEQ ID NO:2.
[0173] According to the present disclosure, host cells for expressing nucleic acids, expression cassettes, and vectors can be eukaryotic or prokaryotic cells, including bacteria, yeast, fungi, plant cells, insect cells, and mammalian cells; and methods for optimizing codon usage in all of these cells, codon-altered nucleic acids, and polypeptides produced by the codon-altered nucleic acids are provided. Exemplary host cells include Gram-negative bacteria and Gram-positive bacteria. Exemplary host cells also include eukaryotic organisms, such as various yeasts, mammalian cells, and insect cells. In some embodiments, the host cell is a cell selected from the group consisting of Pichia pastoris, Bacillus subtilis, Pseudomonas fluorescens, Myceliopthora thermophilefungus, Tricodermea reesei, Escherichia coli, Bacillus licheniformis, Aspergillus niger, Schizosaccharomyces pombe, and Sacaramyces cerevisiae. The nucleic acid encoding the nuclease polypeptide disclosed herein can be located on the genome of the host cell, for example, a portion of a host cell chromosome. In some embodiments, the nucleic acid encoding the nuclease polypeptide is located on an expression vector isolated from the host cell genome. In some embodiments, the method for producing a nuclease polypeptide disclosed herein can include expressing the nucleic acid encoding the nuclease polypeptide under conditions that allow expression of the nuclease polypeptide, thereby producing the nuclease polypeptide. In some embodiments, the nucleic acid encoding the nuclease polypeptide is operably linked to an inducible promoter, e.g., a promoter inducible by changes in temperature and / or pH and / or the presence, absence, or amount / concentration of compounds such as IPTG, arabinose, tetracycline, steroids, and metals.
[0174] The present disclosure also encompasses nucleic acid and polypeptide sequences optimized for expression in the aforementioned organisms and species.
[0175] Composition and use
[0176] Provided herein is a method for degrading polynucleotides using one or more nuclease polypeptides disclosed herein. In some embodiments, the method includes contacting one or more polynucleotide molecules with one or more nuclease polypeptides disclosed herein to degrade the polynucleotide molecules. The polynucleotide molecules may comprise DNA (such as single-stranded or double-stranded DNA), RNA (such as single-stranded or double-stranded RNA) or any combination thereof. The contact reaction can be carried out at a variety of pH values, such as pH 4, pH 5, pH 6, pH 7, pH 8, pH 9, pH 10, pH 11 or the range between any two of the above values. The reaction temperature can also encompass a variety of intervals, such as 10°C, 15°C, 20°C, 25°C, 30°C, 35°C, 40°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C or the range between any two of the above values. The one or more polynucleotide molecules may be present in a variety of carriers, such as a detergent composition (such as a textile, clothing, cotton product, fabric or any combination thereof) or an aqueous solution (such as a reaction mixture).
[0177] Provided herein are compositions and kits comprising one or more nuclease polypeptides disclosed herein. The composition can be an enzyme composition, a detergent composition, a detergent additive, a food, a food supplement, a feed supplement, a feed, a pharmaceutical composition, a fermentation product, a fermentation intermediate, a fermentation downstream reaction mixture, a reaction mixture, or a combination thereof. In some embodiments, the reaction mixture is used for protein expression or purification. For example, the enzyme composition can include one or more nuclease polypeptides disclosed herein and a storage buffer. In some embodiments, the storage buffer includes a pH buffer system (e.g., Tris-HCl) providing a pH of about 6.0-9.0, such as pH 6.0-7.0 or about pH 6.5. For example, the storage buffer can include a stabilizer, such as glycerol, having a concentration of at least about 5%, 10%, 20%, 30%, 40%, 50% or more.
[0178] Also disclosed herein are uses of the polypeptides, nucleic acids, nucleic acid constructs, or host cells disclosed herein for the preparation of pharmaceutical compositions for the prevention or treatment of diseases or conditions, such as wounds, dental plaque, dental caries, periodontitis, natural valve endocarditis, chronic bacterial prostatitis, otitis media, infections associated with medical devices (such as artificial heart valves, artificial pacemakers, contact lenses, artificial joints, sutures, catheters, and arteriovenous shunts), infections associated with wounds, lacerations, ulcers, and mucosal injuries (such as ulcers), oral, oropharyngeal, nasopharyngeal, and hypopharyngeal infections, external ear infections, eye infections, stomach and large and small intestine infections, urinary tract and vaginal infections, skin infections, intranasal infections (such as sinus infections), or combinations thereof. In some embodiments, the method of using the composition comprises contacting the composition with a wound, laceration, ulcer, mucosal injury, or any combination thereof. In some embodiments, the method of using the composition comprises contacting the composition with a medical device. In some embodiments, the method of using the composition comprises contacting the composition with skin, external ear, eye, or a combination thereof. The composition can be formulated into various forms, such as tablets, gels, pills, implants, liquids, sprays, films, micelles, powders, foods, feed pellets, encapsulated forms, or combinations thereof.
[0179] In some embodiments, a composition comprising one or more nuclease polypeptides disclosed herein is used to contact a surface containing a DNA substrate and clean the surface.
[0180] In some embodiments, the nuclease polypeptides disclosed herein can be used alone or in combination with one or more other enzymes in an application or use. Some embodiments provide compositions comprising nucleases disclosed herein and variants thereof. In some embodiments, the composition may further comprise one or more other enzymes, one or more other components, or any combination thereof. The one or more other enzymes may include, but are not limited to, one or more of the following: nucleases, proteases, lipases, cutinases, amylases, carbohydrases, cellulases, pectinases, mannanases, arabinanases, galactanases, xylanases, oxidases (such as laccases and peroxidases) and deoxyribonucleases (DNases).
[0181] Also provided herein are pharmaceutically acceptable prodrugs of pharmaceutical compositions, and methods of treating using such prodrugs. The term "prodrug" refers to a precursor of a given compound that, after being administered to a subject, produces the compound in vivo by chemical or physiological processes (such as solvolysis or enzymatic cleavage) or under physiological conditions (e.g., the prodrug is converted into a drug when physiological pH is reached). A "pharmaceutically acceptable prodrug" is a prodrug that is non-toxic, bio-tolerated, and biologically suitable for administration to a subject. For example, Bundgaard's book "Prodrug Design" (Elsevier, 1985) describes an exemplary method for selecting and preparing suitable prodrug derivatives.
[0182] Also provided herein are pharmaceutically active metabolites of the pharmaceutical compositions and the use of such metabolites in the disclosed methods. A "pharmaceutically active metabolite" refers to a pharmacologically active product produced by the metabolism of a compound or its salt in vivo. Prodrugs and active metabolites of a compound can be determined using conventional techniques known or available in the art, for example, see Bundgaard, "Prodrug Design" (Elsevier, 1985).
[0183] Any suitable formulation of the compounds described herein can be prepared. See generally Remington's Pharmaceutical Sciences (2000, edited by Hoover JE, 20th edition). The choice of formulation must be suitable for the appropriate route of administration, including oral, parenteral, inhalation, topical, rectal, nasal, buccal, vaginal, via an implantable reservoir or other drug delivery method. When the compound has sufficient alkalinity or acidity to form a stable non-toxic acid or base salt, it may be appropriate to administer it in the form of a salt. Examples of pharmaceutically acceptable salts include organic acid salts formed with acids that form physiologically acceptable anions, such as tosylate, methanesulfonate, acetate, citrate, malonate, tartrate, succinate, benzoate, ascorbate, α-ketoglutarate, and α-glycerophosphate. Suitable inorganic salts can also be formed, including hydrochlorides, sulfates, nitrates, bicarbonates, and carbonates. Pharmaceutically acceptable salts can be obtained by standard procedures well known in the art, for example, by reacting a sufficiently alkaline compound (such as an amine) with a suitable acid to obtain a physiologically acceptable anion. Alkali metal (eg, sodium, potassium, or lithium) or alkaline earth metal (eg, calcium) salts of carboxylic acids may also be prepared.
[0184] method
[0185] Disclosed herein is a method of degrading a polynucleotide, comprising contacting the polynucleotide molecule with one or more polypeptides disclosed herein, thereby degrading the polynucleotide molecule.
[0186] Also disclosed herein are methods of degrading DNA or RNA during cell lysis. In some embodiments, the method comprises: lysing a host cell of interest; and adding a polypeptide disclosed herein under conditions that allow the polypeptide to degrade DNA or RNA.
[0187] Also disclosed herein are methods of degrading DNA or RNA during viral vector production. In some embodiments, the method comprises: culturing a host cell comprising a viral vector of interest; and expressing or adding one or more polypeptides disclosed herein under conditions that allow the one or more polypeptides to degrade DNA or RNA of non-viral vectors.
[0188] Also disclosed herein are methods of degrading polynucleotides during expression or production of a protein of interest. In some embodiments, the method comprises: culturing a host cell comprising a nucleic acid encoding a protein of interest; and expressing one or more nuclease polypeptides disclosed herein under conditions that allow at least one of the one or more nuclease polypeptides to degrade polynucleotides. The polynucleotide molecules can comprise DNA, RNA, or any combination thereof. In some embodiments, at least one of the one or more nuclease polypeptides is not expressed by the cell expressing the protein of interest. In some embodiments, the one or more nuclease polypeptides are expressed by a cell that does not express the protein of interest.
[0189] In these methods, the one or more nuclease polypeptides can be expressed from, for example, one or more expression vectors present in the host cell, or nucleic acid sequences in the host cell chromosome. Expression of the protein of interest and / or expression of the nuclease polypeptides can be inducible. For example, the coding sequence for the protein of interest, the coding sequence for the nuclease polypeptide, or both, can be operably linked to an inducible promoter. The inducible promoter can be induced, for example, by the presence, absence, and / or amount of one or more chemical or biological compounds, or a change in pH, temperature, osmotic pressure, ionic strength / concentration, or any combination thereof. Some embodiments provide a host cell comprising a nucleic acid encoding a protein of interest and a nucleic acid encoding one or more nuclease polypeptides disclosed herein.
[0190] Examples
[0191] Aspects of the above-described embodiments can be further described with reference to the following experimental examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the application. Therefore, the application should not be construed as being limited to the following examples, but rather the scope of the application includes all variations and modifications that are apparent to those of ordinary skill in the art in light of the teachings provided herein. Methods and reagents used in the examples are conventional methods and reagents in the art, unless otherwise stated.
[0192] Materials and Methods
[0193] Protein Engineering
[0194] Based on the crystal structure of Serratia marcescens nuclease A (PDB accession number: 1G8T), an electrostatic surface remodeling strategy was employed to improve the salt tolerance of NucA by introducing positively charged residues at specific sites on the protein surface. Eleven residues corresponding to S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156 in wild-type nuclease A were mutated to positively charged residues, either lysine or arginine. An additional 15 variants (B0 to B14) were designed to investigate the combined effects of positively charged residues at multiple sites. Single-site saturation mutation variants (B15 to B90) were also designed to investigate the contributions of different residues to high salt tolerance. The sequences of wild-type NucA, B0, B1, B2, B3, B4, and B5 are as follows:
[0195] Wild-type NucA (SEQ ID NO: 1):
[0196] MRFNNKMLALAALLFAAQASA DTLESIDNCAVGCPTGGSSNVSIVRHAYTLNNNSTTKFANWVAYHITKDTPA S GKTR N WKTDPALNPADTLAP A DY T GA NA ALKVDRGHQAPLASLAGVSDWESLNYLSNITPQK SD LN Q GAWARLE D QERKLI D RADISSVYTVTGPLYERDMGKLPGTQKAHTIPSAYWKVIFINNSPAVNHYAAFLFDQNTPKGADFCQFRVTVDEIEKRTGLIIWAGLPDDVQASLKSKPGVLPELMGCKN
[0197] B0:N79K,A95K,A102K,D149K;
[0198] B1:N79K,A95K,A102K,Q141K,D149K;
[0199] B2:N79K,A95K,N101K,A102K,Q141K,D149K;
[0200] B3:A95K,N101K,Q141K,D149K;
[0201] B4:S74K,N79K,A95K,T98K,N101K,Q141K,D149K;
[0202] B5:S74K,N79K,A95K,T98K,N101K,A102K,S137K,D138K,Q141K,D149K,D156R;
[0203] B6:N79R,A95R,A102R,D149R;
[0204] B7:A95K,A102K,D149K;
[0205] B8:N79K,A102K,D149K;
[0206] B9:N79K,A95K,D149K;
[0207] B10:N79K,A95K,A102K;
[0208] B11:A95R,A102R,D149R;
[0209] B12:N79R,A102R,D149R;
[0210] B13:N79R,A95R,D149R;
[0211] B14:N79R,A95R,A102R;
[0212] B15:N79K; B16:N79R; B17:N79A; B18:N79C; B19:N79D; B20:N79E; B21:N79F; B22:N79G; B23:N79H; B24:N79I; B25:N79L; B26:N79M; B27:N79P; B28:N79Q; B29:N79S; B30:N79T; B31:N79V; B32:N79W; B33:N79Y; B34:A95K; B35:A95R; B36:A95C; B37:A95D; B38:A95E; B39:A95F; B40:A95G; B41:A95H; B42:A95I; B43:A95L; B44:A95M; B45:A95N; B46:A95P; B47:A95Q; B48:A95S; B49:A95T; B50:A95V; B51:A95W; B52:A95Y; B53:A102K; B54:A102 R;B55:A102C;B56:A102D;B57:A102E;B58:A102F;B59:A102G;B60:A102H;B61:A102I;B62:A102L;B63:A102 M; B64:A102N; B65:A102P; B66:A102Q; B67:A102S; B68:A102T; B69:A102V; B70:A102W; B71:A102Y; B72:D149 K;B73:D149R;B74:D149A;B75:D149C;B76:D149E;B77:D149F;B78:D149G;B79:D149H;B80:D149I;B81:D149 L;B82:D149M;B83:D149N;B84:D149P;B85:D149Q;B86:D149S;B87:D149T;B88:D149V;B89:D149W;B90:D149Y
[0213] Plasmid preparation, protein expression and purification
[0214] Genes encoding nuclease A mutants were synthesized after codon optimization (Wuxi WuXi Biologics) and cloned into the pET-28a(+) vector (Novagen). The plasmids were transformed into Escherichia coli BL21(DE3) cells. The overexpressed proteins contained a C-terminal 6×histidine tag and could be purified using Ni-NTA resin (Cytiva). Positive clones were inoculated into 200 ml of LB medium and cultured at 37°C to an OD600 of 0.6-0.8. Protein expression was induced by the addition of isopropylthio-β-galactoside to a final concentration of 0.4 mM. The cultures were then transferred to 18°C and incubated for an additional 20 hours. Cell cultures were harvested by centrifugation at 3500 g for 20 minutes (4°C), and the cell pellets were resuspended in lysis buffer (20 mM Tris, 300 mM NaCl, 10 mM imidazole, 5% v / v glycerol, pH 8.0) and lysed by sonication. The lysate was centrifuged at 16000 g for 20 minutes (4° C.), and the supernatant was loaded onto a Ni-NTA column and washed twice with wash buffer (20 mM Tris, 300 mM NaCl, 40 mM imidazole, 5% v / v glycerol, pH 8.0). The target protein was eluted with elution buffer (20 mM Tris, 300 mM NaCl, 250 mM imidazole, 5% v / v glycerol, pH 8.0). The eluate was further purified by a HiLoad 16 / 600 Superdex 75 pg gel filtration column (Cytiva), and the protein was dialyzed into storage buffer (20 mM Tris-Cl, pH 8.0, 20 mM NaCl, 2 mM MgCl 2 ), concentrated to approximately 1 mg / ml, and stored with 50% glycerol.
[0215] Enzyme specificity assay
[0216] The enzyme specificity of wild-type nuclease A and its mutants was determined as follows: 4 U of enzyme was mixed with 1 μg of various substrates (including total RNA from CHO cells, plasmid DNA-pET28(a+), single-stranded DNA, and λ DNA) in buffers containing varying salt concentrations. The reaction was incubated at 37°C for 30 minutes and terminated by the addition of 10 mM EDTA. The samples were loaded onto 1% or 2% agarose gels and analyzed for digestion after staining with GelRed (Sigma-Aldrich).
[0217] Determination of residual enzyme activity
[0218] Salmon sperm DNA was dissolved in reaction buffer (50 mM Tris, 2 mM MgCl2, pH 8.0) to prepare a solution with a concentration of 1 mg / ml. The enzymatic activity of wild-type nuclease A and its mutants was determined as follows: 2 ng of each enzyme and 8 μl of salmon sperm DNA (1 mg / ml) were added to buffers containing different salt concentrations and pH values, incubated at 37°C for 30 minutes, and the reaction was terminated by adding 10 mM EDTA. The A value of each reaction was measured using a NanoDrop (Thermo Fisher). 260 Absorbance. Relative enzyme activity was calculated using the following formula:
[0219] Relative enzyme activity = (A260 样品 –A260 无酶 ) / (A260 无盐 –A260 无酶 )
[0220] A260 样品 : A260 absorbance of each reaction sample; A260 无酶 : A260 absorbance of a blank reaction without enzyme; A260 no salt: A260 absorbance of a control reaction containing only reaction buffer (50 mM Tris, 2 mM MgCl2, pH 8.0).
[0221] This method quantifies the residual activity of the enzyme under different salt concentrations and pH conditions by detecting changes in ultraviolet absorbance of nucleotides released after DNA degradation, thereby evaluating the salt tolerance and acid-base stability of the mutant.
[0222] Example 1
[0223] To investigate whether High Salt NucA, like wild-type NucA, is a nonspecific nuclease, we used High Salt NucA and wild-type NucA to digest various nucleic acid forms, including supercoiled plasmid DNA, single-stranded DNA, linear double-stranded DNA (λ phage genomic DNA), and total RNA, in buffers containing varying NaCl concentrations. In this and subsequent examples, "Mutant B0" is equivalent to "High Salt NucA," unless otherwise indicated.
[0224] HighSalt NucA can digest linear double-stranded DNA at high salt concentrations
[0225] The lambda phage genomic DNA (New England Biolabs, Cat. No. N3011S) was incubated with HighSalt NucA and wild-type nuclease A. The results showed that the digestion activity of wild-type nuclease A was significantly inhibited in the buffer containing >200mM NaCl ( Figure 2However, HighSalt NucA can still effectively digest λDNA in a buffer containing 500mM NaCl ( Figure 2 ).
[0226] HighSalt NucA can digest supercoiled plasmid DNA at high salt concentrations
[0227] Plasmid pET28a(+) (Novagen) was incubated with HighSalt NucA and wild-type nuclease A. The results were similar to those of the λ DNA digestion experiment: the digestion activity of wild-type nuclease A was significantly inhibited in buffers containing >200 mM NaCl, while HighSalt NucA was still able to effectively digest plasmid DNA at salt concentrations up to 500 mM ( Figure 3 ).
[0228] HighSalt NucA digests RNA at high salt concentrations
[0229] Total RNA was extracted from CHO cells (Wuxi WuXi Biologics Co., Ltd.) using the RNeasy Plus Mini Kit (QIAGEN, Cat. No. 74134) and incubated with HighSalt NucA and wild-type nuclease A, respectively. Ribosomal RNA is the most abundant component of total RNA, with different regions containing double-stranded and single-stranded portions. Digestion results showed that the activity of wild-type nuclease A was significantly inhibited when the salt concentration exceeded 200 mM ( Figure 4 ), while HighSalt NucA can still effectively digest RNA at salt concentrations up to 500mM ( Figure 4 ).
[0230] HighSalt NucA can digest single-stranded DNA at high salt concentrations
[0231] Single-stranded DNA (120-130 nucleotides) was synthesized and incubated with HighSalt NucA and wild-type nuclease A. The results showed that the digestion activity of wild-type nuclease A was significantly inhibited in buffers containing >200mM NaCl ( Figure 5 ), while HighSalt NucA can still effectively digest single-stranded DNA in a buffer containing 500mM NaCl ( Figure 5 ).
[0232] in conclusion
[0233] Wild-type Serratia marcescens nuclease A is a nonspecific nuclease that digests diverse nucleic acid forms. Enzyme specificity assays revealed that the engineered HighSalt NucA maintains nonspecific catalytic activity against diverse nucleic acid forms. Furthermore, HighSalt NucA exhibits significantly greater salt tolerance than wild-type nuclease A, maintaining activity at high salt concentrations such as 500 mM NaCl.
[0234] Example 2
[0235] To investigate the inhibitory effects of different salts on wild-type nuclease A and HighSalt NucA, salmon sperm DNA was incubated with wild-type nuclease A and HighSalt NucA in buffers containing different concentrations of monovalent salts (NaCl, KCl), divalent salts (MgCl2, MnCl2, (NH4)2SO4) and trivalent salts (Na2HPO4), and the residual enzyme activities were determined.
[0236] Inhibitory effect of monovalent salt on HighSalt NucA
[0237] Wild-type nuclease A is sensitive to both NaCl and KCl concentrations. When the monovalent salt concentration exceeds 300 mM, the wild-type enzyme activity drops significantly to below 60% ( Figure 6 However, the modified HighSalt NucA can still maintain >60% enzyme activity at 500 mM salt concentration ( Figure 6 ). Therefore, HighSalt NucA is significantly more tolerant to monovalent salts.
[0238] Inhibitory effect of divalent salts on HighSalt NucA
[0239] Divalent salts provide stronger ionic strength in solution. The experiment further explored the effect of divalent ions on nuclease A and HighSalt NucA. Similar to the performance of HighSalt NucA in monovalent salt buffer, the results showed that it could still maintain 60% of the enzyme activity at 400mM MgCl2, and its tolerance to MgCl2 was better than that of wild-type nuclease A ( Figure 7 In addition, HighSalt NucA was also slightly more tolerant to (NH4)2SO4 than the wild type ( Figure 8 ). However, it is worth noting that the inhibitory effect of MnCl2 on HighSalt NucA seems to be stronger than that on the wild type ( Figure 7 ).
[0240] Inhibitory effect of trivalent salts on HighSalt NucA
[0241] Due to the chelation effect, nucleases that rely on metal ions are sensitive to PO43-. In PO43- buffer, HighSaltNucA behaves similarly to wild-type nuclease A. When the PO43- concentration exceeds 100mM, the enzyme activity of both is significantly inhibited ( Figure 9 ).
[0242] in conclusion
[0243] The enzyme residual activity assay showed that HighSalt NucA tolerated higher concentrations of NaCl, KCl, and (NH4)2SO4 than wild-type nuclease A. Modification of nuclease A by introducing positively charged residues on its specific surface did improve its high salt tolerance.
[0244] Example 3
[0245] To investigate the contribution of single residue mutations at N79, A95, A102, and D149 to salt tolerance, each residue was mutated to 19 other amino acids, and the activities of each mutant were analyzed under different salt concentrations.
[0246] N79 mutation and its contribution to salt tolerance
[0247] The N79 residue was mutated to 19 other amino acids, and the mutants were purified and analyzed for their salt tolerance. The results showed that, except for the mutations of the positively charged residues K (lysine) and R (arginine), many other mutations (such as A, C, G, Q, S, T, V, W, and Y) showed better salt tolerance than the wild-type enzyme ( Figure 10 and Table 1 ).
[0248] Table 1. Salt tolerance of wild-type NucA, HighSalt NucA, and the N79 mutant
[0249]
[0250]
[0251] A95 mutation and its contribution to salt tolerance
[0252] The A95 residue was mutated to 19 other amino acids, and the mutants were purified and analyzed for their salt tolerance. The results showed that except for the mutation of the positively charged residues K (lysine) and R (arginine), other mutations (such as F, N, P, Q, S, T, V, Y) had a positive effect on improving the salt tolerance of the mutants ( Figure 11 and Table 2 ).
[0253] Table 2. Salt tolerance of wild-type NucA, HighSalt NucA, and A95 site single mutant
[0254]
[0255]
[0256] A102 mutation and its contribution to salt tolerance
[0257] The A102 residue was mutated to 19 other amino acids, and the mutants were purified and analyzed for their salt tolerance. Among these mutations, the mutants containing K, R, T, V, W, and Y had better salt tolerance than the wild-type enzyme ( Figure 12 and Table 3 ).
[0258] Table 3. Salt tolerance of wild-type NucA, HighSalt NucA, and A102 site single mutant
[0259]
[0260] D149 mutation and its contribution to salt tolerance
[0261] The D149 residue was mutated to 19 other amino acids, and the mutants were purified and analyzed for their salt tolerance. By mutating this site to most other residues to reduce the acidic surface, the salt tolerance of the enzyme can be improved. This site can be replaced by potential residues such as A, F, H, K, N, Q, R, S, T, V, W, and Y ( Figure 13 and Table 4 ).
[0262] Table 4. Salt tolerance of wild-type NucA, HighSalt NucA, and D149 mutants
[0263]
[0264] Effect of positively charged residue combination on salt tolerance
[0265] Based on the above results, single-site mutations (especially K and R mutations) can improve the salt tolerance of the enzyme, but their effect is limited. Therefore, we hypothesized that salt tolerance can be further enhanced by combining mutations. We selected N79, A95, A102, and D149 sites for combined mutations. The results showed that compared with the wild-type enzyme and single-site mutants, the salt tolerance of the combined mutants was significantly improved ( Figure 14 and Table 5). In addition, the results also verified another hypothesis: the enhancement of enzyme activity in high-salt buffer depends on the balance between substrate binding and release.
[0266] Table 5. Salt tolerance of wild-type NucA, HighSalt NucA, and mutants with positively charged residue combinations
[0267]
[0268]
[0269] in conclusion
[0270] In addition to mutations in the positively charged residues K (lysine) and R (arginine), mutations in other residues also improve salt tolerance. However, single-site mutations are not sufficient to achieve optimal performance, so multi-site mutations (especially multi-site K / R mutations) can significantly improve enzyme activity under high salt conditions.
Claims
1. A polypeptide derived from nuclease A of Serratia marcescens having high salt tolerance, wherein the polypeptide comprises one or more mutations that increase the positively charged surface area of the polypeptide in its three-dimensional structure. Preferably, compared to the nuclease having the sequence of SEQ ID NO: 1, the polypeptide still has at least 60% of the nuclease activity under the condition that the solution ionic strength is greater than 200 mM.
2. A polypeptide comprising: (a) an amino acid sequence that is at least 70% identical to SEQ ID NO: 1 or the mature polypeptide thereof, and has a mutation at one, two, three, four, five, six, seven or more positions, or all of the positions S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, or D156; or (b) an amino acid sequence having at least 70% sequence identity with the sequence described in (a), and having a mutation at the position described in (a) and retaining nuclease activity, Preferably, N79, A95, A102, and D149 are each independently mutated to a non-polar amino acid, such as alanine, valine, isoleucine, proline, phenylalanine, tryptophan, methionine, or glycine, or to an uncharged polar amino acid, such as serine, threonine, glutamine, tyrosine, or cysteine, or to a positively charged polar amino acid, such as histidine, lysine, or arginine. More preferably, S74, N79, A95, T98, N101, A102, S137, D138, Q141, D149, and D156 are mutated to positively charged amino acid residues, preferably to histidine, lysine, or arginine.
3. The polypeptide of claim 2, comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and having mutations at: (1) A95 and at least three positions selected from the group consisting of S74, N79, T98, N101, A102, S137, D138, Q141, D149, and D156; or (2) D149 and at least three positions selected from the group consisting of S74, N79, A95, T98, N101, A102, S137, D138, Q141, and D156; Preferably, the polypeptide has mutations at A95 and D149 and at least two sites selected from S74, N79, T98, N101, A102, S137, D138, Q141, D156; More preferably, the polypeptide has mutations at N79, A95, A102 and D149, and optionally at least one site selected from S74, T98, N101, S137, D138, Q141, D156; further preferably, the polypeptide has mutations at N79, A95, A102 and D149, and optionally at one or two sites selected from Q141 and N101; More preferably, the polypeptide has mutations at A95, N101, Q141 and D149, and optionally has mutations at at least 3 sites selected from S74, N79, T98, A102, S137, D138, D156; further preferably, the polypeptide has mutations at A95, N101, Q141, D149, and optionally has mutations at S74, N79 and T98.
4. The polypeptide of claim 2, comprising an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or the mature polypeptide thereof, and having mutations at at least three positions selected from N79, A95, A102, and D149, and optionally having a mutation at at least one position selected from S74, T98, N101, S137, D138, Q141, D156; Preferably, the polypeptide has a mutation at A95, A102 and D149 and optionally has a mutation at at least one site selected from S74, N79, T98, N101, S137, D138, Q141, D156; or the polypeptide has a mutation at N79, A102 and D149 and optionally has a mutation at at least one site selected from S74, A95, T98, N101, S137, D138, Q141, D156; Or the polypeptide has a mutation at N79, A95 and D149 and optionally has a mutation at at least one site selected from S74, T98, N101, A102, S137, D138, Q141, D156; or the polypeptide has a mutation at N79, A95 and A102 and optionally has a mutation at at least one site selected from S74, T98, N101, A102, S137, D138, Q141, D149, D156.
5. The polypeptide of claim 2, comprising an amino acid sequence having at least 70% sequence identity to SEQ ID NO: 1 or the mature polypeptide thereof, wherein the polypeptide has nuclease activity under high solution ionic strength, and the polypeptide comprises one or more mutations selected from the group consisting of: (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, (5) S74K, (6) T98K, (7) N101K, (8) S137K, (9) D138K, (10) Q141K, and (11) D156R, Preferably, the polypeptide has: (a) A95K or A95R mutation, and at least one selected from (1) N79K or N79R, (2) A102K or A102R, (3) D149K or D149R, (4) S74K, (5) T98K, (6) N101K, (7) S137K, (8) D138K, (9) Q141K and (10) D156R. 3 mutations; or (b) D149K or D149R mutation, and at least 3 mutations selected from (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) S74K, (5) T98K, (6) N101K, (7) S137K, (8) D138K, (9) Q141K, and (10) D156R, More preferably, the polypeptide has at least three mutations selected from (1) N79K or N79R, (2) A95K or A95R, (3) A102K or A102R, (4) D149K or D149R, and optionally has a mutation selected from S74K, T98K, N101K, S137K, D138K, Q141K and D156R, More preferably, the polypeptide comprises an amino acid sequence having at least 70% identity to SEQ ID NO: 1 or its mature polypeptide, and has A95K, N101K, Q141K, D149K mutations, and optionally has a mutation selected from S74K, N79K, T98K, A102K, S137K, D138K and D156R; further preferably, the polypeptide has A95K, N101K, Q141K, D149K mutations, and optionally has S74K, N79K and T98K mutations.
6. A composition or kit comprising one or more polypeptides according to any one of claims 1 to 3, Preferably, the composition is a reaction mixture, a detergent composition, a detergent additive, a food, a food supplement, a feed supplement, a feed, a pharmaceutical composition, a fermentation product, a fermentation intermediate, a fermentation downstream reaction mixture, or a combination thereof.
7. A nucleic acid comprising: (1) a sequence encoding one or more polypeptides according to any one of claims 1 to 3, or a complementary sequence thereof; or (2) A sequence that is at least 50%, 60%, 70%, 80% or 90% identical to (1).
8. A nucleic acid construct comprising the polynucleotide sequence of the nucleic acid according to claim 5, Preferably, the nucleic acid construct is a cloning vector, an expression vector or a recombinant vector. More preferably, the expression vector comprises a viral vector, a plasmid, a phage, a phagemid, a cosmid, a fosmid, a phage, an artificial chromosome or a combination thereof.
9. A reaction mixture, wherein the reaction mixture comprises: (a) one or more polypeptides as described in any one of claims 1 to 3, (b) one or more nucleic acid molecules, and (c) an aqueous solution in which the polypeptide hydrolyzes the one or more nucleic acid molecules.
10. A cell comprising one or more polypeptides according to any one of claims 1 to 3, one or more nucleic acids encoding the polypeptides, one or more nucleic acid constructs comprising the polynucleotide sequences of the nucleic acids, or a combination thereof.
11. A method for producing a polypeptide having nuclease activity under high solution ionic strength, comprising: Expressing a nucleic acid encoding the polypeptide according to any one of claims 1 to 3 under conditions allowing polypeptide expression, thereby producing a recombinant polypeptide having nuclease activity, wherein the nucleic acid is operably linked to a promoter, Preferably, the method comprises: culturing the cell of claim 7 under conditions allowing the cell to express the polypeptide.
12. A method for degrading a polynucleotide, comprising contacting a polynucleotide molecule with one or more polypeptides according to any one of claims 1 to 3, thereby degrading the polynucleotide molecule. Preferably, said contacting occurs at pH 4 to pH 11, Preferably, the contacting occurs at 10°C to 70°C.
13. A method for degrading DNA or RNA during cell lysis, comprising: Lyse target host cells; and contacting the lysate with one or more polypeptides according to any one of claims 1 to 3 under conditions that allow said polypeptides to degrade DNA or RNA, Preferably, the host cell is selected from the group consisting of bacterial cells, mammalian cells, fungal cells, yeast cells, plant cells or insect cells.
14. A method for degrading DNA or RNA during viral vector production, comprising: Cultivating a host cell, wherein the host cell contains a target viral vector; and expressing or adding one or more of the polypeptides according to any one of claims 1 to 3 under conditions that allow the polypeptide to degrade DNA or RNA other than the viral vector, Preferably, the host cell is selected from the group consisting of bacterial cells, mammalian cells, fungal cells, yeast cells, plant cells or insect cells.
15. A method for degrading DNA or RNA during protein production, comprising: Cultivating a host cell, wherein the host cell comprises a nucleic acid encoding a protein of interest; and expressing or adding one or more polypeptides disclosed herein under conditions that allow the polypeptides to degrade DNA or RNA, Preferably, the host cell is selected from the group consisting of bacterial cells, mammalian cells, fungal cells, yeast cells, plant cells or insect cells.
16. A method for degrading DNA or RNA in a protein production mixture, comprising: Cultivating host cells containing nucleic acid encoding a protein of interest; and expressing or adding one or more of the polypeptides according to any one of claims 1 to 3 under conditions that allow the polypeptide to degrade DNA or RNA, Preferably, the host cell is selected from the group consisting of bacterial cells, mammalian cells, fungal cells, yeast cells, plant cells or insect cells.
Citation Information
Cited By
PPR (pentatricopeptide repeats) high-salt-resistant all-potent nuclease mutant and application thereof
CN121610472A
Ppr high-salt-resistant omnipotent nuclease mutant and application thereof
CN121610472B