Method for preparing base mismatched specific binding protein and related biological material and application thereof

By preparing base mismatch-specific binding proteins, using fusion proteins of endonuclease and maltose binding protein domains, the problem of high base error rate in artificially synthesized DNA is solved, and efficient and low-cost DNA error correction and detection applications are achieved.

CN120289647APending Publication Date: 2025-07-11TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410030164.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has a high base error rate in the process of artificial synthesis of DNA, which is difficult to effectively correct, affecting the quality of high-throughput DNA synthesis.

Method used

Base mismatch-specific binding proteins are prepared, and fusion proteins expressing base mismatch-specific endonuclease and maltose binding protein domains are produced using a prokaryotic expression system, and purified by affinity chromatography of specific matrix vectors, and applied to DNA error correction.

Benefits of technology

It significantly reduces the base error rate of artificially synthesized DNA, improves the accuracy and efficiency of DNA synthesis, is suitable for automated operations, reduces production costs, and is used for artificially synthesized DNA error correction, disease-related gene mutation detection, and animal and plant genetic breeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004655823020000101
    Figure BDA0004655823020000101
  • Figure HDA0004655823030000011
    Figure HDA0004655823030000011
  • Figure HDA0004655823030000012
    Figure HDA0004655823030000012
Patent Text Reader

Abstract

The invention discloses a method for preparing a base mismatched specific binding protein as well as a related biological material and application thereof. The invention provides a method for preparing a base mismatched specific binding protein, which comprises the following step: expressing a coding gene of the base mismatched specific binding protein in organisms to obtain the base mismatched specific binding protein, the organisms are microorganisms, plants or non-human animals; the base mismatch specific binding protein comprises a base mismatch specific endonuclease and a structural domain which contributes to the affinity adsorption of the zymoprotein to the matrix carrier or promotes the expression of the zymoprotein. The base mismatch specific endonuclease is a CEL I zymoprotein or a CEL II zymoprotein of apios graveolens, and the structural domain is an MBP protein. The fusion protein disclosed by the invention is applied to methods for detecting DNA base errors or mutation, correcting base errors in a gene synthesis process and the like, and has the advantages of high efficiency, rapidness, convenience in automatic operation, low cost and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and particularly relates to a method for preparing a base mismatch-specific binding protein, and related biological materials and applications thereof. Background Art

[0002] The rapid development of modern biotechnology such as synthetic biology and metabolic engineering has greatly increased the demand for artificially synthesized DNA. At present, the chemical method for synthesizing DNA mainly uses the phosphoramidite method for solid-phase synthesis. Due to factors such as the control technology of the synthesizer, the synthesis scale, and the quality of reagents, there is still a certain error rate (such as <1%) in each step of the synthesis reaction, resulting in possible errors in the finally synthesized DNA. For high-throughput DNA synthesis, the problem of a relatively high error rate is more obvious. Developing efficient error correction technology has become an important part of high-efficiency high-throughput DNA synthesis technology. It has been found that in many organisms, there are some proteins that can specifically bind to base mismatches (base mismatch-specific binding proteins, MSB), such as MutS proteins derived from Escherichia coli, Haemophilus influenzae, Methanosaeta thermophila, etc. Through the binding action of base mismatch-specific binding proteins and appropriate separation means, the base mismatches formed during gene synthesis or replication can be distinguished from the correct fragments, thereby achieving the error correction of DNA base errors in the overall gene molecular library. Base mismatch-specific binding proteins have important applications in fields such as error correction of artificially synthesized DNA bases, detection of genetic disease gene mutations, and localization of mutant genes in animal and plant genetic breeding. Summary of the Invention

[0003] The technical problem to be solved by the present invention is the efficient error correction of artificially synthesized DNA bases.

[0004] To solve the above technical problems, the present invention provides a method for preparing a novel base mismatch-specific binding protein and applications of related biological materials.

[0005] In the first aspect, the present invention claims to protect a method for preparing a base mismatch-specific binding protein.

[0006] The method for preparing a base mismatch-specific binding protein claimed in the present invention may include: expressing the coding gene of the base mismatch-specific binding protein in an organism to obtain the base mismatch-specific binding protein; the organism being a microorganism, a plant or a non-human animal. The base mismatch-specific binding protein comprises a base mismatch-specific endonuclease (Mismatch specific endonuclease, MSE) and a domain that helps the enzyme protein to adsorb affinity to a matrix carrier or promotes the expression of the enzyme protein.

[0007] In the above method, the base mismatch-specific endonuclease is the CELI enzyme protein or the CEL II enzyme protein of celery (Apium graveolens); the domain that helps the enzyme protein to adsorb affinity to a matrix carrier or promotes the expression of the enzyme protein is the MBP protein. The MBP protein is a maltose binding protein. The MBP protein can be a natural protein or a recombinant protein, such as the MBP protein derived from Escherichia coli.

[0008] Furthermore, the MBP protein can be any one of the following:

[0009] (A1) A protein whose amino acid sequence is positions 1-388 of SEQ ID No.2;

[0010] (A2) A protein obtained by substituting and / or deleting and / or adding amino acid residues to the protein defined in (A1) and having more than 75% identity with the protein defined in (A1) and having the same function.

[0011] Furthermore, the CEL I enzyme protein can be any one of the following:

[0012] (B1) A protein whose amino acid sequence is positions 408-702 of SEQ ID No.2;

[0013] (B2) A protein obtained by substituting and / or deleting and / or adding amino acid residues to the protein defined in (B1) and having more than 75% identity with the protein defined in (B1) and having the same function.

[0014] Furthermore, the CEL II enzyme protein can be any one of the following:

[0015] (C1) A protein whose amino acid sequence is positions 408-716 of SEQ ID No.4;

[0016] A protein which is obtained by substitution and / or deletion and / or addition of amino acid residues to the protein defined in (C1) and has an identity of more than 75% with the protein defined in (C1) and has the same function.

[0017] Furthermore, the base mismatch-specific binding protein may be any one of the following:

[0018] (D1) A protein whose amino acid sequence is SEQ ID No.2 (corresponding to MSE32 in the examples) or SEQ ID No.4 (corresponding to MSE42 in the examples);

[0019] (D2) A fusion protein obtained by fusing a protein tag to the carboxyl terminus and / or amino terminus of the protein defined in (D1);

[0020] (D3) A protein which is obtained by substitution and / or deletion and / or addition of amino acid residues to the protein defined in (D1) or (D2) and has an identity of more than 75% with the protein defined in (D1) or (D2) and has the same function.

[0021] In SEQ ID No.2, positions 389-407 are the amino acid sequence encoded by the nucleotide sequence from pET-28a(+) (including a His-Tag) and an amino acid sequence linked to the CELI gene, positions 408-702 are the amino acid sequence of the CEL I enzyme protein derived from celery (Apium graveolens), and positions 1-388 are the amino acid sequence of the Escherichia coli MBP protein.

[0022] In SEQ ID No.4, positions 389-407 are the amino acid sequence encoded by the nucleotide sequence from pET-28a(+) (including a His-Tag) and an amino acid sequence linked to the CELII gene, positions 408-716 are the amino acid sequence of the CEL II enzyme protein derived from celery (Apium graveolens), and positions 1-388 are the amino acid sequence of the Escherichia coli MBP protein.

[0023] The above-mentioned base mismatch-specific binding protein can be artificially synthesized, or its encoding gene can be synthesized first and then obtained by biological expression.

[0024] Among the above base mismatch-specific binding proteins, the protein tag refers to a polypeptide or protein that is fused and expressed with the target protein by using in vitro DNA recombination technology for facilitating the expression, detection, tracing, and / or purification of the target protein. The protein tag can be a Flag tag, His tag, MBP tag, HA tag, myc tag, GST tag, SUMO tag, AviTag tag, eGFP (enhanced green fluorescent protein), eCFP (enhanced cyan fluorescent protein), eYFP (enhanced yellow-green fluorescent protein), and / or mCherry (monomeric red fluorescent protein).

[0025] The base mismatch-specific binding protein has the function of correcting artificial DNA base errors through base mismatch-specific binding.

[0026] In the above method, the organism can be the above-mentioned microorganism. Correspondingly, the expression includes introducing the coding gene of the base mismatch-specific binding protein into a recipient microorganism to obtain a recombinant microorganism expressing the coding gene of the base mismatch-specific binding protein, culturing the recombinant microorganism, and expressing to obtain the base mismatch-specific binding protein.

[0027] In the above method, the recipient microorganism can be any one of the following: prokaryotic microorganism; Gram-negative bacterium; bacterium of the genus Escherichia; Escherichia coli; Escherichia coli BL21(DE3).

[0028] In the above method, the recombinant microorganism may be recombinant Escherichia coli. In one embodiment of the present invention, the recombinant Escherichia coli is a recombinant microorganism obtained by introducing pMSE32 into Escherichia coli BL21(DE3) and expressing the base mismatch-specific binding protein with the amino acid sequence of SEQ ID No.2, and the recombinant microorganism is named BL21(DE3) / pMSE32. pMSE32 is a recombinant expression vector obtained by replacing the fragment between the upstream homologous arm sequence (nucleotide sequence positions 5021-5070 of SEQ ID No.1) and the downstream homologous arm sequence (nucleotide sequence positions 7180-7229 of SEQ ID No.1) of pET-28a(+) with the fragment of nucleotide sequence positions 5071-7179 of SEQ ID No.1 while keeping other nucleotides unchanged. In another embodiment of the present invention, the recombinant Escherichia coli is a recombinant microorganism obtained by introducing pMSE42 into Escherichia coli BL21(DE3) and expressing the fusion protein with the amino acid sequence of SEQ ID No.4, and the recombinant microorganism is named BL21(DE3) / pMSE42. pMSE42 is a recombinant expression vector obtained by replacing the fragment between the upstream homologous arm sequence (nucleotide sequence positions 5021-5070 of SEQ ID No.3) and the downstream homologous arm sequence (nucleotide sequence positions 7222-7271 of SEQ ID No.3) of pET-28a(+) with the fragment of nucleotide sequence positions 5071-7221 of SEQ ID No.3 while keeping other nucleotides unchanged.

[0029] In a second aspect, the present invention claims to protect a base mismatch-specific binding protein.

[0030] The base mismatch-specific binding protein claimed to be protected by the present invention comprises a base mismatch-specific endonuclease (Mismatch specific endonuclease, MSE) and a domain that helps the enzyme protein to adsorb affinity to a specific matrix carrier or promotes the expression of the enzyme protein. Specifically, it may be the base mismatch-specific binding protein described in the first aspect above.

[0031] The present invention uses the base mismatch-specific binding protein to perform DNA error correction (Enzymatic mismatch correction, EMC) through base mismatch-specific binding.

[0032] In a third aspect, the present invention claims to protect a biological material related to the base mismatch-specific binding protein described in the second aspect above.

[0033] The biological material may specifically be any one of the following:

[0034] (E1) A nucleic acid molecule encoding the base mismatch-specific binding protein described in the second aspect of the foregoing text;

[0035] (E2) An expression cassette containing the nucleic acid molecule described in (E1);

[0036] (E3) A recombinant vector containing the nucleic acid molecule described in (E1);

[0037] (E4) A recombinant vector containing the expression cassette described in (E2);

[0038] (E5) A recombinant microorganism containing the nucleic acid molecule described in (E1);

[0039] (E6) A recombinant microorganism containing the expression cassette described in (E2);

[0040] (E7) A recombinant microorganism containing the recombinant vector described in (E3);

[0041] (E8) A recombinant microorganism containing the recombinant vector described in (E4);

[0042] (E9) A transgenic plant cell line containing the nucleic acid molecule described in (E1) or a transgenic plant cell line containing the expression cassette described in (E2);

[0043] (E10) A transgenic plant tissue containing the nucleic acid molecule described in (E1) or a transgenic plant tissue containing the expression cassette described in (E2);

[0044] (E11) A transgenic plant organ containing the nucleic acid molecule described in (E1) or a transgenic plant organ containing the expression cassette described in (E2);

[0045] (E12) A transgenic animal cell line containing the nucleic acid molecule described in (E1) or a transgenic animal cell line containing the expression cassette described in (E2);

[0046] (E13) A transgenic animal tissue containing the nucleic acid molecule described in (E1) or a transgenic animal tissue containing the expression cassette described in (E2);

[0047] (E14) A transgenic animal organ containing the nucleic acid molecule described in (E1) or a transgenic animal organ containing the expression cassette described in (E2).

[0048] Among the above biological materials, the nucleic acid molecule may be DNA, such as cDNA, genomic DNA or recombinant DNA; the nucleic acid molecule may also be RNA.

[0049] Specifically, in (E1), the nucleic acid molecule is any one of the following:

[0050] (e1) The coding sequence is a nucleic acid molecule from positions 5071 to 7179 of SEQ ID No. 1 (corresponding to the coding gene of MSE32 in the examples) or from positions 5071 to 7222 of SEQ ID No. 3 (corresponding to the coding gene of MSE42 in the examples);

[0051] (e2) A nucleic acid molecule having more than 75% identity with the DNA molecule defined in (e1) and encoding the base mismatch specific binding protein.

[0052] Among the above-mentioned biological materials, the expression cassette refers to DNA that can express the base mismatch-specific binding protein in a host cell. The expression cassette may also include a single-stranded or double-stranded nucleic acid molecule containing all the regulatory sequences necessary for the nucleic acid molecule expressing any one of the above proteins. The regulatory sequences can direct the coding sequence to express any one of the above proteins in a suitable host cell under their compatible conditions. The regulatory sequences include, but are not limited to, leader sequences, polyadenylation sequences, propeptide sequences, promoters, signal sequences, and transcription terminators. At a minimum, the regulatory sequences should include a promoter and transcription and translation termination signals. To introduce specific restriction enzyme sites into the vector for ligating the regulatory sequences to the coding region of the nucleic acid sequence encoding the protein, regulatory sequences with linkers can be provided. The regulatory sequence can be a promoter sequence, i.e., a nucleic acid sequence recognizable by the host cell expressing the nucleic acid sequence. The promoter sequence contains transcriptional regulatory sequences that mediate protein expression. The promoter can be any nucleic acid sequence having transcriptional activity in the selected host cell, including mutant, truncated, and hybrid promoters, and can be derived from genes encoding extracellular or intracellular proteins homologous or heterologous to the host cell. The regulatory sequence can also be a transcription termination sequence, i.e., a sequence that can be recognized by the host cell to terminate transcription. The termination sequence is operably linked to the 3'-end of the nucleic acid sequence encoding the protein. Any terminator functional in the selected host cell can be used in the present invention. The regulatory sequence can also be a leader sequence, i.e., an untranslated region of mRNA that is important for translation in the host cell. The leader sequence is operably linked to the 5'-end of the nucleic acid sequence encoding the protein. Any leader sequence functional in the selected host cell can be used in the present invention. The regulatory sequence can also be a signal peptide coding region, which encodes an amino acid sequence linked to the amino terminus of the protein and can direct the encoded protein into the cell secretion pathway. Any signal peptide coding region that can direct the expressed protein into the secretion pathway of the host cell used can be used in the present invention. It may also be necessary to add regulatory sequences that can regulate protein expression according to the growth of the host cell. Examples of regulatory systems are those that can respond to chemical or physical stimuli (including in the presence of regulatory compounds) to turn gene expression on or off. Other examples of regulatory sequences are those that can amplify genes. In these examples, the nucleic acid sequence encoding the protein should be operably linked to the regulatory sequence.

[0053] The present invention also relates to a recombinant expression vector comprising a nucleic acid molecule encoding any one of the above-mentioned proteins, a promoter, and transcriptional and translational termination signals of the present invention. When preparing the expression vector, the nucleic acid molecule encoding any one of the above-mentioned base mismatch-specific binding proteins can be located in the vector so as to be operably linked to an appropriate expression regulatory sequence. The recombinant expression vector can be any vector that facilitates recombinant DNA manipulation and expression of the nucleic acid sequence (such as a plasmid or a virus). The choice of the vector usually depends on the compatibility of the vector with the host cell into which it is to be introduced. The vector can be a linear or closed circular plasmid. The vector can be an autonomously replicating vector (i.e., a complete structure existing outside the chromosome and capable of replicating independently of the chromosome), such as a plasmid, an episome, a minichromosome, or an artificial chromosome. The vector can contain any mechanism that ensures self-replication. Alternatively, the vector is a vector that will integrate into the genome when introduced into the host cell and replicate together with the chromosome into which it is integrated. In addition, a single vector or plasmid, or two or more vectors or plasmids or transposons that collectively contain all the DNA to be introduced into the host cell genome can be applied. The vector contains one or more selectable markers that facilitate the selection of transformed cells. A selectable marker is a gene whose product confers resistance to biocides or viruses, resistance to heavy metals, or prototrophy to auxotrophs, etc. Examples of bacterial selectable markers are the dal gene of Bacillus subtilis or Bacillus licheniformis, or resistance markers to antibiotics such as ampicillin, kanamycin, chloramphenicol, or tetracycline. The vector contains elements that enable the vector to be stably integrated into the host cell genome or ensure autonomous replication of the vector independently of the cell genome in the cell. In the case of autonomous replication, the vector can also contain an origin of replication that enables the vector to replicate autonomously in the target host cell. The origin of replication can carry a mutation that makes it temperature-sensitive in the host cell (see, for example, fEhrlich, 1978, Proceedings of the National Academy of Sciences of the United States of America 75: 1433). One or more copies of the nucleic acid molecule encoding any one of the above-mentioned proteins of the present invention can be inserted into the host cell to increase the yield of the gene product. The increase in the copy number of the nucleic acid molecule can be achieved by inserting at least one additional copy of the nucleic acid molecule into the host cell genome, or by inserting an amplifiable selectable marker together with the nucleic acid molecule, and by culturing the cells in the presence of a suitable selection reagent to select the cells containing the amplified copy of the selectable marker gene and thus containing an additional copy of the nucleic acid molecule. The operations for ligating the above-mentioned elements to construct the recombinant expression vector of the present invention are well known to those skilled in the art (see, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, 1989).

[0054] The term "operably linked" is defined herein as a conformation in which the regulatory sequence is in an appropriate position relative to the coding sequence of the DNA sequence so that the regulatory sequence directs the expression of the protein.

[0055] The vectors described in this article are well-known to those skilled in the art, including but not limited to: plasmids, phages (such as λ phage or M13 filamentous phage, etc.), cosmids (i.e., cosmid plasmids), Ti plasmids or viral vectors.

[0056] Among the above biological materials, the microorganism can be any one of the following: prokaryotic microorganisms; Gram-negative bacteria; Escherichia bacteria; Escherichia coli; Escherichia coli BL21(DE3).

[0057] In this article, the term "identity" refers to the sequence similarity with nucleic acid sequences or amino acids. Identity can be evaluated by the naked eye or computer software. Using computer software, the identity between two or more sequences can be expressed as a percentage (%), which can be used to evaluate the identity between related sequences.

[0058] In this article, an identity of more than 75% can be an identity of 80%, 85%, 90% or more than 95%. The identity of more than 80% can be at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%.

[0059] Those of ordinary skill in the art can easily use known methods, such as directed evolution and point mutation methods, to mutate the protein of a) of the present invention or its coding gene. Those nucleotides that have been artificially modified and have an identity of 75% or higher with the nucleotide sequence of the protein of a) described in the present invention, as long as they encode the DNA error correction enzyme protein and have the DNA error correction enzyme function, are all derived from the nucleotide sequence of the present invention and are equivalent to the sequence of the present invention. The present invention also provides a method for preparing a base mismatch-specific binding protein.

[0060] Fourthly, the present invention also claims any one of the following applications:

[0061] (F1) The application of the method described in the first aspect above in the preparation of a DNA error correction enzyme preparation;

[0062] (F2) The application of the MBP protein described in the first aspect above in enhancing the DNA error correction efficiency of a base mismatch-specific endonuclease or the application in the preparation of a product for enhancing the DNA error correction efficiency of a base mismatch-specific endonuclease;

[0063] (F3) The application of the nucleic acid molecule encoding the MBP protein described in the first aspect above in the preparation of a product for enhancing the DNA error correction efficiency of a base mismatch-specific endonuclease;

[0064] (F4) Use of the base mismatch-specific binding protein described in the second aspect above in DNA error correction, or use in the preparation of a DNA error correction enzyme preparation;

[0065] (F5) Use of the biological material described in the third aspect above in the preparation of a DNA error correction enzyme preparation;

[0066] (F6) Use of the base mismatch-specific binding protein described in the second aspect above in recognizing and binding to DNA mismatches, or use in the preparation of a preparation for recognizing and binding to DNA mismatches;

[0067] (F7) Use of the biological material described in the third aspect above in the preparation of a preparation for recognizing and binding to DNA mismatches.

[0068] Experimental results show that the base mismatch-specific binding proteins MSE32 and MSE42 of the present invention have the function of correcting artificial DNA base errors through base mismatch-specific binding, and are novel base mismatch-specific binding proteins. After error correction treatment, the base errors can be reduced to 8 / 10KB.

[0069] The base mismatch-specific binding proteins MSE32 and MSE42 disclosed in the present invention are expressed and produced using a prokaryotic expression system, and can be conveniently purified by affinity chromatography on a specific matrix carrier, which can reduce production costs such as expression and purification. At the same time, the base mismatch-specific binding proteins MSE32 and MSE42 disclosed in the present invention can be applied to the error correction of base errors in the process of artificial DNA synthesis, the detection of DNA base errors or mutations, etc., and have the advantages of high efficiency, convenience for automated operation, low cost, etc. It can be used in research and application fields including but not limited to the error correction of artificial DNA base errors, the early screening and detection of gene mutations related to important diseases, and the localization of mutant genes in animal and plant genetic breeding. Brief Description of the Drawings

[0070] Figure 1 It is the medium-scale induced expression electrophoresis pattern of MSE32 / MSE42 protein. In the left figure, lanes 1 and 2 are the purified flow-through fractions of MSE32 protein, and lanes 3, 4, and 5 are the purified elution fractions of MSE32 protein. In the right figure, lanes 1, 2, 3, and 4 are the purified flow-through fractions of MSE42 protein, and lane 5 is the purified elution fraction of MSE42 protein. In the figure, M represents the protein molecular weight Marker; among them, the arrow indicates the target band (base mismatch-specific binding protein).

[0071] Figure 2 It is the principle and experimental process of MSE32 / MSE42 for DNA base error correction.

[0072] Figure 3Electrophoresis detection of the products after the test samples were treated with MSE32 / 42 enzyme protein EMC (A) and EMC-PCR (B). In A, lanes 1 and 2 were AmpAN (15 μl) treated with 3 μl of MSE32 enzyme protein added to a 20 μl EMC treatment reaction system; lanes 3 and 4 were AmpAN (15 μl) treated with 2 μl of MSE32 enzyme protein added to a 20 μl EMC treatment reaction system; lanes 5 and 6 were AmpAN (15 μl) treated with 3 μl of MSE42 enzyme protein added to a 20 μl EMC treatment reaction system; lanes 7 and 8 were AmpAN (15 μl) treated with 2 μl of MSE42 enzyme protein added to a 20 μl EMC treatment reaction system. In B, lane 1 was the untreated AmpAN sample; lane 2 was the AmpMU sample; lane 3 was the AmpWT sample; lanes 4 and 5 were AmpAN (15 μl) treated with 2 μl of MSE32 enzyme protein added to a 20 μl EMC-PCR treatment reaction system; lanes 6 and 7 were AmpAN (15 μl) treated with 2 μl of MSE42 enzyme protein added to a 20 μl EMC-PCR treatment reaction system; lane M was a 1Kb DNA ladder DNA marker (ThermoFisher, USA). Detailed implementation manners

[0073] The present invention will be further described in detail below in conjunction with the specific implementation manners. The examples given are only for clarifying the present invention, rather than limiting the scope of the present invention.

[0074] The experimental methods in the following examples are all conventional methods unless otherwise specified.

[0075] The materials, reagents, etc. used in the following examples can be obtained from commercial channels unless otherwise specified.

[0076] In the embodiments of the present invention, the conventional instruments and consumables include a constant temperature shaking incubator, a sterile operating table, an induction cooker, a protein electrophoresis tank (XCell SureLock Mini-Cell, Invitrogen), a decolorizing shaker, a gel imaging system, a disposable SDS-PAGE gel (Invitrogen Precast gel), test tubes, 50 mL and 1 L conical flasks, 12-layer gauze, etc. The common culture media include LB liquid culture medium, LB solid plate culture medium (the formula is shown in "Molecular Cloning Experiment Guide"). The conventional reagents and drugs include 1000× kanamycin solution (50 mg / L), 1 M IPTG solution, 4× NuPAGE LDS loading buffer, NuPAGE Tris-acetate SDS electrophoresis buffer, DNA and protein molecular weight standards (Blue V ProteinMarker (10 - 190 kDa, Transgen), Coomassie Brilliant Blue staining solution (formula: dissolve 0.25 g of Coomassie Brilliant Blue R - 250 in 100 mL of methanol:acetic acid solution (decolorizing solution), filter through filter paper to remove particulate matter), decolorizing solution (formula: 500 mL of methanol, add 400 mL of water, mix well and then mix with 100 mL of glacial acetic acid), binding buffer (formula: 20 mM Na3PO4, 500 mM NaCl, pH 6.0, the rest is water), elution buffer (formula: add imidazole with a final concentration of 500 mM to the binding buffer, the rest is water), 20% ethanol solution, ultrapure water, etc.

[0077] Example 1, Production, Expression and Purification of Base Mismatch - Specific Binding Proteins MSE32 and MSE42

[0078] I. Preparation of BL21(DE3) / pMSE32 and BL21(DE3) / pMSE42

[0079] Using the CEL I gene and CEL II gene sequences of celery (Apium graveolens) in the EBI - ENA database as templates, relevant software for gene synthesis sequence design was used to design the sequences of gene synthesis primers. The gene sequence of the base mismatch - specific endonuclease part encoding the base mismatch - specific binding proteins MSE32 / MSE42 of the present invention was obtained by the method of fusion PCR splicing.

[0080] Using the malE gene sequence (MBP module) from Escherichia coli retrieved from the Genebank database as a template, after sequence optimization, relevant software for gene synthesis sequence design was used to design the sequences of gene synthesis primers. The gene sequence of the domain part encoding the base mismatch - specific binding proteins MSE32 / MSE42 of the present invention, which helps the enzyme protein to adsorb affinity to a specific matrix carrier or promotes the expression of the enzyme protein, was obtained by the method of fusion PCR splicing.

[0081] On this basis, CEL I / CEL II and MBP were fused into an expression cassette by fusion PCR splicing to obtain the expression cassette of the base mismatch-specific binding protein MSE32 / MSE42 of the present invention. The nucleotide sequence of the expression cassette of MSE32 is from position 5071 to 7179 of SEQ ID No.1. Among them, positions 5021-5070 and 7180-7229 of SEQ ID No.1 are the upstream homologous arm sequence and downstream homologous arm sequence of pET-28a(+), respectively. The nucleotide sequence of the expression cassette of MSE42 is from position 5071 to 7221 of SEQ ID No.3. Among them, positions 5021-5070 and 7222-7271 of SEQ ID No.3 are the upstream homologous arm sequence and downstream homologous arm sequence of pET-28a(+), respectively.

[0082] By using cloning methods such as Gibson assembly, the expression cassettes of the base mismatch-specific binding proteins MSE32 and MSE42 were inserted into the pET28a(+) expression vector respectively and placed under the transcriptional regulation of the T7 promoter and terminator, and the recombinant vectors pMSE32 and pMSE42 for expressing the base mismatch-specific binding proteins MSE32 and MSE42 of the present invention were obtained respectively. The recombinant expression vector pMSE32 is a recombinant plasmid obtained by replacing the fragment between the upstream homologous arm sequence (nucleotide sequence is positions 5021-5070 of SEQ ID No.1) and the downstream homologous arm sequence (nucleotide sequence is positions 7180-7229 of SEQ ID No.1) of pET-28a(+) with positions 5071-7179 of SEQ ID No.1, while keeping other nucleotides unchanged. The recombinant expression vector pMSE42 is a recombinant plasmid obtained by replacing the fragment between the upstream homologous arm sequence (nucleotide sequence is positions 5021-5070 of SEQ ID No.3) and the downstream homologous arm sequence (nucleotide sequence is positions 7222-7271 of SEQ ID No.3) of pET-28a(+) with positions 5071-7221 of SEQ ID No.3, while keeping other nucleotides unchanged.

[0083] 1) The recombinant expression vector pMSE32 expresses the base mismatch-specific binding protein MSE32 with the amino acid sequence of SEQ ID No.2. The full sequence of the recombinant expression vector pMSE32 is as shown in SEQ ID No.1, which contains the MSE32 gene with a His tag (positions 5071-7179 of SEQ ID No.1), and positions 5071-7179 of SEQ ID No.1 encode the protein MSE32 shown in SEQ ID No.2.

[0084] In SEQ ID No.1, nucleotides 1-5070, 6235-6264, and 7180-7315 are nucleotide sequences from pET-28a(+), nucleotides 6265-6291 are a sequence linked to the CELI gene, nucleotides 6292-7179 are the gene sequence of CEL I derived from celery (Apium graveolens), and nucleotides 5071-6234 are the MBP gene sequence optimized for E. coli codons.

[0085] In SEQ ID No.2, amino acids 389-407 are the amino acid sequence encoded by the nucleotide sequence from pET-28a(+) (including a His-Tag) and an amino acid sequence linked to the CELI gene, amino acids 408-702 are the amino acid sequence of CEL I derived from celery (Apium graveolens), and amino acids 1-388 are the amino acid sequence of E. coli MBP.

[0086] The recombinant expression vector pMSE32 was transformed into competent E. coli BL21(DE3) cells. The cells were evenly spread on an LB plate containing kanamycin and cultured at 37 °C for 16 hours. Single colonies were cultured overnight with shaking, and plasmids were extracted and sequenced to verify the correct positive clones. The recombinant E. coli containing pMSE32 was named BL21(DE3) / pMSE32. The amino acid sequence expressed by BL21(DE3) / pMSE32 is the base mismatch-specific binding protein MSE32 of SEQ ID No.2.

[0087] 2) The amino acid sequence expressed by the recombinant expression vector pMSE42 is the base mismatch-specific binding protein MSE42 of SEQ ID No.4. The full sequence of the recombinant expression vector pMSE42 is as shown in SEQ ID No.3, which contains the MSE42 gene with a His tag (nucleotides 5071-7221 of SEQ ID No.3). Nucleotides 5071-7221 of SEQ ID No.3 encode the protein MSE42 shown in SEQ ID No.4.

[0088] In SEQ ID No.3, nucleotides 1-5070, 6235-6264, and 7222-7357 are nucleotide sequences from pET-28a(+), nucleotides 6265-6291 are a sequence linked to the CELII gene, nucleotides 6292-7221 are the gene sequence of CEL II derived from celery (Apium graveolens), and nucleotides 5071-6234 are the MBP gene sequence optimized for E. coli codons.

[0089] In SEQ ID No.4, amino acid sequences encoded by nucleotide sequences from pET-28a(+) (including a His-Tag) and an amino acid sequence linked to the CELII gene are located at positions 389-407, the amino acid sequence of CEL II derived from Apium graveolens is located at positions 408-716, and the amino acid sequence of Escherichia coli MBP is located at positions 1-388.

[0090] The recombinant expression vector pMSE42 was transformed into Escherichia coli BL21(DE3) competent cells. The cells were evenly spread on an LB plate containing kanamycin and cultured at 37°C for 16 hours. Single colonies were cultured overnight with shaking, and plasmids were extracted and sequenced to verify the correct positive clones. The recombinant Escherichia coli containing pMSE42 was named BL21(DE3) / pMSE42. The amino acid sequence expressed by BL21(DE3) / pMSE42 is the base mismatch-specific binding protein MSE42 of SEQ ID No.4.

[0091] II. Purification of Base Mismatch-Specific Binding Proteins MSE32 and MSE42

[0092] In this example, the production expression and purification experiments of the base mismatch-specific binding proteins MSE32 / MSE42 include five steps: strain culture and induction expression, cell disruption, enzyme protein affinity purification, enzyme protein SDS-PAGE electrophoresis detection, and enzyme protein concentration and buffer replacement. The experimental steps are detailed as follows:

[0093] 1. Strain culture and induction expression

[0094] (1) Using a sterile inoculation loop, pick a loop of bacteria from the (-80°C) frozen bacterial solution of the recombinant microorganisms BL21(DE3)pMSE32 / BL21(DE3)pMSE42 expressing the base mismatch-specific binding proteins MSE32 / MSE42 and inoculate it into a 50 mL conical flask containing 10 mL of LB medium with 50 μg / L kanamycin. Incubate at 200 rpm and 37°C overnight until saturation (12h - 16h).

[0095] (2) Take 5 mL of the overnight culture and inoculate it into a 1 L conical flask containing 100 mL of LB medium with 50 μg / L kanamycin at an inoculation amount of 5%. Incubate at 200 rpm and 37°C for 3 h (the liquid volume in the conical flask is 1 / 10, and the inoculation amount is 5%). Five parallel cultures are set up for each strain.

[0096] (3) Pipette 1 mL of the uninduced culture into a 1.5 mL EP tube, measure the absorbance at A550nm, centrifuge at 12000 rpm for 1 min at room temperature to collect the bacterial cells, and store them at -20°C for later use.

[0097] (4) Add 100 μL of 1 M IPTG solution to the remaining culture in step 3 to a final concentration of 1 mM, adjust the culture temperature to 16 °C, and shake at 200 rpm for an additional 6 h.

[0098] (5) After 6 h of induction culture, transfer 1 mL of the induced culture to a 1.5 mL EP tube, and measure the absorbance at A550nm; after 24 h of culture, take another 1 mL sample to measure the absorbance at A550nm and terminate the culture.

[0099] 2. Cell disruption

[0100] (1) Combine the induced cultures (500 mL) and pour them into a 500 mL centrifuge tube pre-cooled on ice. Weigh and balance to an error of less than 0.1 g. Centrifuge at 6000 rpm for 30 min at 4 °C to collect the cells, discard the supernatant, and immediately proceed with the subsequent experimental procedure.

[0101] (2) Resuspend the cells in 25 mL of binding buffer and transfer them to a new 50 mL centrifuge tube pre-cooled on ice (calculate 4 mL of binding buffer per 100 mL of collected cell suspension, plus an additional 5 mL to wash the tube wall). Centrifuge at 6000 rpm for 10 min at 4 °C to wash the cells once. Resuspend the cells in an equal volume of binding buffer as required and combine the resuspended cell suspensions as appropriate.

[0102] (3) Check whether the drain valve on the back of the high-pressure cell crusher (JN-3000, Guangzhou Jueneng Nano Biotechnology Co., Ltd.) is closed (perpendicular to the pipeline when closed), keep it in the closed state, and then add distilled water to the water tank to the scale line.

[0103] (4) Turn on the power switch of the main unit on the right side and press the refrigeration button (the default set temperature is generally 4 °C). Wait for the temperature to drop to the set temperature, and at this time, the water flow in the water tank can be seen to be churning.

[0104] (5) Turn on the power of the high-pressure machine on the left side and ensure that the high-pressure pressure maintenance knob is in the closed state. Adjust the oil pressure knob to 20 bar.

[0105] (6) Check whether the feed cup on the instrument is full of water. Fill the feed cup with double-distilled water, press the air pressure start valve (the green button on the left side of the panel) to start working, and the liquid in the feed cup slowly flows into the high-pressure chamber. Wash the pipeline with about 2 - 3 feed cup volumes of double-distilled water.

[0106] (7) Use 2 - 3 times the cup volume of binding buffer to rinse the high-pressure chamber to make its solution composition consistent with the subsequent sample suspension. Note that during the rinsing process, try not to wait until the bottom of the cup is completely dry to avoid air bubbles entering.

[0107] (8) When the rinsing of the binding buffer is almost finished, pour the resuspended bacterial solution into the feed cup, then quickly collect the lysate at the outlet. Then the preliminary lysate can be transferred back into the feed cup for secondary lysis, with a total of 3 lysis steps.

[0108] (9) After lysis, wash the high-pressure cell disruptor with approximately 500 mL of double-distilled water until the effluent is completely clear. Then turn off the power of the right-side refrigerator, slowly adjust the oil pressure regulating knob to reduce the oil pressure / gas pressure to 0, then turn off the gas pressure and oil pressure working buttons, and loosen the pressure maintaining knob to release the pressure. After the pressure is completely released, tighten the pressure maintaining valve. Turn off the power of the high-pressure machine, and at the same time open the drain valve at the back to drain the cooling water in the water tank. Check and replenish double-distilled water to keep the feed cup full.

[0109] 3. Affinity purification of enzyme protein

[0110] (1) Transfer the cell lysate treated by the ultra-high pressure cell disruptor into a clean pre-cooled round-bottom centrifuge tube dedicated for high-speed centrifugation on ice, weigh and balance to an error of less than 0.01 g. Centrifuge at 4 °C and 12,000 rpm for 1 h. Then use a syringe to aspirate the supernatant, filter it through a 0.22 μm filter membrane, and transfer it into a new pre-cooled clean centrifuge tube for sample loading.

[0111] (2) Use a new affinity column for the purification and preparation of the enzyme protein sample, and operate the AKTA protein purification system instrument and software according to the procedure.

[0112] (3) First, rinse the pipeline with high flow rate of ultrapure water (put both A / B pump heads into the ultrapure water bottle, set the B pump opening to 50%) and exhaust air. Use a syringe to draw out several tubes of liquid from the pump head exhaust knob until there are no obvious bubbles in the pipeline. Then reduce the flow rate to 1 mL / min, connect the affinity column (such as HisTrap HP nickel column, Cytival (GE Health) company) to the system, avoiding air bubbles. Set the flow rate to 2 mL / min and generally rinse for 5 - 10 column volumes.

[0113] (4) Transfer the A / B pumps to the binding buffer bottle to rinse the pipeline and the column, set the flow rate to 2 mL / min, and generally rinse for 5 - 10 column volumes to equilibrate the column.

[0114] (5) Pause the pump, transfer the A pump to the sample centrifuge tube, load the sample at a flow rate of 2 mL / min, and collect the flow-through at the same time. When it is almost finished, pour the flow-through into the remaining sample for re-injection. At the same time, set the program to collect the flow-through sample.

[0115] (6) Then, rinse the pump head of Pump A with a small amount of binding buffer (denoted as Solution A), and transfer Pump A to a bottle of binding buffer without imidazole (containing binding buffer, i.e., Solution A). Flatten the flow-through protein peak to the baseline at a flow rate of 2 mL / min, and collect the sample for the first impurity removal. Generally, rinsing 2 - 5 column volumes is sufficient.

[0116] (7) Next, transfer Pump B to a bottle of elution buffer containing high-concentration imidazole (containing elution buffer, denoted as Solution B). Then, turn on Pump B at a certain ratio to mix with Pump A until the imidazole concentration in the mobile phase reaches 50 mM (the mobile phase consists of 10% Solution B and 90% Solution A). Conduct the second impurity removal at a flow rate of 2 mL / min, and collect the impurity removal sample. Generally, rinsing 2 - 5 column volumes can zero the UV absorption baseline.

[0117] (8) Start linear gradient elution (set Pump B so that the volume ratio of Solution B increases from 0 to 100% within 1 h). Elute the bound protein with high-concentration imidazole elution buffer (the mobile phase consists of Solution B and Solution A, and within 1 h, the volume ratio of Solution B increases from 0 to 100%, and the volume ratio of Solution A decreases from 100% to 0) at a flow rate of 2 mL / min. Set the program to collect the elution sample when there is a protein absorption peak until the absorption peak drops to the baseline. Collect all the eluents to obtain 10 collection fractions numbered 1 - 10 (each collection fraction is 5 mL).

[0118] (9) After the above collection is completed, change Pump B to make the proportion of Solution B reach 100% and continue to elute at a flow rate of 2 mL / min, and rinse the column until the baseline completely returns to zero to elute the residual bound protein. Rinse 5 - 10 column volumes.

[0119] 4. SDS-PAGE electrophoresis detection of enzyme protein

[0120] Perform SDS-PAGE electrophoresis detection on each of the above collection fractions, take pictures through a gel imaging system, and then wrap it well with plastic wrap and store it in a 4°C refrigerator for future use. Analyze information such as the size and gray scale of the protein electrophoresis bands.

[0121] The electrophoresis pattern of medium-scale induced expression of MSE32 protein is as shown in Figure 1 the left figure in the middle. The results show that collection fractions 3 and 4 ( Figure 1 (left) electrophoresis bands 3 and 4) contain the base mismatch specific binding protein MSE32 with the amino acid sequence of SEQ ID No. 2 expressed by recombinant Escherichia coli BL21(DE3) / pMSE32.

[0122] The electrophoresis pattern of medium-scale induced expression of MSE42 protein is as shown in Figure 1 the right figure in the middle. The results show that collection fraction 5 ( Figure 1(Right) Electrophoresis band 5) contains the base mismatch specific binding protein MSE42 with the amino acid sequence SEQ ID No. 4 expressed by recombinant Escherichia coli BL21(DE3) / pMSE42.

[0123] In this example, a collection fraction with a relatively high purity of the target protein was obtained (see the SDS-PAGE electrophoresis diagram of each collection fraction in Figure 1 ), and further enzyme protein concentration and buffer replacement can be carried out.

[0124] 5. Enzyme protein concentration and buffer replacement

[0125] (1) Add 10 mL of ultrapure water to a Millipore 10KD ultrafiltration centrifugal tube, centrifuge at 4°C, select 5000 rpm as the working rate after testing different centrifugation rates for centrifugal filtration, stop centrifugation when the ultrapure water is basically filtered, and pour out the small amount of residual liquid at the bottom.

[0126] (2) Combine the collection fractions 3-4 of MSE32 protein into one tube, and the collection fraction 5 of MSE42 protein into a separate tube. Add them to a 10KD ultrafiltration centrifugal tube according to the protein molecular weight size, and centrifuge at 4°C and 5000 rpm to remove the eluent until the remaining volume is less than 1 mL.

[0127] (3) Prepare a 10× working buffer for the base mismatch specific binding proteins MSE32 / MSE42 at a final concentration of 0.5M NaCl, 0.1M Tris-HCl, 0.1M MgCl2, 0.01M ZnCl, and 0.01M DTT, and split the 10× working buffer into two components: 10× EN Buffer I (a metal ion solution containing 0.1M MgCl2 and 0.01M ZnCl, with the rest being water) and 10× EN Buffer II (a buffer salt solution containing 0.5M NaCl, 0.1M Tris-HCl, and 0.01M DTT, with the rest being water) for addition and use.

[0128] (4) Add 10 mL of the pre-prepared 10× EN Buffer II working buffer (a buffer salt solution containing 0.5M NaCl, 0.1M Tris-HCl, and 0.01M DTT, with the rest being water) to the protein concentrate in the ultrafiltration tube, gently mix, and centrifuge at 4°C and 5000 rpm to remove the buffer until the remaining volume is less than 1 mL.

[0129] (5) Repeat the above step (4) three times until the original eluent ratio drops to 1 / 1000, which can be regarded as the completion of buffer replacement. Base mismatch-specific binding proteins MSE32 and MSE42 (protein concentrates) are obtained respectively. Storage solutions of base mismatch-specific binding proteins MSE32 and MSE42 are prepared by adding 500 μL of sterile 75% glycerol solution to each 1 mL of protein concentrate (where the enzyme protein content of MSE32 / MSE42 is about 50 - 500 mg / L), and stored at -20 °C in a refrigerator for future use.

[0130] Example 2. Error correction effect of base mismatch-specific binding proteins MSE32 / MSE42 on synthetic DNA base errors

[0131] In this example, in addition to the aforementioned conventional experimental instruments, consumables, and reagent drugs, experimental materials for error correction reactions need to be prepared. The principle and process of the error correction reaction are as Figure 2 shown. These materials include the Amp resistance gene mutant (AmpMU) fragment for preparing the error correction reaction substrate, whose sequence is shown in SEQ ID No.5; the template plasmid pPIC9K for PCR amplification to prepare the wild-type Amp resistance gene fragment (AmpWT) of the error correction reaction substrate, whose sequence is shown in SEQ ID No.6; the vector p15C-M1-37-Kan for fragment cloning after the error correction reaction, whose sequence is shown in SEQ ID No.7; as well as high-fidelity DNA polymerase, Gibson assembly kit, DH5a competent cells, etc.

[0132] The nucleotide sequence of AmpMU is:

[0133] The steps of this example are described in detail as follows:

[0134] 1. Using the p15C-M1-37-Kan plasmid (SEQ ID No.7) as a template, the p15C-Vec vector fragment was obtained by PCR amplification using the forward and reverse primers p15C-M1-37-Kan-F3 / p15C-M1-37-Kan-R3 (5`-tgcctcactgattaagcattggtaactgtcagaccaagtttactcatatatac-3`) and (5`-gcgacacggaaatgttgaatactcatactcttcctttttcaatattattgaag-3`), and the target fragment was recovered by gel electrophoresis and gel cutting. The obtained p15C-Vec vector fragment was digested with DpnI. 5 μL of 10×FastDigest Buffer and 1 μL of DpnI enzyme solution were added to the prepared p15C-Vec vector fragment solution, and the volume was made up to 50 μL with water, and the reaction was carried out at 37 °C for 1 h; after the reaction, the DNA sample was directly recovered by column. The OD260 / 280 of the prepared p15C-Vec vector fragment was measured to determine the DNA fragment concentration.

[0135] The p15C-M1-37-Kan plasmid is a circular plasmid with the full sequence shown in SEQ ID No.7. The nucleotide sequence of the p15C-Vec vector fragment obtained by PCR amplification using the forward and reverse primers p15C-M1-37-Kan-F3 and p15C-M1-37-Kan-R3 is positions 2572-1761 of the circular SEQ ID No.7. The p15C-Vec vector fragment has a Kan resistance gene.

[0136] 2. Using the pPIC9K plasmid (SEQ ID No.6) as a template, the wild-type Amp resistance gene fragment AmpWT (positions 8106-8966 of SEQ ID No.6) was obtained by PCR amplification using the forward and reverse primers Amp-F / Amp-R (5`-atgagtattcaacatttccgtgtcgccctt-3`) and (5`-ttaccaatgcttaatcagtgaggcacctatc-3`), and the target fragment was recovered by gel electrophoresis and gel cutting. The obtained target fragment was digested with DpnI. 5 μL of 10×FastDigest Buffer and 1 μL of DpnI enzyme solution were added to the prepared target fragment solution, and the volume was made up to 50 μL with water, and the reaction was carried out at 37 °C for 1 h; after the reaction, the DNA sample was directly recovered by column. The OD260 / 280 of the prepared target fragment was measured to determine the DNA fragment concentration.

[0137] 3. According to the designed Amp resistance gene mutant (AmpMU) fragment sequence for preparing the substrate of error correction reaction, primers for AmpMU gene synthesis were designed using relevant gene synthesis primer design software, and the target gene fragment was obtained by splicing through fusion PCR. The target fragment was recovered by cutting and gel electrophoresis, and the OD260 / 280 of the prepared target fragment was measured to determine the DNA fragment concentration, obtaining the AmpMU fragment. The AmpMU fragment was generated by 10 mutations (including 8 different types of base errors) in the AmpWT fragment as follows: 42DelT, C70G, G242C, T358A, 430InsA, 639DelA, C662T, 726InsG, G756A, 786DelG. The AmpMU fragment is a double-stranded DNA fragment with the nucleotide sequence shown in SEQ ID No.5.

[0138] 4. Using the AmpMU fragment (SEQ ID No.5) and AmpWT fragment (positions 8106 - 8966 of SEQ ID No.6) obtained in the above steps, the substrate AmpAN for error correction reaction was prepared. 700 ng of each of the AmpMU fragment (20 μL) and AmpWT fragment (14 μL) were pipetted in an equimolar ratio into a 0.2 mL PCR tube and mixed evenly, and ultrapure water was added to make the total volume about 50 μL; then it was placed in an Eppendorf PCR instrument, incubated at 98 °C for 15 min, and then the power was turned off and it was allowed to cool naturally for 30 min to room temperature. Through this annealing program, a sample (AmpAN) of DNA molecules with DNA base errors was prepared.

[0139] 5. Two EMC treatments and one untreated control were set up, namely: untreated control AmpAN, EMC treatment of MSE32, and EMC treatment of MSE42.

[0140] (1) EMC treatment of MSE32: Take 25 μL of the DNA sample from the sample prepared in step 4 for error correction reaction. The reaction system and procedure are as follows: Pipette 2 μL of 10×EN Buffer-I (see Example 1), 2 μL of 10×EN Buffer-II (see Example 1), 25 μL of the DNA sample, 2 μL or 3 μL (two parallels) of the base mismatch-specific binding protein MSE32 solution purified in Example 1, react at 37 °C for 0.5 h, and then inactivate at 60 °C for 10 min to obtain the error correction reaction solution.

[0141] (2) EMC treatment of MSE42: Replace the base mismatch-specific binding protein MSE32 solution purified in Example 1 in step (1) with an equal volume of the base mismatch-specific binding protein MSE42 solution purified in Example 1, and the others are the same as in step (1).

[0142] (3) Untreated control AmpAN: The sample (AmpAN) of the DNA molecule prepared in step 4 above.

[0143] 6. Directly use the EMC treatment (error correction reaction) solution in step 5 above as a DNA template, and perform PCR amplification using a touchdown PCR program and Amp-F / Amp-R forward and reverse primers (see step 2) to obtain the full-length Amp resistance gene fragment. The reaction system is as follows: 30 μL of DNA template, 2 μL each of Amp-F / Amp-R forward and reverse primers (20 mM), 10 μL of 5×DNA polymerase buffer, 5 μL of dNTP Mix (2.5 mM), and make up to 50 μL with ultrapure water. Touchdown PCR program: 95°C for 2 min; then perform 8 cycles: 95°C for 30 sec, annealing temperature (starting from 67°C and decreasing by 1°C per cycle) for 30 sec, 72°C for 45 sec; then perform the following 20 cycles: 95°C for 30 sec; 58°C for 30 sec; 72°C for 45 sec; 72°C for 10 min. Recover the full-length fragment by gel electrophoresis and cutting the gel, and measure the OD260 / 280 of the purified sample to determine the DNA concentration.

[0144] 7. Take a certain amount of the Amp resistance gene fragment prepared in step 6 above or the untreated control AmpAN in step 5 and perform a Gibson assembly ligation reaction with a certain amount of p15C-Vec vector fragment (the molar ratio of the fragment to the Vs vector is about 1:2), add 10 μL of 2×Gibson enzyme mixture, the reaction system is 20 μL, and the reaction program is to react at 50°C for 45 min.

[0145] 8. Transform the 20 μL Gibson assembly reaction solution into DH5α competent cells (100 μL) according to the conventional molecular cloning procedure, add 800 μL of LB medium, and incubate at 37°C and 150 rpm for 1 h. After centrifugation, remove part of the supernatant, resuspend the cells with the remaining part, and pipette an equal amount (300 μL) of the resuscitated culture and spread it on a Kan+ monoclonal antibody plate (LB solid medium containing 50 μg / L kanamycin) and a Kan+ / Amp+ double-antibody LB plate (LB solid medium containing 50 μg / L kanamycin and 50 μg / L ampicillin) respectively, and culture overnight (12 - 16 h) at 37°C.

[0146] 9. Count the single colonies growing on the overnight culture plates, and calculate the ratio of the number of colonies growing on the Kan+ / Amp+ double-antibody plate to the number of colonies growing on the Kan+ monoclonal antibody plate as reference data. At the same time, pick single colonies (such as 10 - 20) growing on the Kan+ monoclonal antibody LB plate medium and inoculate them into 4 mL of Kan+ resistant LB liquid medium, and culture overnight at 37°C and 220 rpm.

[0147] 10. Centrifuge to collect the overnight culture, extract the plasmid according to the recommended procedure. After detecting plasmid bands by electrophoresis, send it for Sanger sequencing. The sequencing primers used are Amp-SeqF / R (5'-gagacaataaccctgataaatgcttca-3') and (5'-ctgatgtccggcggtgcgtatatatga-3'). Analyze the proportion of correct sequences in the plasmid sequencing results of the error correction reaction-treated samples and untreated control samples, and use this as the basis for evaluating the effect of the error correction function.

[0148] In this example, the base mismatch-specific binding proteins MSE32 / MSE42 showed good error correction performance. Under the test conditions, the proportion of correct sequences (i.e., containing the AmpWT fragment shown at positions 8106 - 8966 of SEQ ID No.6, the same below) of the DNA substrate fragments treated by the error correction reaction increased significantly (see Table 1). The proportions of correct sequences of the base mismatch-specific binding proteins MSE32 / MSE42 after EMC-PCR treatment were 1 / 10 and 3 / 10 respectively, and the proportions of correct sequences after EMC treatment were 5 / 8 (or 6 / 7) and 8 / 9 (or 9 / 9) respectively. While the proportion of correct sequences of the control samples without error correction reaction treatment was 4 / 8, which proved that the proportion of correct sequences after the error correction reaction could be increased after EMC treatment with the error correction enzyme MSE32 / MSE42. At the same time, the electrophoresis results of the EMC reaction products showed that MSE32 / MSE42 had a certain DNA binding ability (as Figure 3 shown). It indicates that MSE32 / MSE42 can be used for the error correction of artificially synthesized DNA base errors or the detection of DNA base errors / mutations.

[0149] Table 1. Proportions of completely correct sequences of the test samples after EMC treatment or EMC-PCR treatment

[0150]

[0151] Note: NC represents the control sample without error correction reaction treatment; the substrate DNA for the error correction reaction contains 8 different types of errors in the entire 861bp gene fragment.

[0152] Example 3. Base Mismatch-Specific Binding Effect of DNA Error Correction Enzyme MSE32 / MSE42 on Artificially Synthesized DNA

[0153] In this example, in addition to preparing the aforementioned conventional experimental instruments, consumables, and reagent drugs, experimental materials for the error correction reaction also need to be prepared. The principle and process of the error correction reaction are as Figure 2As shown. These materials include the Amp resistance gene mutant (AmpMU) fragment for preparing the error correction reaction substrate, the template plasmid pPIC9K for PCR amplification to prepare the wild-type Amp resistance gene fragment (AmpWT) of the error correction reaction substrate, the vector p15C-M1-37-Kan for fragment cloning after the error correction reaction (each sequence information is shown in Example 2), high-fidelity DNA polymerase, Gibson assembly kit, DH5a competent cells, etc.

[0154] The steps of this example are described in detail as follows:

[0155] 1. Referring to the method described in step 4 of Example 2, four parallel denaturation-annealing reactions (numbered A, B, C, D) were carried out simultaneously to prepare test DNA samples.

[0156] 2. Referring to the method described in step 5 of Example 2, the following groups of EMC reactions were carried out: The test DNA samples in tubes A and B were mixed evenly and then dispensed into 4 independent tubes for EMC treatment of MSE32 enzyme protein. The three parallel reactions were numbered AB-1, AB-2, AB-3, and AB-4 was the untreated control. The test DNA samples in tubes C and D were mixed evenly and then dispensed into 4 independent tubes for EMC treatment of MSE42 enzyme protein. The three parallel reactions were numbered CD-1, CD-2, CD-3, and CD-4 was the untreated control.

[0157] 3. After the EMC reaction, the DNA samples treated with MES32 / 42 were directly subjected to agarose gel electrophoresis. The fluorescent chromophores retained at the loading wells (labeled "up") and the target fragment bands in the lanes (labeled "down") found under ultraviolet light irradiation were excised and recovered respectively. The upper and lower DNA bands of the parallel reaction samples numbered AB-2 and AB-3 were respectively combined into samples numbered "AB-23 up" and "AB-23 down"; The upper and lower bands of the parallel reaction samples numbered CD-2 and CD-3 were respectively combined into samples numbered "CD-23 up" and "CD-23 down".

[0158] 4. Referring to steps 7-10 of Example 2, the recovered products of each treated sample were transformed by seamless cloning into the vector and then spread on the resistant plates. The number of growing colonies was counted, and 20 clones were randomly selected for culturing and plasmid extraction and sequencing.

[0159] In this embodiment, the proportion of completely correct sequences (i.e., containing the AmpWT fragment shown in positions 8106 - 8966 of SEQ ID No. 6, the same below) and the base error rate in the test sample after MSE32 treatment are shown in Table 2; the proportion of completely correct sequences and the base error rate in the test sample after MSE42 treatment are shown in Table 3. The experimental results show that the base error rate of the lower bands of the samples treated with MSE32 error correction is lower than that of the upper bands (where the base error rate of the lower band of AB - 1 drops to 8.6 / 10Kb), and the proportion of completely correct sequences is close. The proportion of completely correct sequences in the lower bands of the samples treated with MSE42 error correction is significantly higher than that of the upper bands, while the base error rate is close. Since the MSE32 / MSE42 enzyme proteins used in the EMC reaction are retained near the loading wells due to their large molecular weights during agarose gel electrophoresis, the fluorescence chromophores generated at this location are produced by the DNA - adsorbed nucleic acid dye bound to them, and the lower bands in the lanes are free DNA without binding to MSE32 / MSE42 proteins. The sequencing results show that the base error rate of the upper bands is relatively high or the proportion of correct sequences is relatively low; therefore, during the EMC reaction process, the MSE32 / MSE42 enzyme proteins separate the DNA molecules containing base errors by specifically binding to base mismatches and retaining them together with the enzyme proteins in the loading wells of the gel electrophoresis, thus playing an error - correction role.

[0160] Table 2. Proportion of correct sequences and base error rate of clone transformants of recovered products of each subdivided DNA component after MSE32 error correction

[0161] 1 2 3 4 5(AN) Proportion of correct sequences 9 / 19 7 / 19 11 / 20 11 / 19 10 / 16 Base error rate 0.00130 0.000862 0.00193 0.00117 0.000596

[0162] Note: 1. Upper band of AB - 1 treatment; 2. Lower band of AB - 1 treatment; 3. Upper band of AB - 23 treatment; 4. Lower band of AB - 23 treatment; 5. AB - 4 AmpAN negative control.

[0163] Table 3. Proportion of correct sequences and base error rate of clone transformants of recovered products of each subdivided DNA component after MSE42 error correction

[0164] 6 7 8 9 10(AN) Proportion of correct sequences 2 / 19 12 / 17 12 / 20 17 / 20 11 / 15 Base error rate 0.00116 0.00187 0.00135 0.00164 0.000389

[0165] Note: 6. Upper band of CD - 1 treatment; 7. Lower band of CD - 1 treatment; 8. Upper band of CD - 23 treatment; 9. Lower band of CD - 23 treatment; 10. CD - 4 AmpAN negative control.

[0166] The foregoing has described the present invention in detail. For those skilled in the art, without departing from the spirit and scope of the present invention and without the need for unnecessary experiments, the present invention can be implemented within a relatively wide range under equivalent parameters, concentrations, and conditions. Although specific embodiments of the present invention are given, it should be understood that the present invention can be further improved. In short, in accordance with the principles of the present invention, this application is intended to cover any variations, uses, or improvements of the present invention, including those that depart from the scope disclosed in this application and are made using conventional techniques known in the art. The application of some basic features can be made within the scope of the appended claims below.

Claims

1. A method for preparing a base mismatch-specific binding protein, comprising: Express the coding gene of the base mismatch-specific binding protein in an organism to obtain the base mismatch-specific binding protein; the organism is a microorganism, a plant or a non-human animal; the base mismatch-specific binding protein comprises a base mismatch-specific endonuclease and a domain that helps the enzyme protein to adsorb to the matrix carrier or promotes the expression of the enzyme protein.

2. The method according to claim 1, wherein: The base mismatch-specific endonuclease is the CEL I enzyme protein or CEL II enzyme protein of celery; and / or The domain that helps the enzyme protein to adsorb to the matrix carrier or promotes the expression of the enzyme protein is the MBP protein; Further, the MBP protein is the MBP protein derived from Escherichia coli.

3. The method according to claim 2, characterized in that: The MBP protein is any one of the following: (A1) A protein with an amino acid sequence of positions 1-388 of SEQ ID No. 2; (A2) A protein obtained by substitution and / or deletion and / or addition of amino acid residues of the protein defined in (A1), having more than 75% identity with the protein defined in (A1) and having the same function.

4. The method according to claim 2 or 3, characterized in that: The CEL I enzyme protein is any one of the following: (B1) A protein with an amino acid sequence of positions 408-702 of SEQ ID No. 2; (B2) A protein obtained by substitution and / or deletion and / or addition of amino acid residues of the protein defined in (B1), having more than 75% identity with the protein defined in (B1) and having the same function; and / or The CEL II enzyme protein is any one of the following: (C1) A protein with an amino acid sequence of positions 408-716 of SEQ ID No. 4; (C2) A protein obtained by substitution and / or deletion and / or addition of amino acid residues of the protein defined in (C1), having more than 75% identity with the protein defined in (C1) and having the same function.

5. The method according to any one of claims 1-4, characterized in that: The base mismatch-specific binding protein is any one of the following: (D1) A protein with an amino acid sequence of SEQ ID No. 2 or SEQ ID No. 4; (D2) A fusion protein obtained by fusing a protein tag to the carboxyl terminus or / and amino terminus of the protein defined in (D1); (D3) A protein obtained by substitution and / or deletion and / or addition of amino acid residues of the protein defined in (D1) or (D2), having more than 75% identity with the protein defined in (D1) or (D2) and having the same function.

6. According to the method described in any one of claims 1-5, characterized in that: The organism is a microorganism; the expression includes introducing the coding gene of the base mismatch-specific binding protein into a recipient microorganism to obtain a recombinant microorganism expressing the coding gene of the base mismatch-specific binding protein, culturing the recombinant microorganism, and expressing to obtain the base mismatch-specific binding protein.

7. A base mismatch-specific binding protein, characterized in that: The base mismatch-specific binding protein is the base mismatch-specific binding protein described in any one of claims 1-6.

8. A biological material, characterized in that, The biological material is any one of the following: (E1) A nucleic acid molecule encoding the base mismatch-specific binding protein described in claim 7; (E2) An expression cassette containing the nucleic acid molecule described in (E1); (E3) A recombinant vector containing the nucleic acid molecule described in (E1); (E4) A recombinant vector containing the expression cassette described in (E2); (E5) A recombinant microorganism containing the nucleic acid molecule described in (E1); (E6) A recombinant microorganism containing the expression cassette described in (E2); (E7) A recombinant microorganism containing the recombinant vector described in (E3); (E8) A recombinant microorganism containing the recombinant vector described in (E4); (E9) A transgenic plant cell line containing the nucleic acid molecule described in (E1) or a transgenic plant cell line containing the expression cassette described in (E2); (E10) A transgenic plant tissue containing the nucleic acid molecule described in (E1) or a transgenic plant tissue containing the expression cassette described in (E2); (E11) A transgenic plant organ containing the nucleic acid molecule described in (E1) or a transgenic plant organ containing the expression cassette described in (E2); (E12) A transgenic animal cell line containing the nucleic acid molecule described in (E1) or a transgenic animal cell line containing the expression cassette described in (E2); (E13) A transgenic animal tissue containing the nucleic acid molecule described in (E1) or a transgenic animal tissue containing the expression cassette described in (E2); (E14) A transgenic animal organ containing the nucleic acid molecule described in (E1) or a transgenic animal organ containing the expression cassette described in (E2).

9. The biomaterial according to claim 8, wherein: In (E1), the nucleic acid molecule is any one of the following: (e1) A nucleic acid molecule whose coding sequence is positions 5071 - 7179 of SEQ ID No.1 or positions 5071 - 7222 of SEQ ID No.3; (e2) A nucleic acid molecule having more than 75% identity with the DNA molecule defined in (e1) and encoding the base mismatch specific binding protein.

10. Any one of the following applications: (F1) The application of the method described in any one of claims 1 - 6 in the preparation of a DNA error correction enzyme preparation; (F2) The application of the MBP protein described in claim 2 or 3 in enhancing the DNA error correction efficiency of a base mismatch specific endonuclease or the application in the preparation of a product for enhancing the DNA error correction efficiency of a base mismatch specific endonuclease; (F3) The application of the nucleic acid molecule encoding the MBP protein described in claim 2 or 3 in the preparation of a product for enhancing the DNA error correction efficiency of a base mismatch specific endonuclease; (F4) The application of the base mismatch specific binding protein described in claim 7 in DNA error correction or the application in the preparation of a DNA error correction enzyme preparation; (F5) The application of the biological material described in claim 8 or 9 in the preparation of a DNA error correction enzyme preparation; (F6) The application of the base mismatch specific binding protein described in claim 7 in recognizing and binding to DNA mismatches or the application in the preparation of a preparation for recognizing and binding to DNA mismatches; (F7) The application of the biological material described in claim 8 or 9 in the preparation of a preparation for recognizing and binding to DNA mismatches.