Fusion protein with DNA (Deoxyribose Nucleic Acid) error correction enzyme activity and preparation method and application thereof

By developing a fusion protein containing base mismatch-specific endonuclease and cellulose-specific binding domain, the problem of low error correction efficiency in DNA synthesis is solved, and efficient and low-cost DNA error correction is achieved, which is suitable for DNA base error detection and disease premature screening.

CN120290523APending Publication Date: 2025-07-11TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410036325.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

现有DNA合成技术中,纠错效率低,尤其在高通量DNA合成中错误率较高,现有纠错技术存在效率低、成本高或设备昂贵的问题。

Method used

A fusion protein containing base mismatch-specific endonuclease and cellulose-specific binding domain was developed to enhance DNA error correction efficiency and achieve efficient error correction through expression vectors and immobilized enzyme forms.

Benefits of technology

It significantly improves DNA error correction efficiency, reduces production costs, and supports automated operations, which is suitable for DNA base error detection and early screening detection of mutations related to important disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004657867720000111
    Figure BDA0004657867720000111
  • Figure BDA0004657867720000121
    Figure BDA0004657867720000121
  • Figure HDA0004657867730000011
    Figure HDA0004657867730000011
Patent Text Reader

Abstract

The invention discloses a fusion protein with DNA (Deoxyribose Nucleic Acid) error correction enzyme activity as well as a preparation method and application thereof. The present invention provides a fusion protein comprising a base mismatch specific endonuclease and a domain that contributes to affinity adsorption of a zymoprotein to a matrix carrier or promotes expression of the zymoprotein. The base mismatched specific endonuclease is escherichia coli T7 bacteriophage endonuclease I, and the structural domain facilitating affinity adsorption of zymoprotein to a specific matrix carrier or promoting zymoprotein expression is a cellulose specific binding structural domain CBM3. The method can be used for DNA base error or mutation detection, base error correction in the gene synthesis process and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of enzymes, and specifically relates to a fusion protein with DNA error-correcting enzyme activity, a preparation method thereof, and an application thereof. Background Art

[0002] Existing research has shown that the occurrence of many major diseases is often closely related to mutations in some key genes; rapid detection of these gene mutations can serve as an effective early screening method and play a role in disease prevention. On the other hand, the rapid development of modern biotechnologies such as synthetic biology and metabolic engineering has greatly increased the demand for artificially synthesized DNA. Currently, artificially synthesized DNA mainly uses the phosphoramidite method for solid-phase synthesis. Due to factors such as synthesizer control technology, synthesis scale, and reagent quality, there is still a certain error rate (such as <1%) in each step of the synthesis reaction, resulting in possible errors in the finally synthesized DNA. For high-throughput DNA synthesis, the problem of a relatively high error rate is more obvious, and the development of efficient error-correcting technologies has become an important part of high-efficiency high-throughput DNA synthesis technologies. According to different application scenarios, the error-correcting technologies used for artificially synthesized DNA can be summarized as pre-use error correction and post-use error correction. Pre-use error-correcting technologies include separation and purification according to the oligonucleotide chain length using polyacrylamide gel electrophoresis (PAGE) or high-performance liquid chromatography (HPLC), which can remove most truncated error oligonucleotide chains; however, the removal rate for base substitutions or partial single-base insertions / deletions is very low. Another pre-use error-correcting technology uses a method of sequencing while synthesizing to ensure that each synthesized oligonucleotide chain is verified. However, this technology still faces problems such as low synthesis efficiency, high equipment and operation costs. Post-use error-correcting technologies use conventional cloning and sequencing to detect errors and repair errors through point mutations. Although correct sequences can be obtained, the time and cost are relatively high. Summary of the Invention

[0003] The technical problem to be solved by the present invention is how to improve DNA error-correcting efficiency.

[0004] To solve the above technical problems, the present invention provides an expression vector (nucleotide sequence SEQ ID No.1) of the following fusion protein (i.e., DNA error-correcting enzyme) and the fusion protein.

[0005] In a first aspect, the present invention claims to protect a fusion protein.

[0006] The fusion protein claimed to be protected by the present invention comprises a base mismatch-specific endonuclease (Mismatch specific endonuclease, MSE) and a domain that helps the enzyme protein to adsorb affinity to a specific matrix carrier or promotes the expression of the enzyme protein.

[0007] In the above-mentioned fusion protein, the base mismatch-specific endonuclease may be Escherichia coli T7 phage endonuclease I; the domain that helps the enzyme protein to adsorb to a specific matrix carrier or promotes the expression of the enzyme protein may be the cellulose-specific binding domain CBM3.

[0008] In the above-mentioned fusion protein, the cellulose-specific binding domain CBM3 may be any of the following:

[0009] (A1) A protein with an amino acid sequence at positions 169-329 of SEQ ID No.2;

[0010] (A2) A protein obtained by substitution and / or deletion and / or addition of amino acid residues to the protein defined in (A1), having more than 75% identity with the protein defined in (A1) and having the same function.

[0011] In the above-mentioned fusion protein, the Escherichia coli T7 phage endonuclease I may be any of the following:

[0012] (B1) A protein with an amino acid sequence at positions 1-149 of SEQ ID No.2;

[0013] (B2) A protein obtained by substitution and / or deletion and / or addition of amino acid residues to the protein defined in (B1), having more than 75% identity with the protein defined in (B1) and having the same function.

[0014] In the above-mentioned fusion protein, the fusion protein may be any of the following:

[0015] (C1) A protein with an amino acid sequence of SEQ ID No.2;

[0016] (C2) A fusion protein obtained by fusing a protein tag to the carboxyl terminus and / or amino terminus of the protein in (C1);

[0017] (C3) A protein obtained by substitution and / or deletion and / or addition of amino acid residues to the protein in (C2), having more than 75% identity with the protein in (C1) or (C2) and having the same function.

[0018] Among them, the protein defined in (C1) is named DNA error correction enzyme MSE1C. In SEQ ID No.2, positions 150-168 are the amino acid sequence encoded by the nucleotide sequence from pET-28a(+) (including a His6-Tag) and a linker amino acid sequence, positions 1-149 are the amino acid sequence of the Escherichia coli T7 phage endonuclease I, and positions 169-329 are the amino acid sequence of the cellulose-specific binding domain CBM3.

[0019] The above-mentioned fusion protein can be artificially synthesized, or its coding gene can be synthesized first and then obtained through biological expression.

[0020] In the above-mentioned fusion protein, the protein-tag refers to a polypeptide or protein that is fused and expressed with the target protein by using in vitro DNA recombination technology to facilitate the expression, detection, tracing, and / or purification of the target protein. The protein-tag can be a Flag tag, His tag, MBP tag, HA tag, myc tag, GST tag, SUMO tag, AviTag tag, eGFP (enhanced green fluorescent protein), eCFP (enhanced cyan fluorescent protein), eYFP (enhanced yellow-green fluorescent protein), and / or mCherry (monomeric red fluorescent protein).

[0021] In the above-mentioned fusion protein, the fusion protein has the property of enhancing the DNA base error correction efficiency of the base mismatch-specific endonuclease.

[0022] The present invention uses the base mismatch-specific cleavage action of the fusion protein for DNA error correction (Enzymatic mismatch correction, EMC).

[0023] In a second aspect, the present invention claims to protect biomaterials related to the fusion protein described in the first aspect above.

[0024] The biomaterials claimed to be protected by the present invention can specifically be any of the following:

[0025] (D1) A nucleic acid molecule encoding the fusion protein described in the first aspect above;

[0026] (D2) An expression cassette containing the nucleic acid molecule described in (D1);

[0027] (D3) A recombinant vector containing the nucleic acid molecule described in (D1);

[0028] (D4) A recombinant vector containing the expression cassette described in (D2);

[0029] (D5) A recombinant microorganism containing the nucleic acid molecule described in (D1);

[0030] (D6) A recombinant microorganism containing the expression cassette described in (D2);

[0031] (D7) A recombinant microorganism containing the recombinant vector described in (D3);

[0032] (D8) A recombinant microorganism containing the recombinant vector described in (D4);

[0033] (D9) A transgenic plant cell line containing the nucleic acid molecule described in (D1) or a transgenic plant cell line containing the expression cassette described in (D2);

[0034] (D10) A transgenic plant tissue containing the nucleic acid molecule described in (D1) or a transgenic plant tissue containing the expression cassette described in (D2);

[0035] (D11) A transgenic plant organ containing the nucleic acid molecule described in (D1) or a transgenic plant organ containing the expression cassette described in (D2);

[0036] (D12) A transgenic animal cell line containing the nucleic acid molecule described in (D1) or a transgenic animal cell line containing the expression cassette described in (D2);

[0037] (D13) A transgenic animal tissue containing the nucleic acid molecule described in (D1) or a transgenic animal tissue containing the expression cassette described in (D2);

[0038] (D14) A transgenic animal organ containing the nucleic acid molecule described in (D1) or a transgenic animal organ containing the expression cassette described in (D2).

[0039] Among the above biological materials, the nucleic acid molecule may be DNA, such as cDNA, genomic DNA or recombinant DNA; the nucleic acid molecule may also be RNA.

[0040] Specifically, the nucleic acid molecule described in (D1) may be any of the following:

[0041] (d1) A nucleic acid molecule whose coding sequence is positions 5071 - 6060 of SEQ ID No.1;

[0042] (d2) A nucleic acid molecule having more than 75% identity with the DNA molecule defined in (d1) and encoding the fusion protein.

[0043] Among the above-mentioned biological materials, the expression cassette refers to DNA capable of expressing the fusion protein in a host cell. The expression cassette may also include a single-stranded or double-stranded nucleic acid molecule containing all regulatory sequences necessary for the nucleic acid molecule expressing any one of the above proteins. The regulatory sequences can direct the coding sequence to express any one of the above proteins in a suitable host cell under their compatible conditions. The regulatory sequences include, but are not limited to, leader sequences, polyadenylation sequences, propeptide sequences, promoters, signal sequences, and transcription terminators. At a minimum, the regulatory sequences should include a promoter and transcription and translation termination signals. To introduce specific restriction enzyme sites into the vector for ligating the regulatory sequences with the coding region of the nucleic acid sequence encoding the protein, regulatory sequences with linkers can be provided. The regulatory sequence can be a promoter sequence, i.e., a nucleic acid sequence recognizable by the host cell expressing the nucleic acid sequence. The promoter sequence contains transcriptional regulatory sequences mediating protein expression. The promoter can be any nucleic acid sequence having transcriptional activity in the selected host cell, including mutant, truncated, and chimeric promoters, and can be derived from genes encoding extracellular or intracellular proteins homologous or heterologous to the host cell. The regulatory sequence can also be a transcription termination sequence, i.e., a sequence recognized by the host cell to terminate transcription. The termination sequence is operably linked to the 3'-end of the nucleic acid sequence encoding the protein. Any terminator functional in the selected host cell can be used in the present invention. The regulatory sequence can also be a leader sequence, i.e., an untranslated region of mRNA that is important for translation in the host cell. The leader sequence is operably linked to the 5'-end of the nucleic acid sequence encoding the protein. Any leader sequence functional in the selected host cell can be used in the present invention. It may also be necessary to add regulatory sequences that can regulate protein expression according to the growth of the host cell. Examples of regulatory systems are those that can respond to chemical or physical stimulants (including in the presence of regulatory compounds) to turn on or off gene expression. Other examples of regulatory sequences are those that can amplify genes. In these examples, the nucleic acid sequence encoding the protein should be operably linked to the regulatory sequence.

[0044] The present invention also relates to a recombinant expression vector comprising a nucleic acid molecule encoding any one of the above-mentioned fusion proteins of the present invention, a promoter, and transcriptional and translational termination signals. When preparing the expression vector, the nucleic acid molecule encoding any one of the above-mentioned fusion proteins can be positioned in the vector so as to be operably linked to appropriate expression control sequences. The recombinant expression vector can be any vector (such as a plasmid or a virus) that facilitates recombinant DNA manipulation and expression of the nucleic acid sequence. The choice of the vector usually depends on the compatibility of the vector with the host cell into which it is to be introduced. The vector can be a linear or closed circular plasmid. The vector can be an autonomously replicating vector (i.e., a complete structure existing outside the chromosome and capable of replicating independently of the chromosome), such as a plasmid, an episome, a minichromosome, or an artificial chromosome. The vector can contain any mechanism that ensures self-replication. Alternatively, the vector is a vector that, when introduced into the host cell, will integrate into the genome and replicate together with the chromosome into which it is integrated. In addition, a single vector or plasmid can be used, or two or more vectors or plasmids that collectively contain all the DNA to be introduced into the host cell genome, or a transposon. The vector contains one or more selectable markers that facilitate the selection of transformed cells. A selectable marker is a gene whose product confers resistance to biocides or viruses, resistance to heavy metals, or prototrophy to auxotrophs, etc. Examples of bacterial selectable markers are the dal gene of Bacillus subtilis or Bacillus licheniformis, or resistance markers for antibiotics such as ampicillin, kanamycin, chloramphenicol, or tetracycline. The vector contains elements that enable the vector to stably integrate into the host cell genome or ensure autonomous replication of the vector in the cell independently of the cell genome. In the case of autonomous replication, the vector can also contain an origin of replication that enables the vector to replicate autonomously in the target host cell. The origin of replication can carry a mutation that makes it temperature-sensitive in the host cell (see, for example, fEhrlich, 1978, Proceedings of the National Academy of Sciences of the United States of America 75: 1433). More than one copy of the nucleic acid molecule encoding any one of the above-mentioned proteins of the present invention can be inserted into the host cell to increase the yield of the gene product. The increase in the copy number of the nucleic acid molecule can be achieved by inserting at least one additional copy of the nucleic acid molecule into the host cell genome, or by inserting an amplifiable selectable marker together with the nucleic acid molecule, and by culturing the cells in the presence of a suitable selection reagent to select cells containing the amplified copy of the selectable marker gene and thus containing an additional copy of the nucleic acid molecule. The operations for ligating the above-mentioned elements to construct the recombinant expression vector of the present invention are well known to those skilled in the art (see, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, 1989).

[0045] The term "operably linked" is defined herein as a conformation in which the regulatory sequence is in an appropriate position relative to the coding sequence of the DNA sequence so that the regulatory sequence directs the expression of the protein.

[0046] The vectors described herein are well-known to those skilled in the art and include, but are not limited to: plasmids, phages (such as λ phage or M13 filamentous phage, etc.), cosmids (i.e., cosmid plasmids), Ti plasmids, or viral vectors.

[0047] Among the above biological materials, the microorganism can be any one of the following: prokaryotic microorganisms; Gram-negative bacteria; bacteria of the genus Escherichia; Escherichia coli; Escherichia coli BL21(DE3).

[0048] As used herein, the term "identity" refers to the sequence similarity to a nucleic acid sequence or an amino acid sequence. Identity can be evaluated by the naked eye or by computer software. Using computer software, the identity between two or more sequences can be expressed as a percentage (%), which can be used to evaluate the identity between related sequences.

[0049] As used herein, an identity of more than 75% can be an identity of 80%, 85%, 90% or more than 95%. The identity of more than 80% can be at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%.

[0050] Those of ordinary skill in the art can easily mutate the protein of the present invention (C1) or its coding gene by using known methods, such as directed evolution and site-directed mutagenesis. Those nucleotides that have been artificially modified and have an identity of 75% or higher with the nucleotide sequence of the protein of the present invention (C1), as long as they encode the DNA error-correcting enzyme protein and have the DNA error-correcting enzyme function, are derived from the nucleotide sequence of the present invention and are equivalent to the sequence of the present invention.

[0051] In a third aspect, the present invention claims to protect a method for preparing the fusion protein described in the first aspect above.

[0052] The method for preparing the fusion protein claimed by the present invention and described in the first aspect above includes: expressing the coding gene of the fusion protein in an organism to obtain the fusion protein.

[0053] Among them, the organism can be a microorganism, a plant or a non-human animal.

[0054] Further, the organism is a microorganism. Correspondingly, the expression includes introducing the coding gene of the fusion protein into a recipient microorganism to obtain a recombinant microorganism expressing the coding gene of the fusion protein, culturing the recombinant microorganism, and expressing to obtain the fusion protein.

[0055] In the above method, the recipient microorganism can be any one of the following: prokaryotic microorganism; Gram-negative bacterium; bacterium of the genus Escherichia; Escherichia coli; Escherichia coli BL21(DE3).

[0056] Furthermore, the recombinant microorganism can be recombinant Escherichia coli. In one embodiment of the present invention, the recombinant Escherichia coli is a recombinant microorganism obtained by introducing pMSE1C into Escherichia coli BL21(DE3) and expressing the fusion protein with the amino acid sequence of SEQ ID No.2, and the recombinant microorganism is named BL21(DE3) / pMSE1C. pMSE1C is a recombinant expression vector obtained by replacing the fragment between the upstream homologous arm sequence (nucleotide sequence positions 5021 - 5070 of SEQ ID No.1) and the downstream homologous arm sequence (nucleotide sequence positions 6061 - 6110 of SEQ ID No.1) of pET-28a(+) with the nucleotide sequence positions 5071 - 6060 of SEQ ID No.1 while keeping other nucleotides unchanged.

[0057] Fourthly, the present invention claims to protect an immobilized enzyme.

[0058] The enzyme active component in the immobilized enzyme claimed by the present invention is the fusion protein described in the first aspect above.

[0059] Furthermore, the immobilized enzyme is prepared based on the affinity between the cellulose-specific binding domain CBM3 in the fusion protein and cellulose.

[0060] Even further, the immobilized enzyme is prepared by a method comprising the following steps: First, a microcrystalline cellulose suspension is prepared using microcrystalline cellulose and the working buffers (10×EN-Buffer I and 10×EN-Buffer II) of DNA error correction enzyme MSE1C; then, DNA error correction enzyme MSE1C (i.e., the fusion protein) is added to the microcrystalline cellulose suspension, and after mixing, it is transferred to a hollow column with a glass fiber membrane at the bottom, and the supernatant is removed to obtain the immobilized enzyme.

[0061] Fifthly, the present invention claims to protect any one of the following applications:

[0062] (E1) The application of the fusion protein described in the first aspect above or the immobilized enzyme described in the fourth aspect above in DNA error correction; or the application in the preparation of a DNA error correction enzyme preparation;

[0063] (E2) The application of the biological material described in the second aspect above or the immobilized enzyme described in the fourth aspect above in the preparation of a DNA error correction enzyme preparation;

[0064] (E3) Use of the method described in the third aspect above in the preparation of a DNA error-correcting enzyme preparation;

[0065] (E4) Use of the cellulose-specific binding domain CBM3 described in the first aspect above in enhancing the DNA error-correcting efficiency of a base mismatch-specific endonuclease or in the preparation of a product for enhancing the DNA error-correcting efficiency of a base mismatch-specific endonuclease;

[0066] (E5) Use of a nucleic acid molecule encoding the cellulose-specific binding domain CBM3 described in the first aspect above in the preparation of a product for enhancing the DNA error-correcting efficiency of a base mismatch-specific endonuclease.

[0067] In (E4) and (E5), the base mismatch-specific endonuclease may specifically be Escherichia coli T7 phage endonuclease I. Further, the Escherichia coli T7 phage endonuclease I is any one of the following: (B1) a protein having an amino acid sequence of positions 1-149 of SEQ ID No. 2; (B2) a protein obtained by substitution and / or deletion and / or addition of amino acid residues to the protein defined in (B1), having more than 75% identity with the protein defined in (B1) and having the same function.

[0068] Further, in the above application, the fusion protein can function either in a free form or in the form of an immobilized enzyme. The functioning specifically refers to performing DNA error correction.

[0069] Experiments have proved that the fusion protein (i.e., DNA error-correcting enzyme) of the present invention has an obvious error-correcting effect on artificially synthesized DNA base errors, shows obvious advantages compared with the control DNA error-correcting enzyme, the error-correcting efficiency of its DNA base errors is increased by more than one time, and it can function in an immobilized form.

[0070] The error-correcting enzyme disclosed in the present invention is expressed and produced using a prokaryotic expression system and can be conveniently purified by affinity chromatography on a specific matrix carrier, which can reduce production costs such as expression and purification. At the same time, the error-correcting enzyme disclosed in the present invention can be applied to methods such as the detection of DNA base errors or mutations and the error correction of base errors in the gene synthesis process, and has advantages such as high efficiency, speed, convenience for automated operation, and low cost. It can be used in research and application fields including but not limited to the early screening detection of gene mutations related to important diseases and the error correction of artificially synthesized DNA base errors. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1SDS-PAGE electrophoresis pattern of medium-level induced expression of MSE1C enzyme protein. Among them, the arrow indicates the target band (DNA error correction enzyme MSE1C). Lanes 1-10 are serial elution samples of the purified DNA error correction enzyme MSE1C protein after the second round of purification by nickel column, and the two lanes on the right are Markers.

[0072] Figure 2 Principle and experimental procedure of MSE1C for correcting DNA base errors in artificial synthesis. Specific implementation mode

[0073] The present invention will be further described in detail below in conjunction with specific implementation modes. The examples given are only for clarifying the present invention, rather than limiting the scope of the present invention.

[0074] The experimental methods in the following examples are all conventional methods unless otherwise specified.

[0075] The materials, reagents, etc. used in the following examples can be obtained from commercial sources unless otherwise specified.

[0076] In the embodiments of the present invention, the experimental materials to be prepared include instruments and consumables, culture media, reagents and drugs, etc. Among them, conventional instruments and consumables include a constant temperature shaking incubator, a sterile operating table, an induction cooker, a protein electrophoresis tank (XCellSureLock Mini-Cell, Invitrogen), a decolorizing shaking table, a gel imaging system, a disposable SDS-PAGE gel (Invitrogen Precast gel), test tubes, 50 mL and 1 L conical flasks, 12-layer gauze, etc. Common culture media include LB liquid culture medium, LB solid plate culture medium (formula see "Molecular Cloning Experiment Guide"). Conventional reagents and drugs include 1000× kanamycin solution (50 mg / L), 1 M IPTG solution, 4× NuPAGE LDS loading buffer, NuPAGE Tris-acetate SDS electrophoresis buffer, DNA and protein molecular weight standards (Blue V Protein Marker (10-190 kDa), Transgen), Coomassie brilliant blue staining solution (formula: dissolve 0.25 g of Coomassie brilliant blue R-250 in every 100 mL of methanol: acetic acid solution (decolorizing solution), filter with filter paper to remove particulate matter), decolorizing solution (formula: 500 mL of methanol, add 400 mL of water, mix well and then mix with 100 mL of glacial acetic acid), binding buffer (formula: 20 mM Na3PO4, 500 mM NaCl, pH 6.0, the rest is water), elution buffer (formula: add imidazole with a final concentration of 500 mM to the binding buffer, the rest is water), 20% ethanol solution, ultrapure water, etc.

[0077] Example 1: Production, Expression and Purification of DNA Error-Correcting Enzyme MSE1C

[0078] In this example, the recombinant Escherichia coli BL21(DE3) / pMSE1C was used to prepare the DNA error-correcting enzyme MSE1C.

[0079] The recombinant Escherichia coli BL21(DE3) / pMSE1C was prepared according to the following method:

[0080] (1) Preparation of BL21(DE3) / pMSE1C

[0081] Using the Escherichia coli T7 phage T7EI gene sequence in the Genbank database as a template, gene synthesis primer sequences were designed using gene synthesis sequence design-related software. The gene sequence encoding the base mismatch-specific endonuclease (MSE) in the DNA error-correcting enzyme MSE1C of the present invention was obtained by fusion PCR splicing (named T7EI).

[0082] Using the gene sequence of the CBM3 domain (CtCBM3 motif) from Clostridium thermocellum retrieved from the Uniprot protein database as a template, after optimizing the expression codons for Escherichia coli, gene synthesis primer sequences were designed using gene synthesis sequence design-related software. The gene sequence encoding the domain part (CtCBM3 motif) that helps the enzyme protein to adsorb to a specific matrix carrier or promotes the expression of the enzyme protein in the DNA error-correcting enzyme MSE1C of the present invention was obtained by fusion PCR splicing (named CBM3-linker).

[0083] On this basis, T7EI and CBM3-linker were fused into an expression cassette by fusion PCR splicing to obtain the expression cassette of the DNA error-correcting enzyme MSE1C of the present invention. The nucleotide sequence of this expression cassette is positions 5071 - 6060 of SEQ ID No.1, wherein positions 5021 - 5070 and 6061 - 6110 of SEQ ID No.1 are the upstream homologous arm sequence and the downstream homologous arm sequence of pET-28a(+), respectively.

[0084] The expression cassette of the DNA error-correcting enzyme MSE1C constructed was inserted into the expression vector pET-28a(+) by cloning methods such as Gibson assembly, and placed under the transcriptional regulation of the T7 promoter and T7 terminator to obtain the recombinant vector pMSE1C expressing the DNA error-correcting enzyme MSE1C of the present invention. The recombinant expression vector pMSE1C was obtained by replacing the fragment between the upstream homologous arm sequence (nucleotide sequence positions 5021-5070 of SEQ ID No.1) and the downstream homologous arm sequence (nucleotide sequence positions 6061-6110 of SEQ ID No.1) of pET-28a(+) with the nucleotide sequence positions 5071-6060 of SEQ ID No.1, while keeping other nucleotides unchanged to obtain a recombinant plasmid.

[0085] The recombinant expression vector pMSE1C expresses the DNA error-correcting enzyme MSE1C with the amino acid sequence of SEQ ID No.2. The complete plasmid sequence of the recombinant expression vector pMSE1C is shown in SEQ ID No.1, with a full length of 6196bp. Among them, positions 1-5070, positions 5518-5574 (including the His6-Tag coding sequence), and positions 6061-6196 are nucleotide sequences from pET-28a(+), positions 5071-5517 are the gene sequence of Escherichia coli T7 endonuclease I (T7EI), and positions 5575-6060 are the (CBM3-linker) gene sequence.

[0086] Positions 5071-6060 of SEQ ID No.1 are the coding frame sequence of the MSE1C gene, encoding the DNA error-correcting enzyme MSE1C shown in SEQ ID No.2. In SEQ ID No.2, positions 1-149 are the amino acid sequence derived from Escherichia coli T7 endonuclease I, positions 169-329 are the amino acid sequence of CBM3, and positions 150-168 are the amino acid sequence encoded by the nucleotide sequence from pET-28a(+) (including a His6-Tag and a linker amino acid sequence).

[0087] The recombinant expression vector pMSE1C was transformed into Escherichia coli BL21(DE3) competent cells. It was evenly spread on an LB plate containing kanamycin and cultured at 37°C for 16 hours. Single colonies were cultured overnight with shaking, and plasmids were extracted and sequenced to verify the correct positive clones. The recombinant Escherichia coli containing the pMSE1C vector is the recombinant Escherichia coli BL21(DE3) / pMSE1C. BL21(DE3) / pMSE1C expresses the DNA error-correcting enzyme MSE1C with the amino acid sequence of SEQ ID No.2.

[0088] In this example, the experiment for the production, expression and purification of DNA error-correcting enzyme MSE1C (SEQ ID No. 2) includes five steps: strain culture and induction expression, cell disruption, affinity purification of enzyme protein, SDS-PAGE electrophoresis detection of enzyme protein, and concentration and buffer replacement of enzyme protein. The experimental steps are detailed as follows:

[0089] 1. Strain culture and induction expression

[0090] (1) Dip a loop of bacteria from the frozen bacterial solution (-80°C) of recombinant Escherichia coli BL21(DE3) / pMSE1C expressing DNA error-correcting enzyme MSE1C into a 50 mL conical flask containing 10 mL of LB medium with 50 μg / mL kanamycin, and culture overnight at 37°C with 200 rpm until saturation (12 h - 16 h).

[0091] (2) Take 5 mL of the overnight culture and inoculate it into a 1 L conical flask containing 100 mL of LB medium with 50 μg / mL kanamycin at an inoculation amount of 5%, and culture at 37°C with 200 rpm for 3 h (the liquid volume in the conical flask is 1 / 10, and the inoculation amount is 5%). Five parallel cultures are set up for each strain.

[0092] (3) Pipette 1 mL of the uninduced culture into a 1.5 mL EP tube, measure the absorbance at A550nm, centrifuge at 12000 rpm for 1 min at room temperature to collect the cells, and store them at -20°C for later use.

[0093] (4) Add 100 μL of 1 M IPTG solution to the remaining culture in step (3) to a final concentration of 1 mM, adjust the culture temperature to 16°C, and continue shaking culture at 200 rpm for 6 h.

[0094] (5) After 6 h of induction culture, pipette 1 mL of the induced culture into a 1.5 mL EP tube and measure the absorbance at A550nm; after 24 h of culture, take another 1 mL sample to measure the absorbance at A550nm and terminate the culture.

[0095] 2. Cell disruption

[0096] (1) Combine the five parallel induced cultures (500 mL) and pour them into a 500 mL centrifuge tube pre-cooled on ice, weigh and balance to an error of less than 0.1 g. Centrifuge at 6000 rpm for 30 min at 4°C to collect the cells, pour out the supernatant, and immediately proceed with the subsequent experimental procedures.

[0097] (2) Resuspend the bacterial cells with 25 mL of binding buffer and transfer them into a new 50 mL centrifuge tube pre-cooled on ice (calculate according to adding 4 mL of binding buffer for every 100 mL of collected bacterial liquid, and add an additional 5 mL to wash the tube wall). Centrifuge the bacterial cells at 6000 rpm for 10 min at 4°C to wash them once. Resuspend the bacterial cells by adding an equal volume of binding buffer as required previously.

[0098] (3) Check whether the drain valve on the back of the high-pressure cell crusher (JN-3000, Guangzhou Jueneng Nano Biotechnology Co., Ltd.) is closed (perpendicular to the pipeline when closed), keep it in the closed state, and then add distilled water to the water tank to the scale line.

[0099] (4) Turn on the power switch of the main unit on the right side and start the refrigeration button (the default set temperature is generally 4°C). Wait until the temperature drops to the set temperature. At this time, the water flow in the water tank can be seen to be tumbling.

[0100] (5) Turn on the power of the high-pressure machine on the left side and ensure that the high-pressure pressure maintaining knob is in the closed state. Adjust the oil pressure knob to 20 bar.

[0101] (6) Check whether the feed cup on the instrument is full of water. Fill the feed cup with double-distilled water, press the air pressure start valve (the green button on the left side of the panel) to start working, and the liquid in the feed cup slowly flows into the high-pressure chamber. Wash the pipeline with about 2 - 3 volumes of double-distilled water of the feed cup.

[0102] (7) Use 2 - 3 volumes of the binding buffer to rinse the high-pressure chamber to make its solution composition consistent with the subsequent sample suspension. Note that try not to wait until the bottom of the cup is completely dry during the rinsing process to avoid air bubbles entering.

[0103] (8) When the rinsing with the binding buffer is almost finished, pour the resuspended bacterial liquid into the feed cup, then quickly collect the lysate at the outlet. Then the preliminary lysate can be transferred back into the feed cup for secondary crushing, and crush a total of 3 times.

[0104] (9) After the crushing is completed, wash the high-pressure cell crusher with about 500 mL of double-distilled water until the effluent is completely clear. Then turn off the power of the refrigerator on the right side, slowly adjust the oil pressure regulating knob to reduce the oil pressure / air pressure to 0, then turn off the air pressure and oil pressure working buttons, and loosen the pressure maintaining knob to release the pressure. After the pressure is completely released, tighten the pressure maintaining valve. Turn off the power of the high-pressure machine, and at the same time open the drain valve at the back to drain the cooling water in the water tank. Check and replenish double-distilled water to keep the feed cup full.

[0105] 3. Affinity purification of enzyme protein

[0106] (1) Transfer the cell lysate treated by an ultra-high pressure cell disruptor into a clean round-bottom centrifuge tube for high-speed centrifugation pre-cooled on ice, weigh and balance to an error of less than 0.01 g. Centrifuge at 12,000 rpm for 1 h at 4 °C. Then, use a syringe to aspirate the supernatant, filter it through a 0.22-μm filter membrane, and transfer it into a new pre-cooled clean centrifuge tube for sample loading.

[0107] (2) Use a new affinity column for the purification and preparation of the enzyme protein sample, and operate the AKTA protein purification system instrument and software according to the procedure.

[0108] (3) First, rinse the pipeline at a high flow rate with ultrapure water (put both A / B pump heads into an ultrapure water bottle, set the B pump opening to 50%) and exhaust air. Use a syringe to draw out several tubes of liquid from the pump head exhaust knob until there are no obvious bubbles in the pipeline. Then, reduce the flow rate to 1 mL / min, connect the affinity column (HisTrap HP nickel column, Cytiva (GE Life) company) to the system, and avoid air bubbles entering. Set the flow rate to 2 mL / min and rinse for 5 - 10 column volumes.

[0109] (4) Transfer the A / B pumps to the binding buffer bottle to rinse the pipeline and the column, set the flow rate to 2 mL / min, and rinse for 5 - 10 column volumes to balance the column.

[0110] (5) Pause the pump, transfer the A pump to the sample centrifuge tube, load the sample at a flow rate of 2 mL / min, and collect the flow-through at the same time. When it is almost finished, pour the flow-through into the remaining sample and inject it again. At the same time, set the program to collect the flow-through sample.

[0111] (6) Then, rinse the pump head with a small amount of binding buffer (denoted as solution A) and transfer the A pump to the binding buffer bottle without imidazole (containing the binding buffer, i.e., solution A). Flatten the flow-through protein peak to the baseline at a flow rate of 2 mL / min and collect the first-step impurity removal sample. Rinse for 2 - 5 column volumes.

[0112] (7) Then, transfer the B pump to the elution buffer bottle containing a high concentration of imidazole (containing the elution buffer, denoted as solution B). Then, open the B pump at a certain ratio and mix it with the A pump until the imidazole concentration in the mobile phase is 50 mM (the mobile phase consists of 10% solution B and 90% solution A). Perform the second-step impurity removal at a flow rate of 2 mL / min and collect the impurity removal sample. Rinse for 2 - 5 column volumes, and the ultraviolet absorption baseline can be reset to zero.

[0113] (8) Start the linear gradient elution (set pump B to increase the volume ratio of solution B from 0 to 100% within 1 h), and elute the bound protein with a high-concentration imidazole elution buffer (the mobile phase consists of solution B and solution A, the volume ratio of solution B increases from 0 to 100% within 1 h, and the volume ratio of solution A decreases from 100% to 0) at a flow rate of 2 mL / min. Set the program to collect the elution samples when there is a protein absorption peak until the absorption peak drops to the baseline. Collect all the eluates to obtain 10 collected fractions numbered 1 - 10 (each collected fraction is 5 mL).

[0114] (9) After the above collection is completed, change pump B to make the proportion of solution B reach 100% and continue to elute at a flow rate of 2 mL / min, and rinse the column until the baseline returns to zero completely to elute the residual bound protein. Rinse for 5 - 10 column volumes.

[0115] (10) Transfer pumps A and B to the ultrapure water bottle and rinse the column at a flow rate of 2 mL / min until the conductivity returns to zero completely, and rinse for 10 column volumes.

[0116] (11) Transfer pumps A and B to the 20% ethanol solution bottle, rinse the column at a flow rate of 2 mL / min for 10 column volumes, and stop the machine to seal the pipeline.

[0117] 4. SDS-PAGE electrophoresis detection of enzyme protein

[0118] Perform SDS-PAGE electrophoresis detection on each of the above collected fractions, take pictures with a gel imaging system, and then wrap it with plastic wrap and store it in a refrigerator at 4°C for later use. Analyze information such as the size and gray scale of the protein electrophoresis bands. The results show that collected fractions 5 - 8 contain the fusion protein MSE1C with the amino acid sequence of SEQ ID No.2 expressed by recombinant Escherichia coli BL21(DE3) / pMSE1C.

[0119] In this example, collected fractions with relatively high purity of the target protein were obtained (the SDS-PAGE electrophoresis diagrams of each collected fraction are shown in Figure 1 ), and further enzyme protein concentration and buffer replacement can be carried out.

[0120] 5. Enzyme protein concentration

[0121] (1) Add 10 mL of ultrapure water to a Millipore 10KD ultrafiltration centrifugal tube, centrifuge at 4°C, select 5000 rpm as the working rate for centrifugal filtration after testing different centrifugation rates, stop centrifugation when the ultrapure water is basically filtered, and pour out the small amount of residual liquid at the bottom.

[0122] (2) Combine collected fractions 5 - 8 into one tube, add them to a 10KD ultrafiltration centrifugal tube according to the protein molecular weight size, and centrifuge at 4°C and 5000 rpm to remove the eluate until the remaining volume is less than 1 mL.

[0123] (3) Prepare 10× working buffer for MSE1C enzyme (the rest is water) with a final concentration of 0.5 M NaCl, 0.1 M Tris-HCl, 0.1 M MgCl₂, 0.01 M ZnCl, and 0.01 M DTT, and split the 10× working buffer into two components: 10× EN Buffer I (a metal ion solution containing 0.1 M MgCl₂ and 0.01 M ZnCl, the rest is water) and 10× EN Buffer II (a buffer salt solution containing 0.5 M NaCl, 0.1 M Tris-HCl, and 0.01 M DTT, the rest is water) for addition and use.

[0124] (4) Add 10 mL of the pre-prepared enzyme working buffer 10× EN Buffer II (a buffer salt solution containing 0.5 M NaCl, 0.1 M Tris-HCl, and 0.01 M DTT, the rest is water) to the protein concentrate in the ultrafiltration tube. After gently mixing, centrifuge at 4°C and 5000 rpm to remove the buffer until the remaining volume is less than 1 mL.

[0125] (5) Repeat step (4) three times until the proportion of the original eluate drops to 1 / 1000 to complete the buffer replacement and obtain the DNA error correction enzyme MSE1C solution (protein concentrate). Prepare the MSE1C enzyme protein storage solution (the enzyme protein content is about 50 - 500 mg / L) by adding 500 μL of sterile 75% glycerol solution to every 1 mL of the protein concentrate, and store it in a -20°C refrigerator for later use.

[0126] Example 2: Application of free DNA error correction enzyme MSE1C in correcting base errors of synthetic DNA

[0127] In this example, in addition to the aforementioned conventional experimental instruments, consumables, and reagent drugs, experimental materials for the error correction reaction need to be prepared. The principle and process of the error correction reaction are as Figure 2 shown. These materials include the Amp resistance gene mutant fragment (AmpMU) for preparing the error correction reaction substrate, the Amp resistance gene wild-type fragment (AmpWT) for preparing the error correction reaction substrate, the p15C-M1-37-Kan vector for fragment cloning after the error correction reaction, high-fidelity DNA polymerase, Gibson assembly kit, DH5α competent cells, etc.

[0128] The detailed steps of this example are as follows:

[0129] 1. Using the p15C-M1-37-Kan plasmid as a template, the p15C-Vec vector fragment was obtained by PCR amplification with the forward and reverse primers p15C-M1-37-Kan-F3 (5’-tgcctcactgattaagcattggtaactgtcagaccaagtttactcatatatac-3’) and p15C-M1-37-Kan-R3 (5’-gcgacacggaaatgttgaatactcatactcttcctttttcaatattattgaag-3’). The target fragment was recovered by gel electrophoresis and gel extraction. The obtained p15C-Vec vector fragment was digested with DpnI. 5 μL of 10× FastDigest Buffer (DpnI restriction enzyme reaction buffer, NEB) and 1 μL of DpnI enzyme solution were added to the prepared p15C-Vec vector fragment solution, and the volume was made up to 50 μL with water. The reaction was carried out at 37 °C for 1 h. After the reaction, the DNA sample was directly recovered by column to obtain the p15C-Vec vector fragment. The OD260nm / 280nm of the prepared p15C-Vec vector fragment was measured to determine the DNA fragment concentration.

[0130] The p15C-M1-37-Kan plasmid is a circular plasmid with the full sequence shown in SEQ ID No.3. The nucleotide sequence of the p15C-Vec vector fragment obtained by PCR amplification using the forward and reverse primers p15C-M1-37-Kan-F3 and p15C-M1-37-Kan-R3 is the 2572-1761 positions of the circular SEQ ID No.3. The p15C-Vec vector fragment has a Kan resistance gene.

[0131] 2. Prepare the wild-type fragment of the ampicillin (Amp) resistance gene, namely the AmpWT fragment.

[0132] The AmpWT fragment is a double-stranded DNA fragment with the nucleotide sequence shown in SEQ ID No.4.

[0133] 3. According to the designed Amp resistance gene mutant (AmpMU) fragment sequence for preparing the error correction reaction substrate, primers for AmpMU gene synthesis were designed using relevant gene synthesis primer design software, and the target gene fragment was obtained by fusion PCR splicing. The target fragment was recovered by gel electrophoresis and gel cutting, and the OD260nm / 280nm of the prepared target fragment was measured to determine the DNA fragment concentration, obtaining the AmpMU fragment. The AmpMU fragment was generated from the AmpWT fragment by the following 10 mutations (containing 8 different types of base errors): 42DelT, C70G, G242C, T358A, 430InsA, 639DelA, C662T, 726InsG, G756A, 786DelG.

[0134] The AmpMU fragment is a double-stranded DNA fragment with a nucleotide sequence as shown in SEQ ID No. 5. The expression product of the mutated Amp gene will affect its Amp resistance phenotype, such as not having Amp resistance.

[0135] 4. Using the AmpMU fragment (SEQ ID No. 5) and AmpWT fragment (SEQ ID No. 4) obtained in the above steps, the substrate AmpAN for the error correction reaction was prepared. 700 ng of each of the AmpMU fragment (20 μL) and AmpWT fragment (14 μL) were pipetted into a 0.2 mL PCR tube according to the equimolar ratio and mixed evenly, and ultrapure water was added to make the total volume about 50 μL; then it was placed in an Eppendorf PCR instrument, incubated at 98 °C for 15 min, and then the power was turned off and it was allowed to cool naturally for 30 min to room temperature (25 °C), obtaining a sample (AmpAN) of DNA molecules with DNA base errors.

[0136] 5. Three treatments were set up, namely the AmpAN sample control without EMC (error correction reaction) treatment, MSE1C treatment, and control enzyme MSES treatment.

[0137] (1) MSE1C treatment (EMC (error correction reaction) treatment): Take 25 μL of the DNA sample from the sample prepared in step 4 for the error correction reaction. The reaction system and procedure are as follows: Pipette 2 μL of 10×EN Buffer I (see Example 1), 2 μL of 10×EN Buffer II (see Example 1), 25 μL of the AmpAN DNA sample prepared in the above step 4, 2 μL of the purified DNA error correction enzyme MSE1C solution in Example 1, react at 37 °C for 0.5 h, and then inactivate at 60 °C for 10 min to obtain the error correction reaction solution.

[0138] (2) Control enzyme MSE treatment (EMC (error correction reaction) treatment): Replace the purified DNA error correction enzyme MSE1C solution in Example 1 in (1) with an equal volume of a control DNA error correction enzyme (commercial enzyme MSES, Company D, VII0602X) solution, and keep the others the same as in (1) to obtain an error correction reaction solution.

[0139] (3) AmpAN sample control without EMC (error correction reaction) treatment: The AmpAN DNA sample prepared in Step 4 above.

[0140] 6. Respectively use the error correction reaction solutions in (1) and (2) in Step 5 above directly as DNA templates, and perform PCR amplification (fusion PCR splicing) using the gradient cooling annealing program with primers Amp-F (5’-atgagtattcaacatttccgtgtcgccctt-3’) and Amp-R (5’-ttaccaatgcttaatcagtgaggcacctatc-3’) to obtain the Amp resistance gene fragment. The reaction system is: 30 μL of DNA template, 2 μL each of Amp-F / Amp-R first and last primers (20 mM), 10 μL of 5×DNA polymerase buffer, 5 μL of dNTP Mix (2.5 mM), and supplement with ultrapure water to 50 μL. The fusion PCR program: 95°C for 2 min; then perform the following 8 cycles: 95°C for 30 sec, annealing temperature (the annealing temperature of the first cycle is 67°C, and it decreases by 1°C per cycle starting from 67°C) for 30 sec, 72°C for 45 sec; then perform the following 20 cycles: 95°C for 30 sec; 58°C for 30 sec; 72°C for 45 sec; 72°C for 10 min. Recover the Amp resistance gene fragment by gel electrophoresis and cutting the gel, and measure the OD260nm / 280nm of the purified sample to determine the DNA concentration.

[0141] 7. Gibson assembly: Respectively perform Gibson assembly ligation reactions on the control reaction solution without EMC treatment (AmpAN sample) in (3) of Step 5 and the Amp resistance gene fragment prepared in Step 6 with the p15C-Vec vector fragment prepared in Step 1 according to a molar ratio of 2:1, add 10 μL of 2×Gibson enzyme mixture, the reaction system is 20 μL, and the reaction program is to react at 50°C for 45 min to obtain a Gibson assembly reaction solution.

[0142] 8. Transform 20 μL of the Gibson assembly reaction mixture into DH5α competent cells (100 μL) according to the conventional molecular cloning procedure. Add 800 μL of LB medium and incubate at 37 °C with shaking at 150 rpm for 1 h. After centrifugation, remove a portion of the supernatant, resuspend the cell pellet with the remaining supernatant, and pipette an equal volume (300 μL) of the resuspended culture onto a Kan+ monoclonal antibody plate (LB solid medium containing 50 μg / mL kanamycin) and a Kan+ / Amp+ double antibody plate (LB solid medium containing 50 μg / mL kanamycin and 50 μg / mL ampicillin), and incubate at 37 °C overnight (12 - 16 h).

[0143] 9. Count the single colonies grown on the overnight culture plates and calculate the ratio of the number of colonies grown on the Kan+ / Amp+ double antibody plate to the number of colonies grown on the Kan+ monoclonal antibody plate as reference data. At the same time, pick single colonies grown on the Kan+ monoclonal antibody plate medium (pick 19 single colonies for the AmpAN treatment) and inoculate them into 4 mL of Kan+ resistant LB liquid medium (liquid LB medium containing 50 μg / mL kanamycin), and incubate at 37 °C with shaking at 220 rpm overnight. The experiment is repeated 3 times, and 19 single colonies are picked each time from the Kan+ monoclonal antibody plate of the untreated AmpAN sample control.

[0144] 10. Centrifuge to collect the overnight culture, extract the plasmid according to the recommended procedure. After detecting plasmid bands by electrophoresis, send the sample for Sanger sequencing. The sequencing primers used are Amp-SeqF (5’-gagacaataaccctgataaatgcttca-3’) and Amp-SeqR (5’-ctgatgtccggcggtgcgtatatatga-3’). Analyze the proportion of correct sequences in the plasmid sequencing results of the error correction reaction-treated samples and the untreated control samples, and use this as the basis for evaluating the effect of the error correction. For the untreated AmpAN sample control, among the 19 single colonies picked from the tested Kan+ monoclonal antibody plate, the sequences of 14 single colonies are correct (i.e., containing the AmpWT fragment shown in SEQ ID No.4, the same below).

[0145] The results showed that the DNA error-correcting enzyme MSE1C exhibited good error-correcting performance. Under the test conditions, the proportion of correct sequences in the DNA substrate fragments treated by the error-correcting reaction of MSE1C increased significantly: in the treatment with MSE1C, among the 17 single colonies growing on the Kan+ monoclonal antibody plate media tested, the sequences of 16 single colonies were correct. After being treated with the DNA error-correcting enzyme MSE1C (EMC-PCR), the proportion of correct sequences was 16 / 17 (94.12%); for the control DNA error-correcting enzyme MSES, the proportion of correct sequences after error-correcting (EMC-PCR) treatment was 0 / 18 (0%), while the proportion of correct sequences in the control sample without error-correcting reaction treatment was 14 / 19 (73.68%). It was proved that the treatment with the error-correcting enzyme MSE1C under the test conditions could increase the proportion of correct sequences after the error-correcting reaction, while the use of the control DNA error-correcting enzyme MSES did not show a significant error-correcting treatment effect. Thus, it was concluded that the error-correcting enzyme MSE1C could be used for the error correction of artificially synthesized DNA base errors or the detection of DNA base errors / mutations.

[0146] Table 1. Proportion of completely correct sequences in the test samples after being treated with different types of MSE for error correction (EMC-PCR)

[0147]

[0148] In Table 1, the denominator in the proportion of correct sequences is the number of single colonies picked from N Kan+ monoclonal antibody plate media in step (9), and the numerator is the number of single colonies with correct sequences.

[0149] Example 3. Application of immobilized DNA error-correcting enzyme MSE1C in error correction of artificially synthesized DNA base errors

[0150] The present invention further verified the application of MSE1C in error correction of DNA base errors in different application scenarios. In this example, an immobilization application scheme of the DNA error-correcting enzyme MSE1C based on the CBM3-cellulose affinity was disclosed. The filling amount of microcrystalline cellulose in the immobilized enzyme adsorption column and the immobilized enzyme reaction system were determined through experiments. The experimental steps are described in detail as follows:

[0151] 1. Weigh 0.02 g of microcrystalline cellulose into a 2 mL EP tube, and add 250 μl each of the working buffers 10×EN-Buffer I and 10×EN-Buffer II of the DNA error-correcting enzyme MSE1C (see Example 1), and mix well.

[0152] 2. Add 2 μL, 4 μL, 6 μL, 8 μL, and 10 μL of the DNA error-correcting enzyme MSE1C solution (see Example 1) to the microcrystalline cellulose suspension in step 1 above on ice. After gently mixing, place it on ice for 15 min. Then transfer it to a hollow column with a glass fiber membrane at the bottom; let it stand still on ice for 5 min and then centrifuge at 300 rpm and 4 °C for 1 min to remove the supernatant.

[0153] 3. Then add 25 μL or 50 μL of the error-correcting reaction substrate AmpAN prepared by the method described in Example 2 and supplement an appropriate amount of MSE1C enzyme working buffer (0 or 25 μL, the preparation system is shown in Table 2). Then place the immobilized enzyme column (obtained in step 2) into a 1.5 mL EP tube and incubate at 37 °C for 1 h.

[0154] Table 2. Immobilized enzyme reaction system

[0155]

[0156] 4. After the reaction is completed, centrifuge at 8000 rpm for 1 min to collect the reaction solution. Take 25 μL of the reaction solution as a template and perform fusion PCR amplification to prepare the full-length Amp resistance gene fragment according to a 50 μL system, with two parallels amplified for each sample. After gel electrophoresis of the obtained PCR products, cut the gel to recover the target fragment and measure the OD260 / 280 of the obtained target fragment to determine the DNA concentration.

[0157] 5. Prepare the p15C-Vec vector fragment by the method described in Example 2 and measure the OD260 / 280 of the prepared vector fragment to determine the DNA concentration.

[0158] 6. Take the Amp resistance gene fragment prepared in step 4 and the p15C-Vec vector fragment according to a molar ratio of fragment to vector of 2:1, and perform Gibson assembly ligation reaction. Add 10 μL of 2×Gibson enzyme mixture, the reaction system is 20 μL, and the reaction program is to react at 50 °C for 45 min. At the same time, set a control by replacing the Amp resistance gene fragment prepared in step 4 with an AmpAN sample without error correction treatment and connecting it to the p15C-Vec vector fragment.

[0159] 7. Transform the 20 μL Gibson assembly reaction solution into DH5a competent cells (100 μL) according to the conventional molecular cloning procedure, add 800 μL of LB medium, and resuscitate and culture at 37 °C and 150 rpm for 1 h. After centrifugation, remove part of the supernatant, resuspend the cells with the remaining part, and pipette an equal amount (300 μL) of the resuscitated culture and spread it on Kan+ monoclonal antibody plates and Kan+ / Amp+ double antibody LB plate media respectively, and culture overnight (12 - 16 h) at 37 °C.

[0160] 8. Count the single colonies growing on the overnight culture plates, and calculate the ratio of the number of colonies growing on the Kan+ / Amp+ double-antibody plates to the number of colonies growing on the Kan+ single-antibody plates as reference data. At the same time, pick single colonies (such as 10 - 20) growing on the Kan+ single-antibody LB plate medium and inoculate them into 4 mL of Kan+-resistant LB liquid medium, and culture overnight at 37°C with 220 rpm.

[0161] 9. Centrifuge to collect the overnight culture, extract the plasmid according to the recommended procedure. After detecting plasmid bands by electrophoresis, send it for Sanger sequencing. The sequencing primers used are Amp-SeqF / R (see Example 2). Analyze the proportion of correct sequences in the sequencing results of the plasmid of the sample treated with the error correction reaction and the untreated control sample (directly connecting the untreated AmpAN to the p15C-Vec vector fragment), and use this as the basis for evaluating the effect of the error correction.

[0162] 10. In this example, the error correction performance of the immobilization system of MSE1C was tested according to the above method. Weak bands could be cut and recovered for samples 3# and 4# (corresponding to the treatment groups with IDs 3 and 4 in Table 2), and there were no obvious bands in the PCR amplification of the remaining treatment groups. After the recovered products were ligated and transformed, the number of growing colonies was counted. Among them, the number of colonies growing on the Amp / Kan double-resistant plate for sample 3# was 37, while the number of colonies growing on the Kan single-antibody plate was 74; the number of colonies growing on the Amp / Kan double-resistant plate for sample 4# was 23, while the number of colonies growing on the Kan single-antibody plate was 125. For each sample, 20 colonies (single colonies growing on the Kan single-antibody plate) were picked for culturing and plasmid extraction and sequencing. Excluding the samples with failed sequencing, the proportion of completely correct sequences (i.e., containing the AmpWT fragment shown in SEQ ID No. 4, the same below) in the treated sample of 3# was 12 / 15 (80%), and the proportion of correct sequences in the treated sample of 4# was 8 / 8 (100%). This result indicates that the error correction effect of the immobilized MSE1C enzyme treated with the 4# reaction system is better. The study also found that under the test conditions, the greater the addition amount of MSE1C enzyme in the microcrystalline cellulose packed column, the more beneficial it is to improve the correct rate.

[0163] The results of this example show that good results can also be obtained by applying the immobilized form of MSE1C enzyme for DNA base error correction, which can be used for the error correction of artificially synthesized DNA bases or the detection of DNA base errors / mutations.

[0164] The above has described the present invention in detail. For those skilled in the art, without departing from the gist and scope of the present invention and without the need for unnecessary experiments, the present invention can be implemented within a relatively wide range under equivalent parameters, concentrations, and conditions. Although specific embodiments of the present invention are given, it should be understood that the present invention can be further improved. In short, in accordance with the principle of the present invention, this application intends to cover any modifications, uses, or improvements to the present invention, including those that depart from the scope disclosed in this application but are made by using conventional techniques known in the art. Some basic features can be applied within the scope of the following appended claims.

Claims

1. A fusion protein, characterized in that: The fusion protein contains a base mismatch-specific endonuclease and a domain that helps the enzyme protein to be affinity adsorbed to a matrix carrier or promotes the expression of the enzyme protein.

2. The fusion protein according to claim 1, wherein: The base mismatch-specific endonuclease is Escherichia coli T7 phage endonuclease I; and / or The domain that helps the enzyme protein to be affinity adsorbed to a specific matrix carrier or promotes the expression of the enzyme protein is the cellulose-specific binding domain CBM3.

3. The fusion protein according to claim 2, wherein: The cellulose-specific binding domain CBM3 is any one of the following: (A1) A protein with an amino acid sequence at positions 169-329 of SEQ ID No. 2; (A2) A protein obtained by substituting and / or deleting and / or adding amino acid residues to the protein defined in (A1), having more than 75% identity with the protein defined in (A1) and having the same function; and / or The Escherichia coli T7 phage endonuclease I is any one of the following: (B1) A protein with an amino acid sequence at positions 1-149 of SEQ ID No. 2; (B2) A protein obtained by substituting and / or deleting and / or adding amino acid residues to the protein defined in (B1), having more than 75% identity with the protein defined in (B1) and having the same function.

4. The fusion protein according to any one of claims 1-3, characterized in that: The fusion protein is any one of the following: (C1) A protein with an amino acid sequence of SEQ ID No. 2; (C2) A fusion protein obtained by fusing a protein tag to the carboxyl terminus or / and amino terminus of the protein in (C1); (C3) A protein obtained by substituting and / or deleting and / or adding amino acid residues to the protein in (C2), having more than 75% identity with the protein in (C1) or (C2) and having the same function.

5. A biomaterial, characterized in that, The biological material is any one of the following: (D1) A nucleic acid molecule encoding the fusion protein according to any one of claims 1-4; (D2) An expression cassette containing the nucleic acid molecule in (D1); (D3) A recombinant vector containing the nucleic acid molecule in (D1); (D4) A recombinant vector containing the expression cassette in (D2); (D5) A recombinant microorganism containing the nucleic acid molecule in (D1); (D6) A recombinant microorganism containing the expression cassette in (D2); (D7) A recombinant microorganism containing the recombinant vector in (D3); (D8) A recombinant microorganism containing the recombinant vector in (D4); (D9) A transgenic plant cell line containing the nucleic acid molecule in (D1) or a transgenic plant cell line containing the expression cassette in (D2); (D10) A transgenic plant tissue containing the nucleic acid molecule in (D1) or a transgenic plant tissue containing the expression cassette in (D2); (D11) A transgenic plant organ containing the nucleic acid molecule in (D1) or a transgenic plant organ containing the expression cassette in (D2); (D12) A transgenic animal cell line containing the nucleic acid molecule in (D1) or a transgenic animal cell line containing the expression cassette in (D2); (D13) A transgenic animal tissue containing the nucleic acid molecule in (D1) or a transgenic animal tissue containing the expression cassette in (D2); (D14) A transgenic animal organ containing the nucleic acid molecule described in (D1) or a transgenic animal organ containing the expression cassette described in (D2).

6. The biomaterial according to claim 5, wherein: (D1) The nucleic acid molecule described in (D1) is any one of the following: (d1) A nucleic acid molecule whose coding sequence is positions 5071-6060 of SEQ ID No.1; (d2) A nucleic acid molecule having more than 75% identity with the DNA molecule defined in (d1) and encoding the fusion protein.

7. A method for preparing the fusion protein according to any one of claims 1-4, comprising: Express the coding gene of the fusion protein in an organism to obtain the fusion protein.

8. The method according to claim 7, characterized in that: The organism is a microorganism; the expression includes introducing the coding gene of the fusion protein into a recipient microorganism to obtain a recombinant microorganism expressing the coding gene of the fusion protein, culturing the recombinant microorganism, and expressing to obtain the fusion protein.

9. An immobilized enzyme, characterized in that: The enzymatic active component in the immobilized enzyme is the fusion protein described in any one of claims 1-4.

10. Any one of the following applications: (E1) The application of the fusion protein described in any one of claims 1-4 or the immobilized enzyme described in claim 9 in DNA error correction, or in the preparation of a DNA error correction enzyme preparation; (E2) The application of the biological material described in claim 5 or 6 or the immobilized enzyme described in claim 9 in the preparation of a DNA error correction enzyme preparation; (E3) The application of the method described in any one of claims 7-9 in the preparation of a DNA error correction enzyme preparation; (E4) The application of the cellulose-specific binding domain CBM3 described in claim 2 or 3 in enhancing the DNA error correction efficiency of a base mismatch-specific endonuclease or in the preparation of a product for enhancing the DNA error correction efficiency of a base mismatch-specific endonuclease; (E5) The application of the nucleic acid molecule encoding the cellulose-specific binding domain CBM3 described in claim 2 or 3 in the preparation of a product for enhancing the DNA error correction efficiency of a base mismatch-specific endonuclease.