Protein engineering methods combining single point saturation mutagenesis and multiple site directed combinatorial mutagenesis

By employing single-point saturation mutation scanning and multi-point targeted combined mutation methods, the problem of low efficiency in fluorescent protein modification in existing protein engineering has been solved, achieving efficient modification of fluorescent proteins and enhancing fluorescence intensity. This method is applicable to the targeted modification of various proteins.

CN116265485BActive Publication Date: 2026-04-21SHENZHEN INST OF ADVANCED TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN INST OF ADVANCED TECH
Filing Date
2021-12-17
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing protein engineering methods for modifying fluorescent proteins suffer from problems such as high computational resource dependence, low design success rate, low efficiency of error-prone PCR mutation, insufficient diversity of mutation directions, and difficulty in full-length saturation mutation, resulting in limited improvement in fluorescence intensity, especially in depth imaging where further improvements are still needed.

Method used

A single-amino acid site saturation mutation library of the target protein was constructed using single-point saturation mutation scanning and multi-point directional combination mutation methods. The correspondence between the mutant sequence and the target characteristics was determined by flow cytometry sorting and high-throughput sequencing. A multi-point directional combination mutant gene library was synthesized, and fluorescent proteins were optimized through multiple rounds of enrichment, sorting and screening.

Benefits of technology

It significantly improves the coverage of amino acid sites and mutation directions, enhances the fluorescence intensity and screening efficiency of fluorescent proteins, shortens the experimental cycle and reduces costs, and is suitable for the targeted modification of a variety of proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116265485B_ABST
    Figure CN116265485B_ABST
Patent Text Reader

Abstract

This invention provides a protein engineering method combining single-point saturation mutation scanning and multi-point targeted combined mutation. The method includes: constructing a single-amino acid site saturation mutation library of the target protein, and scanning and analyzing the single-point mutations and their corresponding target characteristics to obtain the correspondence between the mutant sequence and the target characteristics; based on the correspondence between the mutant sequence and the target characteristics, selecting the amino acid sites and amino acid mutations of the dominant mutants of the target characteristics, synthesizing a multi-point targeted combined mutant gene library, further constructing a multi-point targeted combined mutant expression library, transforming it into an expression host to induce the expression of the multi-point combined mutant target protein, and screening for target protein mutants. The dominant mutants obtained by the method of this invention can cover more amino acid sites and mutation directions of the target protein, thus making the selection of modification targets more scientific and rational, and greatly improving the screening efficiency of dominant mutants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a protein engineering method, specifically a protein engineering method that combines single-point saturation mutation scanning and multi-point directed combination mutation, belonging to the field of protein modification technology. Background Technology

[0002] Protein engineering is a technique that uses molecular biology methods to modify proteins to obtain mutants of target proteins. Existing protein engineering methods are mainly divided into rational design, irrational design, and semi-rational design.

[0003] Rational design, based on existing understanding of protein structure, uses computer technology to simulate the evolutionary trajectory of proteins in nature and simulate mutations. By predicting the impact of mutations at specific sites on protein stability, folding, and substrate binding, target mutants are screened. Computer-aided rational design is highly dependent on computational resources, and due to insufficient understanding of the protein sequence-structure-function relationship, the designed new protein structures and stability are relatively poor, resulting in a low success rate.

[0004] Irrational design does not require in-depth knowledge of protein structure and mechanisms; it only requires simulating natural evolution by constructing and screening libraries of random mutations and fragment recombination. A widely used method for constructing mutant libraries in irrational design is error-prone PCR (epPCR). Error-prone PCR technology increases the random mismatch rate of bases by altering the reaction conditions of the PCR system or using low-fidelity DNA polymerases, thereby causing multi-point mutations and producing mutant libraries with sequence diversity. However, error-prone PCR libraries have the following drawbacks: due to the base bias of polymerases (usually AG>TC), low mutation efficiency, and lack of subsequent mutations, it is difficult to saturate specific amino acid residues with all 19 other types of mutations, usually only 5-6. The diversity of mutation directions is low, and for proteins with short amino acid sequences, such as fluorescent proteins, the overall mutation efficiency is low and mutation accumulation is difficult. Because protein sequence space is enormous, for example, for a protein with a length of 200 amino acids, its sequence space can reach 10^6. 260 Theoretically, it is impossible to experimentally construct and screen full-length saturated mutant libraries.

[0005] On the other hand, since the discovery of Green Fluorescent Protein (GFP) in the bioluminescent jellyfish Victoria in the 1960s, fluorescent proteins have rapidly become important tools in biological research, widely used in protein molecular labeling, cell tracking, in vivo tissue imaging, biological screening, and biosensors. GFP contains 238 amino acid residues, with its 65th-67th amino acid residues (Ser65-Tyr66-Gly67) spontaneously forming a fluorescent chromophore, which can be excited by ultraviolet or blue light to produce green fluorescence. Using GFP as a blueprint, by mining GFP homologous proteins or synthesizing mutants using protein engineering methods, its application can be further extended to the orange, red, and far-infrared spectral regions. Furthermore, to achieve live-cell imaging in anaerobic environments, various oxygen-independent fluorescent proteins have been developed, such as flavin-mononucleotide-based fluorescent proteins (FbFP). Although fluorescent proteins are currently diverse and easy to use, various shortcomings still exist in practical applications, especially in depth imaging where fluorescence intensity still needs further improvement. Summary of the Invention

[0006] One object of the present invention is to provide an improved protein engineering method for targeted modification of protein properties to obtain a target protein mutant.

[0007] Another object of the present invention is to provide the obtained protein mutant.

[0008] Specifically, this invention provides a protein engineering method, which involves targeted modification of a protein to obtain a target protein mutant. The method includes:

[0009] Step 1: Construct a single-amino acid site saturation mutant library of the target protein (referred to as a single-site saturation mutant library), and scan and analyze the single-amino acid site mutations of the target protein and the corresponding target characteristics to obtain the correspondence between the mutant sequence and the target characteristics;

[0010] Step 2: Based on the correspondence between mutant sequences and target characteristics, select the amino acid sites and amino acid mutations of the dominant mutants of the target characteristics, and synthesize a multi-point directed combination mutant gene library;

[0011] Step 3: Using the obtained multi-point directed combination mutant gene library as a template, construct a multi-point directed combination mutant expression library;

[0012] Step 4: Transform the multi-point targeted combinatorial mutant expression library into the expression host to induce the expression of the multi-point combinatorial mutant target protein;

[0013] Step 5: Enrich and sort cells expressing multi-point combination mutant target proteins to screen for target protein mutants.

[0014] According to a specific embodiment of the present invention, in the protein engineering method of the present invention, the target protein is a protein that can couple its target characteristics with the intensity of fluorescence signal.

[0015] According to a specific embodiment of the present invention, in the protein engineering method of the present invention, the target characteristic is selected from one or more of the following protein properties:

[0016] Fluorescence intensity; and / or

[0017] Protein activity, such as transcription factor activity and / or protein-protein interaction activity.

[0018] In this invention, the target protein and its properties can be any protein capable of coupling protein activity with fluorescence signal intensity. The target protein can include various fluorescent proteins, as well as various proteins whose mutants can be coupled with fluorescence intensity. For example, directed evolution of transcription factor activity can be performed by utilizing the higher expression level of the fluorescent protein corresponding to stronger transcription factor activity; or directed evolution of protein-protein interactions can be performed by utilizing coupling with fluorescent proteins to increase fluorescence intensity through stronger protein-protein interactions. That is, the target characteristic described in this invention can be either direct fluorescence signal intensity or indirect protein activity that can be converted into fluorescence signal intensity.

[0019] In this invention, the "target characteristic advantage mutant" refers to a mutant that has an advantage in terms of the target characteristic. Preferably, it is a mutant that has a strong advantage in the target characteristic compared to the unmodified protein. In this invention, the selection of the target characteristic advantage mutant can be based on the diversity of multi-point targeted combination mutant libraries.

[0020] According to a specific embodiment of the present invention, the process of constructing a single-amino acid site saturation mutant library of the target protein in the protein engineering method of the present invention includes:

[0021] Using the gene sequence of the target protein (here referring to the protein encoding the protein to be modified) as a template, construct the target protein expression plasmid;

[0022] Using the target protein expression plasmid as a template, a single amino acid site saturation mutant library of the target protein was constructed.

[0023] In this invention, the synthesis of the target protein gene sequence can be accomplished through commercial gene synthesis technology services. Codon optimization can be performed on the gene sequence encoding the target protein.

[0024] In this invention, the method for constructing the target protein expression plasmid can employ conventional methods in the art. A preferred method for constructing the target protein expression plasmid in this invention includes:

[0025] The expression plasmid vector was subjected to double enzyme digestion and enzyme digestion product purification. PCR was performed using the gene sequence encoding the protein to be modified as a template and primers that fused the expression plasmid vector and the target gene specific sequence, and the PCR product was purified.

[0026] The enzyme digestion and PCR purification products were assembled using Gibson, and the assembled products were transformed into expression host cells for culture. Plasmids were then extracted to obtain the target protein expression plasmid.

[0027] In this invention, the process of constructing a single-point saturation mutant library may include: designing degenerate oligonucleotides to introduce all possible codons at selected amino acid sites; performing PCR amplification of the target protein expression plasmid using degenerate oligonucleotide primers; purifying the PCR product and performing Gibson assembly; transforming the assembly product into expression host cells for culture; extracting the plasmid to obtain the target protein single-point saturation mutant plasmid, i.e., the single-point saturation mutant library. The amino acid sites include all amino acid sites except the start codon. The degenerate oligonucleotide is NNN or NNK (N refers to A, T, C, or G, and K refers to T or G), preferably NNK, and preferably included in one of the target gene-specific pre- or post-primers. Preferably, the PCR amplification of the target protein expression plasmid is a two-stage PCR, one stage using a target gene-specific pre-primer and a vector replicon-specific post-primer, and the other stage using a target gene-specific post-primer and a vector replicon-specific pre-primer.

[0028] According to a specific embodiment of the present invention, in step one of the protein engineering methods of the present invention, the correspondence between the mutant sequence and the target characteristic can be obtained by one or more of the following methods:

[0029] Method 1:

[0030] A single-amino acid site saturation mutation library of the target protein was constructed, and a single-point saturation mutation scanning library of the target protein was constructed accordingly.

[0031] The single-point saturation mutant scanning library was transformed into the expression host to induce the expression of the target protein, and cells expressing the single-point mutant of the target protein were obtained.

[0032] Cells that were induced to express single-point mutants of the target protein were sorted to obtain multiple groups of cell samples;

[0033] Each group of cells from the sorted cell samples was used as a template for multiple rounds of PCR amplification and purification.

[0034] The PCR purification products of multiple cell samples were mixed to obtain an amplicon sequencing library.

[0035] The obtained amplicon sequencing library was subjected to high-throughput sequencing to obtain sequencing data, and the target characteristics of each mutant were further estimated to obtain the correspondence between mutant sequences and target characteristics.

[0036] Method 2:

[0037] A single-amino acid site saturated mutant library of the target protein was constructed, transformed into an expression host, and single clones were selected to induce expression of the target protein. The target characteristics were detected, and the sequence of each mutant was determined using pooled sequencing technology to obtain the correspondence between the mutant sequence and the target characteristics.

[0038] Method 1 of the present invention involves mixing equal amounts of single-point saturated mutant libraries into a single-point saturated mutant scanning library, performing high-throughput transformation and collecting transformants, inducing protein expression, and using cell sorting and sequencing methods to rapidly estimate the target-specific intensity of each mutant, thereby obtaining the correspondence between mutant sequences and target-specific information.

[0039] The second method of this invention involves transforming single-site saturated mutant libraries at various amino acid sites into the expression host, selecting single clones to induce target protein expression, performing high-throughput fluorescence detection on each single-clone bacterial culture, and using pooled sequencing technology to determine the sequence of each mutant. Compared with the first method, this method has a larger experimental workload, longer cycle, and higher cost. However, this method does not require estimation of the target characteristics corresponding to each mutant through biostatistical analysis.

[0040] According to a specific embodiment of the present invention, in the protein engineering method of the present invention, the process of constructing a single-point saturation mutation scanning library of the target protein in method one above includes:

[0041] Multiple single-point saturation mutation libraries are mixed to obtain a single-point saturation mutation scanning library;

[0042] According to a specific embodiment of the present invention, in the protein engineering method of the present invention, the process of mixing multiple single-point saturation mutation libraries to obtain a single-point saturation mutation scanning library in method one above includes:

[0043] Quantification of single-point saturated mutant libraries; fluorescence quantification is preferred.

[0044] Multiple single-point saturation mutation libraries were mixed in equal amounts to obtain a single-point saturation mutation scanning library.

[0045] In this invention, the single-point saturation mutation scanning library contains a single-point saturation mutation library of all amino acid sites of the target protein except the start codon.

[0046] According to a specific embodiment of the present invention, in the protein engineering method of the present invention, the process of sorting cells that are induced to express single-point mutants of the target protein in method one above includes:

[0047] Cells expressing single-point mutants of the target protein are sorted according to their target characteristics to obtain multiple groups of cell samples; preferably, 4-16 groups of cell samples are sorted.

[0048] In some specific embodiments of the present invention, the cell sorting method is fluorescence-activated cell sorting, which uses a flow cytometer to sort cells that induce the expression of single-point mutants of the target protein according to fluorescence intensity, thereby obtaining multiple groups of cell samples; wherein, preferably, the number of sorting is 4-16, that is, 4-16 groups of cell samples are sorted.

[0049] According to a specific embodiment of the present invention, in the protein engineering method of the present invention, after sorting the cells that induce the expression of the target protein single-point mutant, the method further includes a process of lysing the sorted cells; preferably, the lysing process is performed by repeated freeze-thaw cycles.

[0050] According to a specific embodiment of the present invention, in the protein engineering method of the present invention, in the process of performing multiple rounds of PCR amplification and purification using each group of cells from multiple groups of cell samples obtained after sorting as templates, the multiple rounds of PCR are two-round PCR amplification methods, wherein:

[0051] The template for the first round of PCR amplification is the cell sample obtained after sorting. The fusion sequence of the adapter sequence, spacer sequence and expression vector-specific sequence obtained by sequencing on a high-throughput sequencing platform is used as primers for PCR amplification and purification. Preferably, the primers for the first round of PCR have a spacer sequence of 0-8 nucleotides added between the adapter sequence obtained by sequencing on the high-throughput sequencing platform and the target gene-specific sequence.

[0052] The second round of PCR uses the purified product from the first round of PCR as a template and amplifies the adapter sequence, the index, and the fusion sequence of the sequencing adapter sequence using a high-throughput sequencing platform as primers. Preferably, the primers for the second round of PCR use a 6-10 nucleotide sequence as the index sequence.

[0053] According to a specific embodiment of the present invention, in the protein engineering method of the present invention, the process of mixing multiple groups of cell sample PCR purification products to obtain an amplicon sequencing library in method one above includes:

[0054] Quantification of PCR purified products from each group of cell samples was performed; fluorescence quantification was preferred.

[0055] The PCR purification products of multiple cell samples were mixed in equal amounts to obtain the amplicon sequencing library;

[0056] Preferably, the method further includes quantification and / or fragment size detection of the amplicon sequencing library.

[0057] According to a specific embodiment of the present invention, in the protein engineering method of the present invention, the process of estimating the target characteristics of each mutant in method one above includes:

[0058] Estimate the target characteristics of each mutant and obtain the correspondence between the mutant sequence and the target characteristics.

[0059] Mutation determination and data statistics were performed on the high-throughput sequencing data of the obtained amplicon sequencing library to obtain the normalized number distribution of each mutant in each sorting group;

[0060] The target characteristic intensity of each mutant is calculated by weighted average based on the average value or boundary of the target characteristic of each sorting group and the number distribution of each mutant in each sorting group.

[0061] Assuming that the number distribution of each target protein mutant in each sorting group is Gaussian or gamma in logarithmic coordinate system, the target characteristic intensity of each mutant is estimated by the maximum likelihood estimation method.

[0062] According to a specific embodiment of the present invention, in step two of the protein engineering method of the present invention, the process of synthesizing a multi-point directed combination mutant gene library includes:

[0063] Based on the correspondence between mutant sequences and target characteristics, the amino acid sites of mutants with the highest target characteristic intensity are selected as the target sites for the combinatorial mutant library. At the same time, the amino acid mutations with the highest target characteristics in the same amino acid site are selected to synthesize a multi-point targeted combinatorial mutant gene library.

[0064] According to a specific embodiment of the present invention, the process of constructing a multi-point directed combinatorial mutant expression library using the obtained multi-point directed combinatorial mutant gene library as a template in the protein engineering method of the present invention includes:

[0065] Using the obtained multi-point directional combined mutant gene library as a template, the Gibson method was used to assemble the enzyme-digested linearized expression vector and the multi-point directional combined mutant gene library. The assembly product was transformed into expression host cells for culture, and plasmids were extracted to obtain the target protein expression plasmid, i.e., the multi-point directional combined mutant expression library.

[0066] According to a specific embodiment of the present invention, the process of enriching and sorting cells expressing multi-point combination mutant target proteins to screen for target protein mutants in the protein engineering method of the present invention includes:

[0067] Cells with target characteristic advantages are enriched and sorted in a single round or in multiple consecutive rounds;

[0068] Preferably, at least three rounds of sorting are performed: the first round sorts the 5%-10% of cells with the highest target characteristic intensity, and the cells sorted in the first round are analyzed by flow cytometry to sort the 10%-20% of cells with the highest target characteristic intensity; the second round sorts the cells are analyzed by flow cytometry to sort the 20%-30% of cells with the highest target characteristic intensity; the cells sorted in the third round are plated on screening culture plates and cultured overnight.

[0069] After culturing the obtained monoclonal antibodies to the logarithmic growth phase, the target protein was induced to express. The strength of the target characteristics was then detected, or the plasmid was further extracted and sequenced to screen for mutants of the target protein.

[0070] According to a specific embodiment of the present invention, in the protein engineering method of the present invention, the expression host preferably includes Escherichia coli, Saccharomyces cerevisiae, or Pichia pastoris.

[0071] According to a specific embodiment of the present invention, in the protein engineering method of the present invention, the transformation method is preferably an electroconversion method. The preparation of electroconverted competent cells can be carried out using conventional methods in the art.

[0072] According to specific embodiments of the present invention, the induction conditions in the protein engineering method of the present invention can employ conventional methods in the art. In some specific embodiments of the present invention, the preferred induction conditions for the expression host *Escherichia coli* are: culturing the bacterial culture to OD0.05. 600 When the value is 0.4-0.6, use 0.1-1mM IPTG for induction, with an induction temperature of 16-37℃ and an induction time of 6-16h.

[0073] On the other hand, the present invention also provides a protein mutant, which is a target protein mutant obtained by targeted modification of the protein according to the method described in the present invention.

[0074] In some specific embodiments of the present invention, the fluorescent properties of fluorescent proteins are specifically modified. In more specific embodiments, the present invention obtains mutants of one or more of the fluorescent proteins CreiLOV among the S1-S20 mutants shown in Table 3.

[0075] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of the present invention.

[0076] The reagents and raw materials used in the technical solution of this invention are all commercially available.

[0077] In the protein engineering method of this invention, the construction of a single-point saturation mutation scanning library allows for the saturation mutation of each amino acid site into one of the other 19 amino acids. Furthermore, for short-sequence proteins such as fluorescent proteins, it maximizes the coverage of amino acid mutations across the full-length sequence. Compared to error-prone PCR, single-point saturation mutation scanning libraries offer maximized improvements in both amino acid site coverage and amino acid mutation direction coverage for single-point amino acid mutations.

[0078] In the protein engineering method of this invention, flow cytometry sorting and sequencing, along with biostatistical analysis, are used to comprehensively analyze the relationship between protein amino acid mutations and fluorescence intensity based on single-point saturation mutation scanning libraries, guiding the rational selection of combined modification targets. Based on the results of single-point saturation mutation scanning, this invention rationally selects the optimal combination of amino acid sites and directions to construct multi-point directional combined mutation libraries, balancing the number of amino acid targets and mutation directions, thereby improving the efficiency of screening for advantageous mutants.

[0079] Overall, this invention, based on a single-point saturation mutation scanning and multi-point targeted combined mutation screening strategy, yields superior mutants that cover more amino acid sites and mutation directions in the target protein. This allows for a more scientific and rational selection of modification targets, significantly improving the screening efficiency of superior mutants. Simultaneously, combining flow cytometry-based gate sorting sequencing and multi-round enrichment sorting culture strategies greatly increases the throughput and speed of mutant phenotype screening or detection, reducing experimental costs and shortening the experimental cycle. The method of this invention has wide applications and can be easily extended to various proteins that can couple mutants with fluorescence intensity, serving as an effective and universal method for directed protein evolution. Attached Figure Description

[0080] Figure 1 This is a flowchart illustrating a specific embodiment of the protein engineering method of the present invention.

[0081] Figure 2 This is a schematic diagram of the process for constructing the expression plasmid of the target protein in the protein engineering method of the present invention.

[0082] Figure 3 This is a schematic diagram of the process for constructing a single-point saturated mutant library in the protein engineering method of the present invention.

[0083] Figure 4 This is a thermogram of a single-point saturation mutation of CreiLOV and its fluorescence intensity in a specific embodiment of the present invention. The horizontal axis represents the CreiLOV amino acid site, and the vertical axis represents the amino acid mutation at that site.

[0084] Figure 5The structure of a multi-point directed combinatorial mutant gene library obtained in a specific embodiment of the present invention is shown in graphical form. Numbers 1-15 in the figure indicate that the library contains 15 amino acid sites, whose positions, amino acid types, and codons in the wild-type CreiLOV protein are shown as "Amino Acid Sites," "Wild-Type Amino Acids," and "Wild-Type Codons," respectively. AY indicates the amino acid orientation at each amino acid site in the library; the number of amino acid orientations, including wild-type amino acids, is shown in the "Amino Acid Orientation" row.

[0085] Figure 6 This is a schematic diagram of another scheme for obtaining the correspondence between mutant sequences and target characteristics using the protein engineering method of the present invention. Detailed Implementation

[0086] The present invention will be further described below with reference to specific embodiments and accompanying drawings, but this does not limit the invention to the scope of the described embodiments. Experimental methods not specified in the detailed embodiments are performed according to conventional methods and conditions, or according to the instructions of the selected product.

[0087] In some specific embodiments of the present invention, a protein engineering method is provided to directionally modify proteins to obtain target protein mutants. Specifically, this can be performed according to the following steps (which can be combined with...). Figure 1 (For reference only):

[0088] (1) Synthesize the gene sequence of the target protein:

[0089] In this invention, the synthesis of the target protein gene sequence can be completed using commercial gene synthesis technology services. Codon optimization can be performed on the gene sequence encoding the target protein. Codon optimization is a conventional technique in the art, referring to designing the gene sequence of the target protein based on the amino acid sequence of the target protein and targeting the codon bias used by the host expressing the target protein for its amino acid encoding. The expression host is preferably *Escherichia coli* or *Saccharomyces cerevisiae*. Codon optimization methods can utilize commercial tools provided by gene synthesis companies. In this invention, the target protein may include various fluorescent proteins, and may also include various proteins whose mutants can be coupled to the intensity of fluorescent proteins.

[0090] (2) Using the target protein gene sequence as a template, construct the target protein expression plasmid:

[0091] In this process, the method for constructing the target protein expression plasmid can adopt conventional methods in the art, and the present invention preferably includes the following steps:

[0092] The expression plasmid vector was subjected to double enzyme digestion and enzyme digestion product purification. Using the target protein gene sequence as a template, PCR was performed with primers that fused the expression plasmid vector and the target gene specific sequence, and the PCR product was purified. The enzyme digestion and PCR purified products were assembled by Gibson. The assembled product was transformed into E. coli competent cells and cultured on plates. Single clones were picked, cultured in liquid, and the plasmid was extracted and sequenced for verification to obtain the target protein expression plasmid.

[0093] In this invention, low-copy plasmids are preferably selected as the expression plasmid vector. For double digestion, commonly used restriction endonucleases in the multiple cloning site region of the expression plasmid vector are preferred, such as BamHI, XhoI, EcoRI, and HindIII. The preferred double digestion reaction system and conditions are: 2-2.5 μg expression plasmid vector, 5 μL 10× digestion buffer, 2 μL restriction enzyme 1, 2 μL restriction enzyme 2, and ultrapure water to a final volume of 50 μL. The mixture is incubated in a 37°C water bath for 3 hours to overnight. The preferred purification method for the digestion products is to perform agarose gel electrophoresis followed by purification using various commercially available gel extraction kits. The preferred agarose gel electrophoresis conditions are: 1% agarose, 120-170 V, for 20-30 minutes. The preferred kit is the Axygen AxyPrepDNA Gel Extraction Kit, which uses 20-25 μL of ultrapure water for elution. For specific instructions, please refer to the kit's instruction manual.

[0094] In this invention, the preferred method for designing fusion primers is as follows: the upstream 15-20 nt sequence of the endonuclease 1 recognition site in the expression plasmid vector is fused with the start codon of the target gene and a sequence of appropriate length thereafter to form a front primer; the downstream 15-20 nt sequence of the endonuclease 2 recognition site in the expression plasmid vector is fused with the stop codon of the target gene and a sequence of appropriate length thereafter to form a back primer.

[0095] In this invention, PCR can be performed using commercially available high-fidelity polymerase or premixed solution, preferably Takara 2×PrimeSTAR master mix. The preferred PCR reaction system is: 20 μL ultrapure water, 2 μL front primer (10 μM), 2 μL back primer (10 μM), 25 μL 2×PrimeSTAR master mix, 1 μL plasmid DNA (1 ng) containing the target gene sequence, totaling 50 μL. The preferred PCR reaction procedure is: open the heat-sealing lid and preheat to 105℃, (1) denature at 98℃ for 5 min, (2) denature at 98℃ for 15 s, (3) anneal at 50-65℃ for 15 s, (4) extend at 72℃ for 30 s-1 min, wherein steps (2)-(4) are performed for a total of 30-35 cycles, and (5) extend at 72℃ for 5-10 min. PCR purification is preferably performed using a commercially available gel extraction kit, with the Axygen AxyPrep gel extraction kit being a preferred choice. Specific gel extraction methods can be found in the instruction manual of the relevant kit.

[0096] In this invention, the preferred Gibson assembly reaction system and conditions are as follows: 3.75-15 μL of Gibson reaction mixture, 1.25-5 μL of equimolar amounts of enzyme digestion purification product and PCR purification product, totaling 5-20 μL of reaction mixture, incubated at 50°C for 15-60 min. The preferred composition of the Gibson reaction premix is: 320 μL of 5× isothermal reaction buffer, 0.64 μL of 10 U / μL T5 exonuclease, 20 μL of 2 U / μL Phusion polymerase, 160 μL of 40 U / μL Taq ligase, and ultrapure water added to 1.2 mL per 1.2 mL. The preferred 5× isothermal reaction buffer is as follows: per 6 mL, 3 mL of 1M Tris-HCl (pH 7.5), 150 μL of 2M MgCl2, 60 μL of 100mM dGTP, 60 μL of 100mM dATP, 60 μL of 100mM dTTP, 60 μL of 100mM dCTP, 300 μL of 1M DTT, 1.5 g of PEG-8000, 300 μL of 100mM NAD, and ultrapure water to a final volume of 6 mL.

[0097] In this invention, the Gibson assembly product transformation method is preferably a chemical transformation method. The *E. coli* competent cells are preferably commercially available DH5α or Trans5α competent cells. The plasmid extraction method is preferably a commercially available plasmid extraction kit, preferably the Axygen AxyPrep plasmid mini-extraction kit. Specific methods for transformation and plasmid extraction can be found in the instruction manuals of the relevant kits. The plasmid sequencing method is preferably Sanger sequencing, preferably performed using a commercially available Sanger sequencing service.

[0098] According to a specific embodiment of the present invention, the technical process for constructing a target protein expression plasmid using the target protein gene sequence as a template can be found in [reference needed]. Figure 2 As shown.

[0099] (3) Using the target protein expression plasmid as a template, construct a single-point saturation mutant library:

[0100] In this invention, the method for constructing a single-point saturated mutant library involves introducing all possible codons at a single amino acid site using a customized degenerate oligonucleotide primer. The target protein expression plasmid is then amplified by PCR using this degenerate oligonucleotide primer. The single-point saturated mutant library is constructed via Gibson assembly. The assembly product is transformed into competent *E. coli* cells and cultured on plates. The plasmid is then extracted after cell collection. The amino acid site contains all amino acid sites except the start codon. The degenerate oligonucleotide is NNN or NNK (N refers to A, T, C, or G, and K refers to T or G), preferably NNK, and preferably included in one of the target gene-specific pre- or post-primers. The PCR amplification of the target protein expression plasmid is preferably performed in two stages: one stage uses a target gene-specific pre-primer and a vector replicon-specific post-primer, and the other stage uses a target gene-specific post-primer and a vector replicon-specific pre-primer. High-fidelity DNA polymerase or a premix is ​​preferably used for PCR amplification, preferably Takara 2×PrimeSTAR master mix. The PCR reaction system and conditions can be found in step (2). After extracting the single-point saturated mutant library plasmid, it is preferable to also perform sequencing verification of the plasmid. The sequencing method is Sanger sequencing, preferably performed through a commercial Sanger sequencing service.

[0101] According to a specific embodiment of the present invention, the technical process for constructing a single-point saturation mutant library using the target protein expression plasmid as a template can be found in [reference needed]. Figure 3 As shown.

[0102] (4) Perform fluorescence quantification and equal-volume mixing on the single-point saturated mutation library to obtain the single-point saturated mutation scanning library:

[0103] In this invention, the quantification method for the single-point saturated mutant library is preferably fluorescence quantification, and more preferably using the Invitrogen Qubit dsDNA BR detection kit and the Tecan Infinite 200Pro microplate reader. Specific methods for fluorescence quantification using the kit can be found in the kit's instruction manual. The single-point saturated mutant library is mixed in equal volumes. A single-point saturated mutant library refers to a single-point saturated mutant library containing all amino acid sites of the target protein except the start codon.

[0104] (5) The single-point saturation mutation scanning library was converted into the expression host in a high-throughput manner to induce the expression of the target protein:

[0105] In this process of the present invention, the expression host includes *Escherichia coli*, *Saccharomyces cerevisiae*, *Pichia pastoris*, etc., preferably *Escherichia coli*. The transformation method is preferably electrotransformation. The preparation of electrotransformation competent cells is a conventional method in the art, preferably including the following steps: picking a single *E. coli* clone and incubating it overnight at 37°C with shaking at 200 rpm; then transferring 2 mL of the bacterial culture to 200 mL of fresh culture medium and incubating with shaking until OD... 600 The absorbance at 600 nm is approximately 0.6 (about 2-2.5 h). Transfer the bacterial culture to a pre-chilled sterile centrifuge tube on ice, centrifuge at 4000 rpm for 10 min at 4°C, remove the supernatant, wash once with 200 mL of pre-chilled 10% glycerol, then wash twice with 40 mL of pre-chilled 10% glycerol. Resuspend the cell pellet in 2 mL of 10% glycerol and aliquot 50 μL into sterile centrifuge tubes. Quick-freeze in liquid nitrogen and store at -80°C for long-term storage. The preferred electroporation method includes the following steps: Thaw the electroporation competent cells on ice for 5-10 min and pre-chill the electroporation cuvette. Add 1-2 μL of plasmid library and gently mix. Transfer to the electroporation cuvette and perform electroporation at 1.2-2.4 kV. Immediately add 1 mL of LB medium to the cuvette and recover the cells into a 1.5 mL centrifuge tube. Incubate at 37°C with shaking for 1 h, then plate onto LB agar plates containing appropriate antibiotics for overnight culture. Preferably, 0.1 cm spacing electroporation cuvettes are used, and a BioRad Gene Pulser electroporator is preferred. After electroporation of the single-point saturation mutant scanning library, the process preferably also includes transformant collection and OD. 600 Detection. A preferred method for inducing the expression of the target protein is to inoculate and culture bacterial suspensions until OD200 reaches a certain level. 600 The induction temperature is 16-37℃, and the induction time is 6-16h. The induction value is 0.4-0.6.

[0106] (6) Cells expressing single-point mutants of the target protein were sorted by flow cytometry based on fluorescence intensity, and the collected cells were pretreated:

[0107] In this process, the flow cytometer can be a commercially available mainstream flow cytometer, preferably a BDFACSAria III cell sorter. The gate sorting parameters preferably include: 4-16 gates; each gate containing an equal proportion of cells or with a roughly uniform distribution of cells in a logarithmic coordinate system determining the width and boundary of each gate; a total number of cells sorted that is 100-1000 times the library diversity; a 70μm or 85μm nozzle; 4-way sorting; "Purity" or "4-Way Purity" mode; and a cell density of 102. 6 -10 7 / mL, sample temperature maintained at 4℃. The pretreatment method for sorted cells is as follows: transfer sorted cells to centrifuge tubes, centrifuge at 4000-12000g for 3-5min, remove and retain about 10-20μL of supernatant, vortex to mix, repeatedly freeze and thaw 3 times with liquid nitrogen, incubate at 95℃ for 3-5min, and then store at -20℃.

[0108] (7) Using the cells obtained from the sorting process as templates, two rounds of PCR amplification and purification were performed:

[0109] In this invention, the PCR method is a two-round PCR amplification method. The first round of PCR amplification uses sorted cell samples pretreated with cell lysis as templates. A fusion sequence of the adapter sequence, spacer sequence, and expression vector-specific sequence is used as primers for PCR amplification and purification. The first round of PCR primarily enriches and recovers the gene sequences of the target protein mutants contained in each sorted cell sample. The second round of PCR uses the purified product from the first round of PCR as a template and amplifies the fusion sequence of the adapter sequence, index, and sequencing adapter sequence using primers from a high-throughput sequencing platform. The second round of PCR primarily introduces index sequences and high-throughput sequencing platform amplification and sequencing primer sequences labeled with different sorted cell samples, facilitating subsequent mixed-sample high-throughput sequencing. Preferably, the first round of PCR primers include a 0-8 nucleotide spacer sequence between the high-throughput sequencing platform's sequencing adapter sequence and the target gene-specific sequence to reduce the bias of the target gene sequence and improve sequencing quality and output. The second round of PCR primers preferably use a 6-10 nucleotide index sequence, preferably using the index sequence provided by the high-throughput sequencing platform. Two rounds of PCR amplification are preferably performed using high-fidelity DNA polymerase or premixed solution, with Takara 2×PrimeSTAR master mix being the preferred choice. The preferred PCR reaction system for the first round is: 8.5 μL ultrapure water, 1 μL of the first primer (10 μM), 1 μL of the second primer (10 μM), 12.5 μL of 2×PrimeSTAR mastermix, 2 μL of pretreated cell sample, for a total of 25 μL. The preferred PCR reaction procedure for the first round is: open the heat-sealing lid and preheat to 105°C, (1) denature at 98°C for 5 min, (2) denature at 98°C for 15 s, (3) anneal at 50-65°C for 15 s, (4) extend at 72°C for 30 s-1 min, wherein steps (2)-(4) are performed for a total of 15-25 cycles, and (5) extend at 72°C for 5-10 min. The preferred second-round PCR reaction system is as follows: 1 μL of the first primer (10 μM), 1 μL of the second primer (10 μM), 12.5 μL of 2×PrimeSTAR master mix, 5-10 μL of the first-round PCR purified product, and ultrapure water to 25 μL. The preferred second-round PCR reaction procedure is as follows: open the heat-sealing lid and preheat to 105℃, (1) denature at 98℃ for 5 min, (2) denature at 98℃ for 15 s, (3) anneal at 50-65℃ for 15 s, (4) extend at 72℃ for 30 s-1 min, wherein steps (2)-(4) are performed for a total of 8-15 cycles, and (5) extend at 72℃ for 5-10 min. The purification of the two-round PCR amplification products can be carried out using conventional purification methods in the art, preferably using a DNA gel extraction kit or magnetic beads.The DNA gel extraction kit is preferably the Axygen AxyPrep gel extraction kit; specific methods can be found in the kit's instruction manual. The magnetic bead purification method preferably includes the following steps: Add 20 μL of Beckman AMPureXP beads equilibrated at room temperature to the PCR reaction system, mix well, and incubate at room temperature for 10-15 min. Place the tube on a magnetic rack and let it stand for 5 min until the liquid becomes clear, then aspirate the supernatant. Wash twice with 200 μL of freshly prepared 80% ethanol, incubating for 30 s each time at room temperature, and incubate on a magnetic rack until the liquid becomes clear, then aspirate the supernatant. Then, incubate on a magnetic rack for 5-15 min until the magnetic beads are dry, add 20-25 μL of ultrapure water and mix well. Incubate at room temperature for 10 min, then place the PCR tube on a magnetic rack until the liquid becomes clear, and recover 20-25 μL of the PCR product.

[0110] (8) Perform fluorescence quantification and equal-volume mixing on the obtained second-round PCR purification products to obtain the amplicon sequencing library:

[0111] In this invention, the quantification method for the PCR purified products is preferably fluorescence quantification, and more preferably using the Invitrogen Qubit dsDNA BR detection kit and the Tecan Infinite 200Pro microplate reader. Specific methods for fluorescence quantification using the kit can be found in the kit's instruction manual. The PCR purified products are mixed in equal volumes. An amplicon sequencing library refers to a PCR product obtained through two rounds of PCR amplification, containing the sequences of high-throughput sequencing platform-specific amplification primers and sequencing primers at both ends. After obtaining the amplicon sequencing library, the process preferably also includes quantification and fragment size detection of the sequencing library. The quantification method is preferably fluorescence quantification, and the fragment size detection method is preferably capillary electrophoresis, preferably using an Ailent 2100 bioanalyzer or a Bioptic Qsep400 nucleic acid fragment analyzer.

[0112] (9) Perform high-throughput sequencing on the obtained amplicon sequencing library to obtain sequencing data:

[0113] In this process of the present invention, high-throughput sequencing can be a conventional sequencing technology in the art, preferably using various commercial high-throughput sequencers, preferably the Illumina NovaSeq 6000 sequencer, and preferably the selected sequencing read length is 2×250bp paired-end sequencing.

[0114] (10) Mutation determination and data statistics were performed on the obtained sequencing data to obtain the distribution of the number of each mutant in each phylum:

[0115] In this process of the present invention, mutant identification and statistics can be conventional techniques in the art. It is preferable to use open-source high-throughput sequencing data sequence alignment and mutant identification software, and more preferably to use Enrich2 software for sequence alignment and mutant identification. The preferred parameters of this software are: Minimum Quality 20-25, Maximum Mismatches 0, Maximum Mutations 3, Remove Unresolvable Overlaps.

[0116] (11) Based on the average or boundary fluorescence intensity of each phylum, estimate the fluorescence intensity of each mutant using a biostatistical model:

[0117] In this invention, the average fluorescence intensity or boundary of each gate refers to the average fluorescence value of each gate in the fluorescence channel during cell sorting, or the upper and lower boundary fluorescence values ​​of each gate. The biostatistical model preferably assumes that the distribution of fluorescence values ​​of each target protein mutant in each gate is a Gaussian or gamma distribution in a logarithmic coordinate system. The fluorescence intensity estimation method is preferably a simple average estimation method or a maximum likelihood estimation method. The simple average estimation method preferably uses a weighted average of the number of cells of the mutant in each gate and the average fluorescence intensity of that gate. The maximum likelihood estimation method is preferably implemented using the open-source Python toolkit SciPy, and preferably provides the initial expectation and variance based on the simple average estimation.

[0118] (12) Based on the amino acid sites and amino acid mutations of the strongly fluorescent mutants, a multi-site directed combination mutant gene library was synthesized:

[0119] In this invention, the selection method for strong fluorescent mutants involves ranking the estimated fluorescence intensity of each mutant and selecting the amino acid site of the mutant with the highest ranking as the target site for the combined mutant library. Simultaneously, the amino acid mutations with the highest estimated fluorescence intensity at that amino acid site are selected. For example, the top n amino acid sites can be selected, and then the m mutation directions with the highest fluorescence intensity at those n amino acid sites can be determined. Here, n can be any integer from 5 to 30, preferably any integer from 10 to 20, such as 5, 10, 15, 20, etc.; m can be any integer from 5 to 60, preferably any integer from 10 to 30, such as 15, 20, 25, 30, etc. In this invention, m is typically greater than or equal to n and less than or equal to 19n, but is not limited to this. The synthesis method for the multi-point directed combined mutant gene library is gene chip synthesis, preferably completed through commercial gene library synthesis technology services. The two ends of the multi-point directed combined mutant gene library sequence preferably contain adapter sequences required for subsequent expression plasmid library construction.

[0120] (13) Using the obtained multi-point directed combination mutant gene library as a template, construct a multi-point directed combination mutant expression library:

[0121] In this invention, the method for constructing the multi-point directed combination mutant expression library involves double digestion and purification of the expression plasmid vector, assembly of the enzyme-digested linearized expression vector and the multi-point directed combination mutant gene library using the Gibson method, purification and concentration of the assembly product, transformation of the purified and concentrated product into competent E. coli cells using electroporation, overnight plate culture, collection of bacterial cells, and plasmid extraction. The double digestion method and Gibson assembly method for the expression plasmid are described in the aforementioned section "(2) Constructing the target protein expression plasmid using the target protein gene sequence as a template". The Gibson assembly product purification and concentration method is preferably column purification, preferably using the Zymo Research DNA Purification and Concentration Kit. Specific methods for purification and concentration using the kit can be found in the kit's instruction manual. Electroporation preferably uses commercially available, highly efficient electroporation competent cells, particularly Invitrogen MegaX DH10B T1 or NEB 10-beta electroporation competent cells. Specific methods for electroporation using these cells can be found in the instructions for use of the reagents. The number of transformants obtained by electroporation is preferably at least 6 times, and more preferably at least 10 times, greater than the diversity of the multi-site directed combinatorial mutant library.

[0122] (14) High-throughput conversion of multi-point targeted combined mutant expression libraries into expression hosts to induce expression of target proteins:

[0123] In this process of the present invention, the high-throughput conversion method can be electroconversion, and the number of transformants is preferably more than 6 times the diversity of the multi-point directed combination mutant library, and more preferably more than 10 times. The expression host includes Escherichia coli, Saccharomyces cerevisiae, Pichia pastoris, etc., preferably Escherichia coli. For the specific methods of electroconversion and target protein induction expression, please refer to the above "(5) high-throughput conversion of the single-point saturation mutant scanning library into the expression host to induce target protein expression".

[0124] (15) Cells expressing multi-point combination mutant target proteins were continuously enriched and sorted in multiple rounds using flow cytometry based on fluorescence intensity. The collected cells were then plated onto screening plates for culture to obtain single clones.

[0125] In this process of the present invention, sorting can be performed using commercially available mainstream flow cytometers, preferably the BDFACSAria III cell sorter. The sorting method involves continuous multi-round enrichment and sorting of strongly fluorescent cells, preferably as follows: the first round sorts the 5%-10% of cells with the highest fluorescence intensity; the first round sorts the cells by flow cytometry and sorts the 10%-20% of cells with the highest fluorescence intensity; the second round sorts the cells by flow cytometry and sorts the 20%-30% of cells with the highest fluorescence intensity; and the third round sorted cells are plated on screening culture plates and cultured overnight.

[0126] (16) The obtained monoclonal antibodies were cultured in liquid medium with shaking until saturation, then inoculated into fresh medium and cultured with shaking until the logarithmic phase, after which the expression of the target protein was induced:

[0127] In this process of the present invention, the specific methods for inoculating, culturing and inducing expression of single-clone bacterial culture can be found in the aforementioned “(5) High-throughput conversion of single-point saturation mutation scanning library into expression host to induce expression of target protein”.

[0128] (17) Detect the fluorescence intensity of cells expressing multi-point combination mutants using flow cytometry or an enzyme-linked immunosorbent assay (ELISA) reader:

[0129] In this process, the flow cytometer can be any flow cytometer, preferably the BDFACSCelesta flow cytometer. Specific methods for using the flow cytometer can be found in the instrument's instruction manual. The microplate reader can be any microplate reader, preferably the Tecan Infinite 200Pro microplate reader. The method for detecting cell fluorescence intensity using the microplate reader preferably involves simultaneously detecting the absorbance of the cell culture medium at a wavelength of 600 nm and the fluorescence intensity of the corresponding fluorescent protein. The absorbance value is used to calibrate the fluorescence intensity of different clones. Specific instrument usage methods can be found in the instrument's instruction manual.

[0130] (18) After extracting plasmids containing strongly fluorescent mutant cells, sequencing was performed to obtain the gene sequence of the strongly fluorescent mutant:

[0131] In this process, plasmid extraction can be performed using conventional extraction methods in the art, preferably using a commercially available plasmid extraction kit, and more preferably using the Axygen AxyPrep plasmid mini-extraction kit. Specific plasmid extraction methods can be found in the instruction manual of the relevant kit. Plasmid sequencing can be performed using conventional sequencing methods in the art, and more preferably using commercial Sanger sequencing services.

[0132] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain various preferred embodiments of the present invention.

[0133] Example 1

[0134] This embodiment provides a method for directed evolution of the CreiLOV fluorescent protein that combines single-point saturation mutation scanning with multi-point directed combined mutation.

[0135] The target protein, CreiLOV, is an oxygen-independent fluorescent protein derived from the LOV1 domain of the Phot protein from Chlamydomonas reinhardtii. It has an amino acid sequence length of 119 aa and a nucleic acid sequence length of 360 bp. The amino acid sequence of the wild-type CreiLOV fluorescent protein was obtained by searching the NCBI database. The amino acid sequence is: MAGLRHTFVVADATLPDCPLVYASEGFYAMTGYGPDEVLGHNARFLQGEGTDPKEVQKIRDAIKKGEACSVRLLNYRKDGTPFWNLLTVTPIKTPDGRVSKFVGVQVDVTSKTEGKALA* (SEQ ID NO:2).

[0136] The CreiLOV fluorescent protein directed evolution method combining single-point saturation mutation scanning and multi-point directional combined mutation in this embodiment is specifically performed as follows.

[0137] 1. Synthesize the CreiLOV codon to optimize the gene sequence.

[0138] The synthesis of the CreiLOV codon-optimized gene sequence for E. coli was completed using gene synthesis technology services from Genewiz Biotechnology Co., Ltd., and the gene sequence was delivered in the form of pUC57-CreiLOV plasmid.

[0139] The optimized codon gene sequence of CreiLOV E. coli is as follows:

[0140] ATGGCTGGCCTTCGCCATACATTTGTTGTTGCTGATGCTACACTTCCTGATTGTCCTCTTGTTTATGCTTCAGAAGGCTTCTACGCCATGACAGGCTATGGCCCTGATGAAGTTTCTTGGCCATAACGCTCGCTTTCTTCAGGGCGAAGGCACAGATCCTAAAGAAGTTCAGAAGATACGAGA TGCTATTAAGAAGGGTGAGGCTTGTTCAGTTCGCCTTCTTAACTATCGCAAAGATGGCACACCTTTCTGGAATCTTCTTACAGTTACACCTATTAAGACTCCTGACGGCCGCGTTTCAAAGTTCGTCGGCGTTCAGGTTGATGTTACATCAAAGACTGAGGGCAAAGCTCTTGCTTGA(SEQ ID NO:1).

[0141] 2. Constructing the CreiLOV expression plasmid

[0142] (1) Double digestion of pET28 expression vector

[0143] The enzyme digestion reaction system consisted of: 2 μg pET28 plasmid, 5 μL 10×Fast Digest buffer, 2 μL XhoI restriction enzyme, 2 μL BamHI restriction enzyme, and ultrapure water to a final volume of 50 μL. The enzyme digestion reaction conditions were: 37℃ water bath for 3 hours.

[0144] (2) PCR of CreiLOV gene insertion fragment

[0145] The PCR reaction system consisted of: 20 μL ultrapure water, 2 μL of the first primer (10 μM), 2 μL of the second primer (10 μM), 25 μL of 2×PrimeSTAR master mix, and 1 μL (1 ng) of pUC57-CreiLOV plasmid, totaling 50 μL. The first primer was 28-CreiLOV-F: 5'-TAAGAAGGAGATATACCATGGCTGGCCTTCGCCATAC-3' (SEQ ID NO:3); the second primer was 28-CreiLOV-R: 5'-CAGTGGTGGTGGTGGTGGTGTCAAGCAAGAGCTTTGCCCTCAG-3' (SEQ ID NO:4).

[0146] The PCR reaction procedure is as follows: open the hot lid and preheat to 105℃, (1) denature at 98℃ for 5 min, (2) denature at 98℃ for 15 s, (3) anneal at 56℃ for 15 s, (4) extend at 72℃ for 30 s, and perform 35 cycles of steps (2)-(4), (5) extend at 72℃ for 5 min, and (6) hold at 16℃.

[0147] (3) pET28 double digestion and CreiLOV PCR product gel purification

[0148] The double enzyme digestion and PCR products were subjected to 1% agarose gel electrophoresis at 150V for 30 min to excise target fragments of approximately 5200 bp and 400 bp, respectively. The target fragments were then purified using the Axygen AxyPrep DNA Gel Extraction Kit, following the instructions for use. Elution was performed with 25 μL of ultrapure water.

[0149] (4) Gibson assembly

[0150] The Gibson assembly reaction mixture consisted of 7.5 μL of Gibson reaction mixture, 93 ng of pET28 enzyme digestion and purification product, 7 ng of CreiLOV PCR purification product, and ultrapure water to a final volume of 10 μL. The Gibson assembly reaction conditions were: incubation at 50°C for 60 min.

[0151] (5) E. coli transformation culture and plasmid extraction

[0152] Transform 5 μL of Gibson assembly reaction solution into 50 μL of DH5α chemically competent cells, following the DH5α competent cell instruction manual for specific transformation methods. Pick single colonies and culture them overnight in 3 mL of LB medium containing kanamycin. Extract plasmids using the AxygenAxyPrep plasmid mini-extraction kit, following the kit's instruction manual for specific extraction methods. Elute with 50 μL of ultrapure water.

[0153] 3. Constructing a CreiLOV single-point saturation mutant library

[0154] (1) Design and synthesis of oligonucleotide primers

[0155] A front primer (Cre-F) containing the degenerate codon NNK was designed for each of the 118 amino acid sites of the CreiLOV protein, excluding the start and stop codons. A back primer (Cre-R) was shared by five adjacent amino acid sites, with an 18-nt overlap region between the back primer and the front primer. Additionally, a pair of front and back primers (ori-F and ori-R) with an 18-nt overlap region were designed for the ori region of the replicon. The primer sequences are shown in Table 1.

[0156] Table 1

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164] (2) Two-segment PCR

[0165] The PCR reaction system for fragment 1 (P1) was as follows: 20 μL ultrapure water, 2 μL front primer (10 μM), 2 μL back primer (10 μM), 25 μL 2×PrimeSTAR master mix, 1 μL (1 ng) pET28-CreiLOV plasmid, totaling 50 μL. The front primer was Cre-F and the back primer was ori-R, and their sequences are shown in Table 1. The PCR reaction program for fragment 1 (P1) was as follows: open the heat cap and preheat to 105℃, (1) denature at 98℃ for 5 min, (2) denature at 98℃ for 15 s, (3) anneal at 54℃ for 15 s, (4) extend at 72℃ for 75 s, with a total of 35 cycles for steps (2)-(4), (5) extend at 72℃ for 10 min, and (6) hold at 16℃.

[0166] The fragment 2 (P2) PCR reaction system was as follows: 20 μL ultrapure water, 2 μL front primer (10 μM), 2 μL back primer (10 μM), 25 μL 2×PrimeSTAR master mix, 1 μL (1 ng) pET28-CreiLOV plasmid, totaling 50 μL. The front primer was ori-F and the back primer was Cre-R, and their sequences are shown in Table 1. The fragment 2 (P2) PCR reaction program was as follows: open the heat cap and preheat to 105℃, (1) denature at 98℃ for 5 min, (2) denature at 98℃ for 15 s, (3) anneal at 54℃ for 15 s, (4) extend at 72℃ for 120 s, with a total of 35 cycles for steps (2)-(4), (5) extend at 72℃ for 10 min, and (6) hold at 16℃.

[0167] (3) Purification of PCR products

[0168] The two PCR products were subjected to 1% agarose gel electrophoresis at 150V for 30 min. Target fragments of approximately 2000-2300 bp and 3300-3600 bp were excised, respectively. The target fragments were then purified using the Axygen AxyPrep DNA Gel Extraction Kit, following the kit's instructions. Elution was performed with 25 μL of ultrapure water.

[0169] (4) Gibson assembly

[0170] The Gibson assembly reaction mixture consisted of 7.5 μL of Gibson reaction mixture, 35 ng of PCR purified product of fragment 1, 65 ng of PCR purified product of fragment 2, and ultrapure water to a final volume of 10 μL. The Gibson assembly reaction conditions were: incubation at 50 °C for 60 min.

[0171] (5) E. coli transformation culture and plasmid extraction

[0172] Transform 5 μL of Gibson assembly reaction solution into 50 μL of DH5α chemically competent cells, following the DH5α competent cell instruction manual for specific transformation methods. Scrape the overnight plate-cultured clones into 3 mL of LB medium and extract plasmids using the Axygen AxyPrep plasmid mini-extraction kit, following the kit's instruction manual for specific extraction methods. Elute with 50 μL of ultrapure water.

[0173] 4. Quantitative analysis of single-point saturation mutant libraries and mixed-source libraries

[0174] Prepare fluorescent dye buffer and serially diluted DNA standards according to the Invitrogen Qubit dsDNA BR assay kit instructions, and construct a standard curve. Add 1 μL of a single-point saturated mutant plasmid to 199 μL of fluorescent dye buffer and mix thoroughly. Detect the fluorescence intensity of each plasmid sample using a Tecan Infinite 200Pro microplate reader under 488 nm excitation and 525 nm emission light. Calculate the concentration of each plasmid sample based on the standard curve. Dilute all plasmid samples to the same concentration and mix in equal volumes to obtain the CreiLOV single-point saturated mutant scanning library.

[0175] 5. High-throughput conversion of single-point saturation mutation scanning libraries and CreiLOV protein-induced expression

[0176] 50 μL of *E. coli* BL21(DE3) competent cells were thawed on ice. 2 μL of the CreiLOV single-point saturation mutant scanning library plasmid was added to the competent cells and gently mixed. The cells were transferred to pre-chilled 0.1 cm spaced electroporation cuvettes and electroporated at 1.25 kV using a BioRad Gene Pulser electroporator. Immediately afterward, 1 mL of pre-warmed 37°C LB medium was added to the cuvettes, and the cells were collected into 1.5 mL centrifuge tubes. After incubation at 37°C with shaking for 1 h, the cells were plated onto LB agar plates containing kanamycin and cultured overnight. All transformants were collected using a scraper to 3 mL of LB medium, and OD was measured. 600 With the initial OD 600 Inoculate 0.05 g into fresh LB medium and incubate until OD500. 600 Between 0.4 and 0.6 (approximately 2 hours), IPTG was added to a final concentration of 0.5 mM, and the mixture was cultured at 37°C with shaking for another 6 hours to induce the expression of CreiLOV fluorescent protein.

[0177] 6. Flow cytometry sorting of E. coli cells and pretreatment of sorted cells

[0178] The E. coli cell culture medium that induces CreiLOV expression was diluted to 10. 7 Cells were analyzed by flow cytometry using the FITC channel of a BDFACSAria III cell sorter, following the instructions in the instrument's manual. Cells were divided into eight gates (P3, P4, P5, P6, P7, P8, P9, and P10) based on the fluorescence intensity distribution of the CreiLOV single-point saturation mutation scanning library, ensuring approximately equal cell proportions in each gate. Cells were then simultaneously sorted using an 85 μm diameter nozzle in four channels, in "Purity" mode, and the samples were maintained at 4°C. Three biological replicates were performed, with total cell counts sorted in each replicate being 1.68 million, 1.73 million, and 1.81 million, respectively, representing 454-fold, 467-fold, and 488-fold increases in library diversity. A total of 24 sorted cell samples were obtained. Each sorted cell sample and an appropriate amount of unsorted cell sample were transferred sequentially to 1.5 mL and 0.2 mL centrifuge tubes, centrifuged at 7500 g for 3 min, removed and retained about 10 μL of supernatant, vortexed to mix, and then repeatedly frozen and thawed in liquid nitrogen 3 times, and incubated at 95 °C for 3 min.

[0179] 7. Two rounds of PCR amplification and purification

[0180] (1) Design and synthesis of oligonucleotide primers: The first-round PCR primers were fusions of the high-throughput sequencing platform's sequencing adapter sequence, spacer sequence (0-3 nt), and pET28-CreiLOV plasmid-specific sequence (FP-Cre and RP-Cre). The second-round PCR primers were fusions of the high-throughput sequencing platform's amplification adapter sequence, index (8 nt), and sequencing adapter sequence (P5 and P7). The primer sequences are shown in Table 2.

[0181] Table 2

[0182]

[0183] (2) First round of PCR and purification

[0184] The first round of PCR reaction system consisted of: 8.5 μL ultrapure water, 1 μL of the first primer (10 μM), 1 μL of the second primer (10 μM), 12.5 μL of 2×PrimeSTAR master mix, and 2 μL of pretreated cell sample, for a total of 25 μL. The first primer was FP-Cre, and the second primer was RP-Cre. FP-Cre and RP-Cre with the same serial number were used in pairs. The first round of PCR reaction procedure was as follows: open the heat-sealing lid and preheat to 105℃, (1) denature at 98℃ for 5 min, (2) denature at 98℃ for 15 s, (3) anneal at 54℃ for 15 s, (4) extend at 72℃ for 30 s, with steps (2)-(4) performed for a total of 20 cycles, (5) extend at 72℃ for 5 min, and (6) hold at 16℃. The first-round PCR products were purified using an Axygen AxyPrep gel extraction kit after 1.5% agarose gel electrophoresis. Refer to the kit's instruction manual for specific methods. Elute with 20 μL of ultrapure water.

[0185] (3) Second round of PCR and purification

[0186] The second round PCR reaction system consisted of: 5.5 μL ultrapure water, 1 μL of the first primer (10 μM), 1 μL of the second primer (10 μM), 12.5 μL of 2×PrimeSTAR master mix, and 5 μL of the first round PCR purified product, for a total of 25 μL. The first primer was P5, and the second primer was P7. The 24 sorted samples used a pairwise combination of P5-1 to P5-4 and P7-1 to P7-6, while the original samples used P5-5 and P7-7. The second round PCR reaction procedure was as follows: open the heat-sealing lid and preheat to 105℃, (1) denature at 98℃ for 5 min, (2) denature at 98℃ for 15 s, (3) anneal at 50℃ for 15 s, (4) extend at 72℃ for 40 s, with steps (2)-(4) performed for a total of 10 cycles, (5) extend at 72℃ for 5 min, and (6) hold at 16℃. The second round of PCR products were purified using an Axygen AxyPrep gel extraction kit after 1.5% agarose gel electrophoresis. Refer to the kit's instruction manual for specific methods. Elute with 25 μL of ultrapure water.

[0187] 8. The second round of PCR purification products were quantified and mixed to obtain the amplicon sequencing library.

[0188] For specific methods of quantification and mixing, please refer to "4. Quantification and Mixing of Single-Point Saturated Mutant Libraries" above. After mixing, fragment size was detected and quantified using a Bioptic Qsep400 nucleic acid fragment analyzer. For specific methods, please refer to the instrument's instruction manual.

[0189] 9. The amplicon sequencing library was sequenced using an Illumina NovaSeq 6000 sequencer.

[0190] For specific instructions, please refer to the user manual of the Illumina NovaSeq 6000 sequencer.

[0191] 10. Mutant identification and statistics

[0192] (1) Removal of spacer and primer sequences: The specific method is to use the cutadapt software to remove spacer and primer sequences with the parameter -u 27. For details, please refer to the software's instruction manual.

[0193] (2) Mutant identification and statistics: The method is to use Enrich2 software for identification and statistics, with parameters as MinimumQuality 20, Maximum N's 0, Overlap 95, Maximum Mismatches 0, Maximum Mutations 3. For specific methods, please refer to the software's user manual.

[0194] 11. Estimation of fluorescence intensity in mutants

[0195] (1) Calculate the cell number distribution of each mutant in each phylum according to the following formula: C ik =(N ik / N k )×F k ×C. Where, N ik N represents the number of sequencing fragments of mutant i in phylum k. k F represents the total number of sequencing fragments in gate k. k Let C be the proportion of cells belonging to gate k, and C be the total number of cells sorted. Then C ik denoted as the number of cells in phylum k for mutant i.

[0196] (2) Calculate the fluorescence intensity and standard deviation of each mutant by weighted average based on the average fluorescence intensity of each phylum and the cell number distribution of each mutant in each phylum.

[0197] (3) Assuming that the cell number distribution of each mutant in each phylum follows a Gaussian distribution, and using the weighted average calculation result as the initial expectation and variance, the fluorescence intensity and standard deviation of each mutant were estimated using the open-source Python toolkit SciPy, based on the Gaussian distribution model and the upper and lower boundaries of the fluorescence intensity of each phylum. The mutants and their fluorescence intensities are shown below. Figure 4 As shown in the figure, the horizontal axis represents the amino acid sites of CreiLOV, and the vertical axis represents the amino acid mutations at those sites. The redder the color, the stronger the fluorescence of the mutant; the bluer the color, the weaker the fluorescence of the mutant. The dots represent the amino acid types of the wild-type CreiLOV protein.

[0198] 12. Synthesizing multi-point directed combination mutant gene libraries

[0199] The mutants were sorted from highest to lowest fluorescence intensity, and the top 15 amino acid sites were identified. Then, the top 20 mutation directions at these 15 amino acid sites were determined. The mutation direction for each amino acid site was defined as an amino acid from the wild-type CreiLOV protein and the mutation direction contained in these 20 mutants. The diversity of this multi-site directed combinatorial mutant library is approximately 1.8 × 10⁻⁶. 5 The gene library was synthesized using the TWIST Precise Diversity Scattered Mutant Gene Library (SOLD) product; please refer to the product's instruction manual for specific methods. Compared to the wild-type protein sequence, the structure of this gene library is as follows: Figure 5 As shown, numbers 1-15 indicate that the library contains 15 amino acid sites, and their positions, amino acid types, and codons in the wild-type CreiLOV protein are shown as "Amino Acid Sites," "Wild-Type Amino Acids," and "Wild-Type Codons," respectively. AY indicates the amino acid orientation of each amino acid site in the library, and the number of amino acid orientations, including wild-type amino acids, is shown in the "Amino Acid Orientation" row.

[0200] 13. Constructing a multi-point directed combination mutant expression library

[0201] (1) Double digestion and purification of pET28 expression vector. For specific methods, please refer to “2. Construction of CreiLOV expression plasmid” above.

[0202] (2) Gibson assembly: 15 μL Gibson reaction mixture, 186 ng pET28 digestion and purification product, 14 ng CreiLOV multi-point directed combination mutant gene library, and ultrapure water to 20 μL. Three reactions were performed. The mixture was incubated at 50℃ for 60 min in a PCR instrument.

[0203] (3) The assembled product was purified and concentrated using the Zymo Research DNA Purification and Concentration Kit (D4013). For specific methods, please refer to the instruction manual of the kit. Elution was performed using 6 μL of ultrapure water.

[0204] (4) E. coli transformation culture and plasmid extraction:

[0205] 2 μL of the purified Gibson assembly product was electroporated into 50 μL of Invitrogen MegaX DH10B T1 competent cells. Refer to the instructions for use of these competent cells for specific transformation methods. The transformed bacterial culture was then plated onto 6 LB agar plates containing kanamycin and incubated overnight. The total number of transformants was approximately 4 × 10⁻⁶. 6 This method covers more than 20 times the diversity of the library. All transformants were collected into 20 mL of LB medium containing kanamycin. An appropriate amount of bacterial culture was used to extract plasmids using the Axygen AxyPrep plasmid mini-extraction kit. The specific extraction method was described in the kit's instructions. Elution was performed with 50 μL of ultrapure water.

[0206] 14. A multi-site targeted combined mutant expression library was transformed into BL21(DE3) electrocompetent cells in high throughput, with a total of approximately 10 transformants. 7 This method covers more than 50 times the diversity of the library and induces CreiLOV protein expression. For specific methods, please refer to "5. High-throughput conversion of single-point saturation mutation scanning libraries and CreiLOV protein induction expression" above.

[0207] 15. Perform three consecutive rounds of enrichment and sorting using a BD FACSAria III cell sorter. In the first round, sort the top 10% of cells by fluorescence intensity. Then, perform flow cytometry analysis on the cells from the first round and sort the top 20% by fluorescence intensity. In the second round, perform flow cytometry analysis on the cells from the second round and sort the top 30% by fluorescence intensity. Refer to the instrument's instruction manual for specific methods. Spread the third round of sorted cells (approximately 7500 cells) onto LB agar plates containing kanamycin and incubate overnight.

[0208] 16. Pick 88 single colonies and culture them overnight in 1 mL of LB medium containing kanamycin. Inoculate 20 μL into 1 mL of fresh LB medium containing kanamycin, and culture at 37°C with shaking for 2 h. Then add IPTG to a final concentration of 0.5 mM and continue culturing for 6 h to obtain the bacterial culture. Use pET28 expression vector and pET28-CreiLOV plasmid to transform Escherichia coli as controls.

[0209] 17. Use a Tecan Infinite 200Pro microplate reader to detect the absorbance of the bacterial culture at 600 nm and the fluorescence intensity at 450 nm excitation and 495 nm emission. Refer to the instrument's instruction manual for specific methods. Calibrate the OD using blank culture medium. 600 Utilizing OD 600 The fluorescence intensity of each sample was calibrated by transforming E. coli with the pET28 expression vector to calibrate the fluorescence intensity of each mutant.

[0210] Of the 88 mutants, 57 had fluorescence intensities higher than the wild-type CreiLOV protein. Among them, 50 CreiLOV mutants had fluorescence intensities exceeding the wild-type by more than 10%, and 20 CreiLOV mutants had fluorescence intensities exceeding the wild-type by more than 40%.

[0211] 18. Extract the first 20 mutant plasmids using the Axygen AxyPre plasmid mini-extraction kit. Refer to the kit's instruction manual for specific methods. Perform Sanger sequencing and sequence alignment on the plasmids to obtain the nucleotide sequences and amino acid mutation combinations of these 20 CreiLOV mutants.

[0212] The results are shown in Table 3.

[0213] Table 3

[0214]

[0215] As shown in Table 3, the CreiLOV fluorescence-enhanced mutants screened using the method described in this invention exhibit high fluorescence intensity, large quantity, and high proportion. The number of mutants with a 40% increase in fluorescence intensity compared to wild-type CreiLOV reaches approximately 57%, significantly improving the screening efficiency of superior mutants. Furthermore, the method contains a large number of amino acid mutation sites with diverse mutation directions. Compared to the semi-rational design of saturation mutations targeting specific amino acids, it greatly increases the number of target amino acid sites, and compared to error-prone PCR and other random mutagenesis techniques, it significantly improves the direction of amino acid mutations.

[0216] Other embodiments

[0217] Figure 6This paper presents another flowchart of the protein engineering method of the present invention for obtaining the correspondence between mutant sequences and target characteristics. In this method, single-site saturated mutant libraries at each amino acid site are transformed into expression hosts. 192 single clones are picked from each transformation plate and cultured in two 96-well deep-well plates. After inoculation into two new 96-well deep-well plates, protein expression is induced. High-throughput fluorescence detection of each single clone is performed using a microplate reader and 96-well plates. Mixed-cell sequencing technology is used to determine the sequence of each mutant. This method eliminates the need for biostatistical analysis to estimate the fluorescence intensity of each mutant, enabling the scanning and analysis of single-amino acid site mutations and their corresponding fluorescence intensities in the target protein.

[0218] Furthermore, the selection of the target protein and its properties in the above embodiments can be extended to other proteins that can couple the target characteristics of the protein with the fluorescence signal. For example, the expression level of the fluorescent protein can be higher with stronger transcription factor activity to carry out directed evolution of transcription factor activity; or, the expression level of the protein can be higher with stronger protein-protein interaction by coupling with the fluorescent protein to carry out directed evolution of protein-protein interaction.

Claims

1. A protein mutant having amino acid mutations relative to wild-type CreiLOV protein as shown in any of S1-S20, wherein the amino acid sequence of the wild-type CreiLOV protein is shown in SEQ ID NO:2, and the amino acid mutations are any of the following S1-S20: S1: G3E, L4N, R5D, T7S, A29H, R60N, I92C, R98D, V107M, T113E; S2: G3E, L4N, R5D, T7H, A29H, G34T, R60N, D61Q, I92C; S3: T7S, G34T, R60N, D61Q, I92C, D96Q, R98D, V107M; S4: L4N, G34T, I92C, D96Q, V107M; S5: L4N, R5D, D61Q, I92C, D96Q, V107M; S6: T7S, A29K, I92C, D96Q, R98D, V107M, T113E; S7: L4N, T7S, R60N, I92C, D96Q; S8: R5D, T7H, G34T, Q47R, I92C, R98D, V107M, T113E; S9: L4N, T7H, I92C, D96Q, R98D, V107M; S10: R5D, T7H, A29H, G34T, R60N, D61Q, I92C; S11: R5D, T7H, A29K, G34T, I92C, D96Q; S12: A29H, G34T, D61Q, I92C, D96Q, V107M; S13: G3E, R5D, T7H, R60N, D61Q, I92C, R98D; S14: R5D, T7H, A29H, G34T, R60N, D61Q, I92C, R98D, V109Q; S15: T7S, G34T, R60N, D61Q, I92C, D96Q, R98D, V107M, V109Q; S16: L4N, T7S, A29H, R60N, I92C, D96Q, V109Q; S17: R5D, A29H, I92C, R98D, V107M, T113E; S18: R5D, D61Q, I92C, D96Q, V107M, T113E; S19: L4N, T7S, A29H, R60N, D61Q, I92C, D96Q; S20: R5D, T7S, A29H, R60N, I92C, D96Q, R98D, V107M, T113E.

2. A method for targeted modification of a protein to obtain a target protein mutant, wherein the target protein mutant is the mutant described in claim 1, the method comprising: Step 1: Construct a single-amino acid site saturated mutant library of the target protein, and scan and analyze the single-amino acid site mutations of the target protein and the corresponding target characteristics to obtain the correspondence between the mutant sequence and the target characteristics; Step 2: Based on the correspondence between mutant sequences and target characteristics, select the amino acid sites and amino acid mutations of the dominant mutants of the target characteristics, and synthesize a multi-point directed combination mutant gene library; Step 3: Using the obtained multi-point directed combination mutant gene library as a template, construct a multi-point directed combination mutant expression library; Step 4: Transform the multi-point targeted combinatorial mutant expression library into the expression host to induce the expression of the multi-point combinatorial mutant target protein; Step 5: Enrich and sort cells expressing multi-point combination mutant target proteins to screen for target protein mutants; The target protein is a protein that can couple its target characteristics with the intensity of a fluorescence signal; The target characteristic is fluorescence intensity.

3. The method according to claim 2, wherein, Step one, the process of constructing a single-amino acid site saturation mutant library of the target protein, includes: Using the gene sequence encoding the protein to be modified as a template, an expression plasmid for the target protein was constructed. Using the target protein expression plasmid as a template, a single amino acid site saturation mutant library of the target protein was constructed.

4. The method according to claim 3, wherein, The process of constructing a target protein expression plasmid includes: The expression plasmid vector was subjected to double enzyme digestion and enzyme digestion product purification. PCR was performed using the gene sequence encoding the protein to be modified as a template and primers that fused the expression plasmid vector and the target gene specific sequence, and the PCR product was purified. The enzyme digestion and PCR purification products were assembled using Gibson, and the assembled products were transformed into expression host cells for culture. Plasmids were then extracted to obtain the target protein expression plasmid.

5. The method according to claim 3, wherein, The process of constructing a single-amino acid site saturation mutant library of the target protein includes: Degenerate oligonucleotides are designed to introduce all possible codons at selected amino acid sites. Degenerate oligonucleotide primers are used to amplify the target protein expression plasmid by PCR. After purification of the PCR product, Gibson assembly is performed. The assembly product is transformed into expression host cells for culture, and the plasmid is extracted to obtain the target protein single-point saturation mutant plasmid, i.e., single-point saturation mutant library.

6. The method according to claim 5, wherein, The PCR amplification of the target protein expression plasmid consists of two PCR segments: one segment uses a target gene-specific front primer and a vector replicon-specific back primer, and the other segment uses a target gene-specific back primer and a vector replicon-specific front primer.

7. The method according to claim 2, wherein, In step one, one or more of the following methods are used to obtain the correspondence between the mutant sequence and the target characteristic: Method 1: A single-amino acid site saturation mutation library of the target protein was constructed, and a single-point saturation mutation scanning library of the target protein was constructed accordingly. The single-point saturation mutant scanning library was transformed into the expression host to induce the expression of the target protein, and cells expressing the single-point mutant of the target protein were obtained. Cells that were induced to express single-point mutants of the target protein were sorted to obtain multiple groups of cell samples; Each group of cells from the sorted cell samples was used as a template for multiple rounds of PCR amplification and purification. The PCR purification products of multiple cell samples were mixed to obtain an amplicon sequencing library. The obtained amplicon sequencing library was subjected to high-throughput sequencing to obtain sequencing data, and the target characteristics of each mutant were further estimated to obtain the correspondence between mutant sequences and target characteristics. Method 2: A single-amino acid site saturated mutant library of the target protein was constructed, transformed into an expression host, and single clones were selected to induce expression of the target protein. The target characteristics were detected, and the sequence of each mutant was determined using pooled sequencing technology to obtain the correspondence between the mutant sequence and the target characteristics.

8. The method according to claim 7, wherein, The process of constructing a single-point saturation mutant scanning library of the target protein includes: Multiple single-point saturation mutation libraries are mixed to obtain a single-point saturation mutation scanning library.

9. The method according to claim 8, wherein, The process of mixing multiple single-point saturation mutation libraries to obtain a single-point saturation mutation scanning library includes: Quantitative analysis of single-point saturation mutant libraries; Multiple single-point saturation mutation libraries are mixed in equal amounts to obtain a single-point saturation mutation scanning library; The single-point saturation mutation scanning library contains single-point saturation mutation libraries of all amino acid sites of the target protein except the start codon.

10. The method according to claim 9, wherein, When quantifying single-point saturated mutant libraries, fluorescence quantification is used.

11. The method according to claim 7, wherein, The process of sorting cells that express single-point mutants of the target protein includes: Cells expressing single-point mutants of the target protein were sorted according to their target characteristics to obtain multiple groups of cell samples.

12. The method according to claim 11, wherein, The cell sorting method was fluorescence-activated cell sorting, which used flow cytometry to sort cells that expressed single-point mutants of the target protein based on fluorescence intensity.

13. The method according to claim 12, wherein, 4-16 groups of cell samples were sorted out.

14. The method according to claim 12, further comprising a process of lysing the sorted cells.

15. The method according to claim 14, wherein, The pyrolysis process was carried out using a repeated freeze-thaw cycle.

16. The method according to claim 7, wherein, In the process of performing multiple rounds of PCR amplification and purification using each group of cells from the sorted cell samples as a template, the multiple rounds of PCR are two-round PCR amplification methods, wherein: The template for the first round of PCR amplification was the cell samples obtained after sorting. The fusion sequence of the adapter sequence, spacer sequence and expression vector-specific sequence was used as primers for PCR amplification and purification. The second round of PCR used the purified product from the first round of PCR as a template, and amplified the adapter sequence, index, and the fusion sequence of the sequencing adapter sequence using a high-throughput sequencing platform as primers.

17. The method of claim 16, wherein: The first round of PCR primers adds a spacer sequence of 0-8 nucleotides between the sequencing adapter sequence and the target gene-specific sequence on the high-throughput sequencing platform; The second round of PCR primers uses 6-10 nucleotide sequences as index sequences.

18. The method according to claim 7, wherein, The process of mixing PCR purification products from multiple cell samples to obtain amplicon sequencing libraries includes: Quantification of PCR purified products from each group of cell samples was performed. The PCR purification products from multiple cell samples were mixed in equal volumes to obtain an amplicon sequencing library.

19. The method of claim 18, further comprising quantifying and / or detecting fragment size of the amplicon sequencing library.

20. The method according to claim 7, wherein, The process of estimating the target characteristics of each mutant includes: Mutation determination and data statistics were performed on the high-throughput sequencing data of the obtained amplicon sequencing library to obtain the normalized number distribution of each mutant in each sorting group; The target characteristic intensity of each mutant is calculated by weighted average based on the average value or boundary of the target characteristic of each sorting group and the number distribution of each mutant in each sorting group. Assuming that the number distribution of each target protein mutant in each sorting group is Gaussian or gamma in logarithmic coordinate system, the target characteristic intensity of each mutant is estimated by the maximum likelihood estimation method.

21. The method according to claim 2, wherein, Step two, the process of synthesizing a multi-point directed combination mutant gene library, includes: Based on the correspondence between mutant sequences and target characteristics, the amino acid sites of mutants with the highest target characteristic intensity are selected as the target sites for the combinatorial mutant library. At the same time, the amino acid mutations with the highest target characteristics in the same amino acid site are selected to synthesize a multi-point targeted combinatorial mutant gene library.

22. The method according to claim 2, wherein, The process of enriching and sorting cells expressing multi-point combination mutant target proteins to screen for target protein mutants includes: Cells with target characteristic advantages are enriched and sorted in a single round or in multiple consecutive rounds; The process involves at least three rounds of sorting: the first round sorts the 5%-10% of cells with the highest target characteristic intensity, and then performs flow cytometry analysis on the first-round sorted cells to sort the remaining 10%-20% of cells with the highest target characteristic intensity; the second round sorts the cells, and then performs flow cytometry analysis on the remaining 20%-30% of cells with the highest target characteristic intensity; finally, the cells sorted in the third round are plated on screening culture plates and cultured overnight. After culturing the obtained monoclonal antibodies to the logarithmic growth phase, the target protein was induced to express. The strength of the target characteristics was then detected, or the plasmid was further extracted and sequenced to screen for mutants of the target protein.

23. The method according to any one of claims 2-7, wherein: The expression host includes *Escherichia coli*, *Saccharomyces cerevisiae*, or *Pichia pastoris*; and / or The conversion method is an electroconversion method.

Citation Information

Patent Citations

  • Optically activated receptors

    CN106103487A

  • Method for carrying out directional evolution on gene promoter

    CN107338241A