Design method of micromolecular binding protein
By employing a progressively precise screening funnel and multi-dimensional performance evaluation, the problems of low sequence generation efficiency and high false positives in the design of small molecule binding proteins have been solved, achieving efficient and accurate protein design and improving the design success rate and practical value.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional small molecule binding protein design methods suffer from low sequence diversity and poor quality, high false positive rates, and a lack of comprehensive evaluation of multiple performance indicators, resulting in long design cycles and extremely low success rates.
A progressively precise screening funnel was used, combined with the PLIP tool and an AI generation model. Candidate protein sequences were generated through medium-precision primary screening and high-precision fine screening strategies, and multi-dimensional performance screening was performed. Finally, wet experiments were conducted for verification.
It significantly improved the design success rate of small molecule binding proteins, reaching 14%, saving experimental costs and time, and ensuring that the designed proteins have good expression ability and structural robustness.
Smart Images

Figure CN121789784A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computational biology and protein engineering, and particularly to a method for designing small molecule binding proteins. Background Technology Small molecule binding proteins have wide applications in drug delivery, biosensing, and industrial biocatalysis. However, traditional protein design methods rely heavily on researchers' experience and suffer from the following inherent drawbacks, resulting in lengthy design cycles and extremely low success rates (typically less than one in a thousand): (1) Low sequence generation diversity and poor quality: It relies on homology modeling or random mutation, making it difficult to explore entirely new protein sequence space, and the generated sequences are often structurally unstable or insoluble.
[0002] (2) Virtual screening is difficult to balance accuracy and efficiency: Although physical methods such as molecular dynamics simulation are highly accurate, the calculation time is as long as several days or even weeks, which cannot be used for large-scale screening; while fast scoring functions are not accurate enough and have a very high false positive rate, which makes a large number of invalid sequences enter the time-consuming and laborious experimental verification stage.
[0003] (3) Single evaluation dimension: Traditional methods usually only focus on binding affinity (binding energy), ignoring the solubility, structural stability and other properties of proteins that are crucial to their practical application. This results in many "excellent" design sequences in calculations failing to be successfully expressed or maintain their function in experiments.
[0004] Therefore, there is an urgent need in this field for a collaborative optimization design method that can generate high-quality sequences in high throughput and perform multi-dimensional performance screening efficiently and accurately, so as to fundamentally improve the success rate and efficiency of small molecule binding protein design. Summary of the Invention
[0005] The purpose of this invention is to provide a method for designing small molecule binding proteins to solve the problems of low sequence generation efficiency, high false positive rate in screening, and lack of comprehensive evaluation of multiple performance indicators in the prior art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A method for designing small molecule binding proteins includes the following steps: S1. Structure acquisition steps: Obtain the structure of the complex of the target small molecule ligand and the protein; S2. Candidate sequence generation step: Based on the complex structure described in step S1, a protein generation model is used, and the parameters corresponding to the protein generation model are adjusted to generate a set of candidate protein sequences. S3. Initial screening step: Based on the candidate protein sequence set described in step S2 and the small molecule ligand described in step S1, predict the structure of the complex and screen out potential sequences from the structure of the complex. S4. Fine screening stage: Based on step S3, screen the potential sequence and the small molecule ligand in step S1 to predict the complex structure 2, and screen the complex structure 2 to select the preferred protein sequence 1. S5. Wet experiment verification step: Based on the preferred protein sequence selected in step S4, perform soluble expression test and activity test to obtain small molecule binding protein.
[0007] Preferably, in step S1, obtaining the structure of the target small molecule ligand-protein complex includes at least one of a protein structure database and a structure prediction tool.
[0008] Preferably, in step S2, the protein generation model includes a diffusion-based generation model and a structural constraint-based language model, wherein the parameter of the diffusion-based generation model is the mask ratio, and the parameter of the structural constraint-based language model is the sampling temperature.
[0009] Preferably, step S3 includes the following sub-steps: S31. Using a structure prediction method with moderate computational accuracy, the candidate protein sequence set is combined with the small molecule ligand in step S1 to predict the structure of a complex with one number of sequences. S32, Analysis step S31 predicts the structure of the complex and extracts intermolecular interaction parameters using the PLIP tool; S33. Perform Z-score normalization on the intermolecular interaction parameters extracted in step S32, calculate the weighted score according to the weight ratio, and retain the complex sequence with two entries in descending order of score. S34. Obtain the first sequence of the complex according to step S33, and use the FPocket tool to calculate the pocket index. S35. Perform Z-score standardization on the pocket index calculated in step S34, calculate the weighted score according to the weight ratio of 2, and retain the potential sequence of number 3 from high to low according to the score.
[0010] Preferably, in step S33, the intermolecular interaction parameters include the number of hydrogen bonds, the number of salt bridges, the number of hydrophobic interactions, and the number of π-π interactions, with the weight ratios being 30% for the number of hydrogen bonds, 25% for the number of salt bridges, 25% for the number of hydrophobic interactions, and 20% for the number of π-π interactions.
[0011] Preferably, in step S34, the pocket indicators are pocket volume, hydrophobicity, and polarity.
[0012] Preferably, in step S35, the weight ratio two is a volume weight of 40%, a hydrophobicity weight of 35%, and a polarity weight of 25%.
[0013] Preferably, step S4 includes the following sub-steps: S41. Using a high-precision structure prediction method, the potential sequence screened in step S3 is combined with the small molecule ligand in step S1 to predict the structure of the complex. S42. Based on step S41, predict the second complex structure, filter the second complex structure according to the local conformation confidence threshold one and the global structure quality threshold two, and retain the second complex sequence with four entries in descending order of local conformation confidence. S43. Based on the complex sequence two described in step S42, calculate the binding energy, number of interface atoms, and positive and negative charge pairing ratio of the interface using APBS. Based on the number of interface atoms and the charge pairing ratio, retain five complex sequences three in order of increasing electrostatic binding energy; the number of interface atoms is greater than 50, and the charge pairing ratio is greater than 0.6. S44. Based on the complex sequence three retained in step S43, global solubility fraction is predicted using the SolMPNN model. The sequences with six soluble fractions are selected as the preferred protein sequence one and sorted from high to low.
[0014] Preferably, the first local conformation confidence threshold is greater than 0.8, and the second single-chain global structure quality threshold is greater than 0.8.
[0015] Preferably, step S5 includes the following sub-steps: S51. The preferred protein sequence described in step S4 is cloned into an expression vector, and the expression vector is transferred to a host cell to obtain recombinant engineered bacteria; S52. The recombinant engineered bacteria obtained in step S51 are induced and cultured to achieve the expression of the target protein; S53. Collect the recombinant engineered bacteria from step S52 after induced expression, and separate them by centrifugation after ultrasonic disruption, collecting the supernatant and precipitate respectively. S54. The soluble target protein in the supernatant collected in step S53 is purified by nickel column affinity chromatography to obtain the purified target protein. S55. Perform SDS-PAGE analysis on the target protein described in step S54 to obtain a highly soluble expression of the preferred protein sequence with seven sequences, which can achieve highly soluble expression. S56. The preferred protein sequence obtained in step S55 is incubated with the small molecule ligand obtained in step S1. The binding product is detected by phase chromatography-mass spectrometry to verify that the obtained small molecule binding protein has the activity of the small molecule ligand obtained in step S1, and the small molecule binding protein is obtained.
[0016] The beneficial effects of this invention are: 1. This invention utilizes a deep integration of AI-generated data (PLIP tools) and physical screening to construct a progressively precise screening funnel, significantly reducing the false positive rate. In the specific operation, only 7 final screening sequences were verified using wet experiments, successfully obtaining proteins with binding activity, achieving a success rate of approximately 14% (1 / 7). This represents a leap of one to two orders of magnitude compared to the success rates of traditional methods or single-model designs, which are typically below 1%, or even only one in a thousand to one in ten thousand, significantly saving experimental costs and time.
[0017] 2. By combining "medium-precision initial screening" with "high-precision fine screening", the calculation accuracy of key steps is guaranteed while avoiding high-cost calculations on all massive candidate sequences. This makes it possible to perform efficient and reliable fine screening of ultra-large-scale sequence libraries (such as tens of thousands to hundreds of thousands) with limited computing power.
[0018] 3. This invention breaks through the traditional approach of focusing only on binding affinity. It incorporates multiple indicators that are crucial for the practical application of proteins, such as binding activity, structural stability, and solubility, into a unified framework for synergistic screening and optimization. This ensures that the final designed protein sequence can not only bind to the target molecule but also has good expression ability and structural robustness, greatly improving the practical value of the design results.
[0019] 4. By integrating various AI models and physical calculation tools such as PLIP through modular design, an end-to-end automated pipeline from sequence generation to priority sequence output has been formed, reducing reliance on expert experience, improving the repeatability and efficiency of the design process, and laying the foundation for high-throughput design of small molecule binding proteins. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating a small molecule binding protein design method according to the present invention; Figure 2 This is a high-soluble expression diagram of the present invention; Figure 3 This is a diagram of the heme binding activity of the present invention; Figure 4 This is a screenshot of the sequence library of the present invention; Figure 5 This is a screenshot of the calculation results from the PLIP tool of this invention; Figure 6This is a screenshot of the calculation results from the FPocket tool of this invention; Figure 7 This is a structural quality ranking diagram of the present invention; Figure 8 This is a screenshot of the core content of your_input_file.in in this invention; Figure 9 These are the calculation results from the APBS tool of this invention; Figure 10 This is the result of the SolMPNN project address calculation in this invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Please see Figures 1-10 This invention provides a method for designing small molecule binding proteins, comprising the following steps: S1. Structure acquisition steps: Obtain the structure of the complex of the target small molecule ligand and the protein; S2. Candidate sequence generation step: Based on the complex structure described in step S1, a protein generation model is used, and the parameters corresponding to the protein generation model are adjusted to generate a set of candidate protein sequences. S3. Initial screening step: Based on the candidate protein sequence set described in step S2 and the small molecule ligand described in step S1, predict the structure of the complex and screen out potential sequences from the structure of the complex. S4. Fine screening stage: Based on step S3, screen the potential sequence and the small molecule ligand in step S1 to predict the complex structure 2, and screen the complex structure 2 to select the preferred protein sequence 1. S5. Wet experiment verification step: Based on the preferred protein sequence selected in step S4, perform soluble expression test and activity test to obtain small molecule binding protein.
[0023] Specifically, in step S1, obtaining the structure of the target small molecule ligand-protein complex includes at least one of a protein structure database and a structure prediction tool.
[0024] The structure of a known complex can be downloaded from a protein structure database (such as RCSB PDB), or the structure prediction tool (such as AlphaFold Server) can be used to predict the complex structure of the target small molecule and protein.
[0025] Specifically, in step S2, the protein generation model includes a diffusion-based generation model (such as Chroma) and a structure-constrained language model (such as ProteinMPNN). The parameter of the diffusion-based generation model is the mask ratio, and the parameter of the structure-constrained language model is the sampling temperature.
[0026] The protein generation model includes diffusion-based generative models (such as Chroma) or structurally constrained language models (such as ProteinMPNN). Sequences with different levels of diversity can be generated by adjusting model parameters (such as mask ratio and sampling temperature), and then merged to form a candidate sequence library.
[0027] Specifically, step S3 includes the following sub-steps: S31. Using a structure prediction method with moderate computational accuracy (such as Chai-Lab), the candidate protein sequence set is combined with the small molecule ligand in step S1 to predict the structure of a complex with one number of sequences. S32, Analysis step S31 predicts the structure of the complex and extracts intermolecular interaction parameters using the PLIP tool; S33. Perform Z-score normalization on the intermolecular interaction parameters extracted in step S32, calculate the weighted score according to the weight ratio, and retain the complex sequence with two entries in descending order of score. S34. Obtain the first sequence of the complex according to step S33, and use the FPocket tool to calculate the pocket index. S35. Perform Z-score standardization on the pocket index calculated in step S34, calculate the weighted score according to the weight ratio of 2, and retain the potential sequence of number 3 from high to low according to the score.
[0028] Specifically, in step S33, the intermolecular interaction parameters include the number of hydrogen bonds, the number of salt bridges, the number of hydrophobic interactions, and the number of π-π interactions, with the weight ratios being 30% for the number of hydrogen bonds, 25% for the number of salt bridges, 25% for the number of hydrophobic interactions, and 20% for the number of π-π interactions.
[0029] Specifically, in step S34, the pocket indicators are pocket volume, hydrophobicity, and polarity.
[0030] Specifically, in step S35, the weight ratio two is 40% for volume, 35% for hydrophobicity, and 25% for polarity.
[0031] Specifically, step S4 includes the following sub-steps: S41. Using a high-precision structure prediction method (such as AlphaFold3), the potential sequence screened in step S3 is combined with the small molecule ligand in step S1 to predict the complex structure II. S42. Based on step S41, predict the second complex structure, filter the second complex structure according to the local conformation confidence threshold one and the global structure quality threshold two, and retain the second complex sequence with four entries in descending order of local conformation confidence. S43. Based on the complex sequence two described in step S42, calculate the binding energy, number of interface atoms, and positive and negative charge pairing ratio of the interface using APBS. Based on the number of interface atoms and the charge pairing ratio, retain five complex sequences three in order of increasing electrostatic binding energy; the number of interface atoms is greater than 50, and the charge pairing ratio is greater than 0.6. S44. Based on the complex sequence three retained in step S43, global solubility fraction is predicted using the SolMPNN model. The sequences with six soluble fractions are selected as the preferred protein sequence one and sorted from high to low.
[0032] Specifically, the first local conformation confidence threshold is greater than 0.8, and the second single-chain global structure quality threshold is greater than 0.8.
[0033] Specifically, step S5 includes the following sub-steps: S51. The preferred protein sequence described in step S4 is cloned into an expression vector, and the expression vector is transferred to a host cell to obtain recombinant engineered bacteria; S52. The recombinant engineered bacteria obtained in step S51 are induced and cultured to achieve the expression of the target protein; S53. Collect the recombinant engineered bacteria from step S52 after induced expression, and separate them by centrifugation after ultrasonic disruption, collecting the supernatant and precipitate respectively. S54. The soluble target protein in the supernatant collected in step S53 is purified by nickel column affinity chromatography to obtain the purified target protein. S55. Perform SDS-PAGE analysis on the target protein described in step S54 to obtain a highly soluble expression of the preferred protein sequence with seven sequences, which can achieve highly soluble expression. S56. The preferred protein sequence obtained in step S55 is incubated with the small molecule ligand obtained in step S1. The binding product is detected by phase chromatography-mass spectrometry to verify that the obtained small molecule binding protein has the activity of the small molecule ligand obtained in step S1, and the small molecule binding protein is obtained.
[0034] The specific operation process based on the above embodiments is as follows: Design and Validation of an Active Heme-Binding Protein This embodiment uses the design of a protein that binds to heme (HEM) small molecules as an example to demonstrate the specific implementation process and effects of the present invention.
[0035] S1. Structure Acquisition: Obtain the known heme-protein complex structure (PDB ID: 2VEE) from the Protein Database (RCSB PDB). Preprocess the structure using molecular editing software (such as PyMOL), retaining only protein chain A and the heme ligand (HEM) as structural constraint templates for subsequent sequence generation, or use the af3 prediction tool (https: / / alphafoldserver.com / ) to predict the complex structure by inputting the sequence.
[0036] S2, Candidate Sequence Generation Based on the structural template obtained in S21, we jointly use two different types of protein generation models to fully leverage their complementary advantages and generate a diverse candidate sequence library: Chroma based on diffusion model: Set different mask ratios (mask_ratio = 0.2, 0.3, 0.4, 0.5, 0.6), each ratio generates 1000 sequences, for a total of 5000 sequences.
[0037] Chroma project address: https: / / github.com / generatebio / chroma Execute the command: python chroma_substructure_constraint.py --input_pathdataset / 1PET --output_path dataset / test --mask_ratio 0.2 ProteinMPNN based on a language model: Set different sampling temperatures (sampling_temp = 0.2, 0.3, 0.4, 0.5, 0.6), generate 1000 sequences for each temperature, for a total of 5000 sequences.
[0038] The ProteinMPNN project can be found at: https: / / github.com / dauparas / ProteinMPNN Execute command: python protein_mpnn_run.py --pdb_path path_to_PDB --pdb_path_chains chains_to_design --out_folder output_dir --num_seq_per_target 2 --sampling_temp "0.1" --seed 37 --batch_size 1 The sequences generated by the two models were merged and deduplicated, resulting in approximately 10,000 initial candidate sequences, forming a candidate sequence library. Figure 4 .
[0039] S3. Initial Screening Stage – Structural Chemical Feature-Based Weighted Screening Logic The goal of this phase is to quickly and efficiently screen out hundreds of potential sequences from tens of thousands of sequences.
[0040] S31: Rapid Structure Prediction: Using the computationally efficient Chai-Lab tool, the structure of complexes of 10,000 candidate sequences with heme is rapidly predicted.
[0041] Chai-Lab tools: https: / / github.com / chaidiscovery / chai-lab Execute the command: python examples / predict_structure.py S32: Molecular interaction screening: Each predicted complex was analyzed using the PLIP tool to extract the number of hydrogen bonds (h), salt bridges (sb), hydrophobic interactions (hy), and π-π interactions (p). These metrics were Z-score normalized and weighted (Wsb=0.25, Why=0.25, Wp=0.2) to calculate a weighted score S1. The top 1000 sequences with the highest S1 scores were retained.
[0042] PLIP tool: https: / / github.com / pharmai / plip Execute the command: `plip -f your_complex.pdb -o output_directory` The calculation results are as follows Figure 5 S33: Binding Pocket Properties Screening: The binding pockets of the 1000 complexes were analyzed using the FPocket tool, and pocket volume (v), hydrophobicity (hy), and polarity (p) were calculated. After standardization, a weighted score S2 was calculated based on weights (Wv=0.4, Why=0.35, Wp=0.25). The 100 sequences with the highest S2 scores were retained.
[0043] FPocket tool: https: / / github.com / Discngine / fpocket Execute the command: fpocket -f your_protein.pdb The calculation results are as follows Figure 6 S4. Fine Screening Stage – High-Precision Interface Physical Screening Logic The goal of this phase is to conduct a high-cost, high-precision, in-depth evaluation of the hundred-digit sequences in order to select the optimal single-digit sequences.
[0044] S41: High-precision structure prediction: AlphaFold3 was used to perform high-precision prediction of the complex structure of the above 100 sequences with heme.
[0045] AlphaFold3 project site: https: / / github.com / google-deepmind / alphafold3 Execute the command: export mnt= / mnt / sdb4 docker run -it \ --volume $mnt / af_input: / root / af_input \ --volume $mnt / af_output: / root / af_output \ --volume $mnt / alphafold3-model: / root / models \ --volume $mnt / alphafold3-db: / root / public_databases \ --gpus all --privileged \ af3:latest \ python run_alphafold.py \ --json_path= / root / af_input / cal.json \ --model_dir= / root / models --output_dir= / root / af_output S42: Structural Quality Filtering: Analyze the prediction results to obtain Local Conformation Confidence Degree (pLDDT) and Global Structural Quality (PTM). Set thresholds: pLDDT > 0.80 and PTM > 0.8 to ensure protein folding stability. Sequences meeting the criteria are sorted from highest to lowest pLDDT, and the top 50 are retained. The structural quality ranking is as follows: Figure 7 .
[0046] S43: Screening for interfacial physical interactions: The electrostatic binding energy, number of interfacial atoms (distance from ligand ≤ 10 Å), and positive-to-negative charge pairing ratio of these 50 complexes were calculated using the APBS tool. Thresholds were set: number of interfacial atoms > 50, charge pairing ratio > 0.6, to ensure tight and complementary binding interfaces. Sequences meeting the criteria were sorted by electrostatic binding energy (the more negative, the better), and the first 15 were retained.
[0047] APBS tool: https: / / github.com / Electrostatics / apbs Execute the command: apbs your_input_file.in The core content of your_input_file.in is as follows: Figure 8 The calculation results are as follows Figure 9 .
[0048] S44: Solubility Screening: Finally, the SolMPNN model was used to predict the global solubility score of the above 15 sequences. The sequences were sorted from highest to lowest solubility score, and the top 7 sequences were selected for experimental verification.
[0049] SolMPNN project address: https: / / github.com / dauparas / ProteinMPNN (select solvation weights) Execute command python . / protein_mpnn_run.py \ --use_soluble_model \ --path_to_fasta $path_to_fasta \ --pdb_path $path_to_PDB \ --pdb_path_chains "$chains_to_design" \ --out_folder $output_dir \ --num_seq_per_target 5 \ --sampling_temp "0.1" \ --score_only 1 \ --seed 13 \ --batch_size 1 The calculation results are as follows Figure 10 .
[0050] Thus, through four steps of fine screening, seven final candidate sequences were identified from 100 sequences, all of which demonstrated excellent performance in terms of structural stability, binding affinity, and solubility.
[0051] S5, Wet Test Verification To verify the effectiveness of this design method, we performed in vitro expression and function tests on the above 7 final design sequences (numbers: 9854, 4613, 8798, 13013, 19137, 18423, 5685).
[0052] 1. Protein expression and solubility verification: Methods: Seven genes were cloned into expression vectors and transformed into E. coli for induced expression. The supernatant (soluble components) and precipitate (inclusion bodies) were separated by sonication and centrifugation, and the soluble proteins were purified by nickel column affinity chromatography.
[0053] Results and Comparison: As shown in the figure, SDS-PAGE analysis showed that four sequences (4613, 8798, 19137, 18423) achieved highly soluble expression.
[0054] Results of wet experiments, such as Figure 2 .
[0055] 2. Verification of heme binding activity: Methods: The four highly expressed sequences were co-incubated with heme, and the binding products were detected by high performance liquid chromatography-mass spectrometry (HPLC-MS).
[0056] Results and Comparison: The characteristic peaks of the product of the reaction between protein and heme were clearly detected in sample 19137. The retention time and mass-to-charge ratio were consistent with the expectations, confirming that it has heme binding activity.
[0057] Wet test results, such as heme binding activity Figure 3 .
[0058] The computer-readable storage medium can be an internal storage unit of any data processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.
[0059] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions or improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for designing small molecule binding proteins, characterized in that: Includes the following steps: S1. Structure acquisition steps: Obtain the structure of the complex of the target small molecule ligand and the protein; S2. Candidate sequence generation step: Based on the complex structure described in step S1, a protein generation model is used, and the parameters corresponding to the protein generation model are adjusted to generate a set of candidate protein sequences. S3. Initial screening step: Based on the candidate protein sequence set described in step S2 and the small molecule ligand described in step S1, predict the structure of the complex and screen out potential sequences from the structure of the complex. S4. Fine screening stage: Based on step S3, screen the potential sequence and the small molecule ligand in step S1 to predict the complex structure 2, and screen the complex structure 2 to select the preferred protein sequence 1. S5. Wet experiment verification step: Based on the preferred protein sequence selected in step S4, perform soluble expression test and activity test to obtain small molecule binding protein.
2. The method for designing small molecule binding proteins according to claim 1, characterized in that: In step S1, obtaining the structure of the target small molecule ligand-protein complex includes at least one of a protein structure database and a structure prediction tool.
3. The method for designing small molecule binding proteins according to claim 1, characterized in that: In step S2, the protein generation model includes a diffusion-based generation model and a structural constraint-based language model. The parameter of the diffusion-based generation model is the mask ratio, and the parameter of the structural constraint-based language model is the sampling temperature.
4. The method for designing small molecule binding proteins according to claim 1, characterized in that: Step S3 includes the following sub-steps: S31. Using a structure prediction method with moderate computational accuracy, the candidate protein sequence set is combined with the small molecule ligand in step S1 to predict the structure of a complex with one number of sequences. S32, Analysis step S31 predicts the structure of the complex and extracts intermolecular interaction parameters using the PLIP tool; S33. Perform Z-score normalization on the intermolecular interaction parameters extracted in step S32, calculate the weighted score according to the weight ratio, and retain the complex sequence with two entries in descending order of score. S34. Obtain the first sequence of the complex according to step S33, and use the FPocket tool to calculate the pocket index. S35. Perform Z-score standardization on the pocket index calculated in step S34, calculate the weighted score according to the weight ratio of 2, and retain the potential sequence of number 3 from high to low according to the score.
5. The method for designing small molecule binding proteins according to claim 4, characterized in that: In step S33, the intermolecular interaction parameters include the number of hydrogen bonds, the number of salt bridges, the number of hydrophobic interactions, and the number of π-π interactions. The weight ratio is 30% for the number of hydrogen bonds, 25% for the number of salt bridges, 25% for the number of hydrophobic interactions, and 20% for the number of π-π interactions.
6. The method and apparatus for designing small molecule binding proteins according to claim 4, characterized in that: In step S34, the pocket indicators are pocket volume, hydrophobicity, and polarity.
7. The method for designing small molecule binding proteins according to claim 4, characterized in that: In step S35, the weight ratio two is 40% for volume, 35% for hydrophobicity, and 25% for polarity.
8. The method for designing small molecule binding proteins according to claim 1, characterized in that: Step S4 includes the following sub-steps: S41. Using a high-precision structure prediction method, the potential sequence screened in step S3 is combined with the small molecule ligand in step S1 to predict the structure of the complex. S42. Based on step S41, predict the second complex structure, filter the second complex structure according to the local conformation confidence threshold and the global structure quality threshold, and retain the second complex sequence with four entries in descending order of local conformation confidence. S43. Based on the complex sequence two described in step S42, calculate the binding energy, number of interface atoms, and positive and negative charge pairing ratio of the interface using APBS. Based on the number of interface atoms and the charge pairing ratio, retain five complex sequences three in order of increasing electrostatic binding energy; the number of interface atoms is greater than 50, and the charge pairing ratio is greater than 0.
6. S44. Based on the complex sequence three retained in step S43, global solubility fraction is predicted using the SolMPNN model. The sequences with six soluble fractions are selected as the preferred protein sequence one and sorted from high to low.
9. The method for designing small molecule binding proteins according to claim 8, characterized in that: The first local conformation confidence threshold is greater than 0.8, and the second single-chain global structure quality threshold is greater than 0.
8.
10. The method for designing small molecule binding proteins according to claim 1, characterized in that: Step S5 includes the following sub-steps: S51. The preferred protein sequence described in step S4 is cloned into an expression vector, and the expression vector is transferred to a host cell to obtain recombinant engineered bacteria; S52. The recombinant engineered bacteria obtained in step S51 are induced and cultured to achieve the expression of the target protein; S53. Collect the recombinant engineered bacteria from step S52 after induced expression, and separate them by centrifugation after ultrasonic disruption, collecting the supernatant and precipitate respectively. S54. The soluble target protein in the supernatant collected in step S53 is purified by nickel column affinity chromatography to obtain the purified target protein. S55. Perform SDS-PAGE analysis on the target protein described in step S54 to obtain a highly soluble expression of the preferred protein sequence with seven sequences, which can achieve highly soluble expression. S56. The preferred protein sequence obtained in step S55 is incubated with the small molecule ligand obtained in step S1. The binding product is detected by phase chromatography-mass spectrometry to verify that the obtained small molecule binding protein has the activity of the small molecule ligand obtained in step S1, and the small molecule binding protein is obtained.