A method for in silico screening of synthetic proteins based on pre-trained models

Through pre-trained WGAN model and multimodal data integration technology, synthetic proteins with similar structure and physical and chemical characteristics to the target protein are screened out, solving the adaptability and cost problems of screening methods in the existing technology, and achieving efficient and accurate synthetic protein design.

CN117894374BActive Publication Date: 2025-07-11BEIJING INSTITUTE OF PETROCHEMICAL TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410107064.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-07-11
Estimated Expiration
2044-01-25

AI Technical Summary

Technical Problem

The existing high-throughput screening methods are insufficient to adapt to different types of compounds and proteins, and the experimental conditions are complex and costly, making it difficult to effectively screen synthetic proteins with similar structure and physical and chemical characteristics.

Method used

The pre-trained Wasserstein Generative Adversarial Network (WGAN) model was used to generate synthetic protein sequences, and combined with AlphaFold2 and ESM2 tools to calculate the skeleton distance and characteristic distance, and screened out synthetic proteins with similar structure and physical and chemical characteristics to the target protein.

Benefits of technology

It realizes efficient and accurate screening of synthetic proteins similar to the target protein, reducing experimental complexity and cost, and improving the efficiency and reliability of synthetic protein design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117894374B_ABST
    Figure CN117894374B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for synthetic protein dry screening based on a pre-trained model. First, a protein sequence is input into a pre-trained WGAN model to form a pre-trained generator; a Mask operation is performed on specific amino acids in the protein sequence, and the generator is used to imitate the target sequence to generate a batch of synthetic protein sequences; it is checked whether the solubility of the synthetic protein sequences is greater than 0.70 and whether the number of special amino acids exceeds the number of special amino acids in the target sequence plus 1; if the conditions are met, first calculate the backbone distance and the feature distance between the synthetic protein sequence and the target sequence, sort these two distances in ascending order respectively, and select the corresponding synthetic protein according to the sorting results as the final screening result. This method utilizes the comprehensive information of structural and physicochemical properties and can effectively predict and screen synthetic proteins that are more similar to the target sequence in terms of structure and physicochemical properties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of bioinformatics and computational biology, and in particular to a method for in silico screening of synthetic proteins based on a pre-trained model. Background Art

[0002] Bioinformatics and computational biology have developed rapidly in the past few decades. Through tools such as the Human Genome Project, high-throughput sequencing technology, and machine learning, the genome, protein structure, gene expression, and metabolic networks have been successfully decoded, driving the frontiers of biological research. The progress in this interdisciplinary field provides key tools and insights for a deeper understanding of living systems, disease mechanisms, and personalized medicine.

[0003] High-throughput screening in the prior art uses automated equipment and large-scale experimental platforms. HTS allows for rapid and parallel functional screening of a large number of protein variants, which includes applications in biology, drug discovery, and enzyme engineering. Specifically, using a kinetic target-guided synthesis (TGS) platform, a library was designed through high-throughput screening, and four PPIMs were assembled on the target protein Bcl-XL, achieving the regulation of the Bcl-XL / BH3 interaction, highlighting the effectiveness of kinetic TGS in the identification and synthesis of high-quality PPIMs. The specific implementation steps are as follows:

[0004] 1. Synthesis of reactive fragments and acylsulfonamides: Synthesize reactive fragments and acylsulfonamides and ensure their purity and structure meet expectations.

[0005] 2. Expression and purification of Bcl-XL protein: Express wild-type and mutant Bcl-XL fusion proteins and purify them.

[0006] 3. Incubation of Bcl-XL with reactive fragments: In a 96-well plate, add acyl sulfate and sulfonyl azide building blocks to the Bcl-XL solution, and then incubate for 6 hours at a temperature of 37°C.

[0007] 4. Liquid chromatography-mass spectrometry analysis (LC / MS-SIM): Perform LC / MS-SIM analysis on the incubated samples, using a ZorbaxSB-C18 column, a Phenomenex C18 guard column, and positive selected ion mode for mass spectrometry detection. By analyzing the mass spectrometry data, identify TGS hit compounds, including mass and retention time.

[0008] 5. Liquid chromatography-mass spectrometry analysis of the control group: Incubate the same combination of acyl sulfate and sulfonyl azide building blocks in a buffer without Bcl-XL. Then perform LC / MS-SIM analysis for comparison with the chromatogram of the incubation containing Bcl-XL.

[0009] 6. Compare the synthetic sulfonamide with the sulfonamide in the incubation with Bcl-XL: The synthetic sulfonamide is subjected to LC / MS-SIM analysis and compared with the sulfonamide in the incubation containing Bcl-XL.

[0010] 7. Fluorescence polarization competitive binding assay: A fluorescence polarization competitive binding assay method is used for functional evaluation.

[0011] However, the above-mentioned scheme has the following disadvantages in the implementation process:

[0012] 1. Technical limitations: The existing technology may be limited by specific sample types, and different optimizations may be required for different types of compounds and proteins;

[0013] 2. Selection of experimental conditions: Incubation conditions and experimental designs may need to be carefully optimized, otherwise it may lead to low reaction efficiency or uncontrolled side reactions;

[0014] 3. Complexity and cost: The use of advanced technologies such as liquid chromatography-mass spectrometry analysis may increase the complexity and cost of the experiment, and professional equipment and skills are required. Summary of the Invention

[0015] The object of the present invention is to provide a method for dry screening of synthetic proteins based on a pre-trained model, which utilizes comprehensive information on structure and physicochemical properties and can effectively predict and screen synthetic proteins that are more similar to the target sequence in terms of structure and physicochemical properties.

[0016] The object of the present invention is achieved by the following technical solutions:

[0017] A method for dry screening of synthetic proteins based on a pre-trained model, the method comprising:

[0018] Step 1. First, input the target into a pre-trained Wasserstein Generative Adversarial Network (WGAN) model, save the pre-trained parameters of the generator, and form a pre-trained generator;

[0019] Step 2. Perform a Mask operation on specific amino acids in the target protein sequence, use the pre-trained generator in Step 1 to imitate the target protein sequence to generate a batch of synthetic protein sequences, and number them;

[0020] Step 3. After obtaining the synthetic protein sequences, first check whether the solubility of the synthetic protein sequences is greater than 0.70 and whether the number of special amino acids exceeds the number of special amino acids in the target protein sequence plus 1;

[0021] Step 4: If the conditions in Step 3 are met, proceed to Step 5 for rescreening the synthetic protein sequences; if the conditions in Step 3 are not met, return to Step 2 for generating protein sequences with masks.

[0022] Step 5: In the rescreening operation phase, first calculate the backbone distance and feature distance between the synthetic protein sequence and the target protein sequence, sort these two distances in ascending order respectively, and select the corresponding synthetic protein as the final screening result according to the sorting results.

[0023] As can be seen from the technical solution provided by the present invention above, the above method utilizes comprehensive information on structure and physicochemical properties, can effectively predict and screen synthetic proteins that are more similar to the target sequence in terms of structure and physicochemical properties, provides a comprehensive and efficient strategy for the design and optimization of synthetic proteins, and is expected to promote further exploration and innovation in the field of protein engineering. Brief Description of the Drawings

[0024] In order to more clearly illustrate the technical solution of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 Schematic flowchart of the method for dry screening of synthetic proteins based on a pre-trained model provided by an embodiment of the present invention;

[0026] Figure 2 Schematic structural diagram of the WGAN model described in an embodiment of the present invention. Detailed Embodiments

[0027] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments, which does not constitute a limitation to the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention.

[0028] As Figure 1 shown is a schematic flowchart of the method for dry screening of synthetic proteins based on a pre-trained model provided by an embodiment of the present invention, and the method includes:

[0029] Step 1: First, input the target into a pre-trained Wasserstein Generative Adversarial Network (WGAN) model, save the pre-trained parameters of the generator, and form a pre-trained generator;

[0030] In this step, as Figure 2 shown in the structural schematic diagram of the WGAN model described in the embodiment of the present invention, the WGAN model is composed of a generator G and a discriminator D. In this structure, the generator G receives a random noise vector z as input, which is sampled from a normal distribution; the output of the generator G will be normalized through an activation function (such as sigmoid or tanh). Its goal is to generate samples similar to real data by learning the distribution characteristics of the training data, so as to deceive the discriminator D;

[0031] The discriminator D is a binary classifier used to accurately classify the given input samples into real samples from the training data and synthetic samples from the generator G;

[0032] During the training process of the WGAN, the generator G and the discriminator D engage in a game, continuously improving their own capabilities through alternating training. The generator G is committed to generating more and more realistic samples to fool the discriminator D, while the discriminator D tries to improve its accuracy so that it is difficult to distinguish between real samples and synthetic samples. This game process continues until the generator can generate samples similar enough to the real data, and at the same time, the discriminator becomes increasingly unable to accurately distinguish between real samples and synthetic samples. Through this adversarial training, the goal of the WGAN is to achieve a stable and efficient generator that can produce synthetic data with highly realistic properties. The entire training process enables the generator to learn the distribution characteristics of the real data, thereby improving the quality and diversity of the generated samples;

[0033] The distribution of the generator G on real samples is P G, by mapping the random noise vector z to the sample space in the data space, denoted as D(z; θg); G is a tunable function represented by a multi-layer perceptron with parameters θg, and its output is denoted as G(z); the discriminator D maps the input x to the data space, denoted as D(z; θd), where x is the input of the discriminator D and θd represents the parameters of the discriminator D, and its output is a scalar; when x comes from real samples, the output of the discriminator D is D(x), and when x comes from the samples generated by the generator G, the output of the discriminator D is D(G(z)); when training the discriminator D, the goal is to make the value of D(x) as large as possible and the value of D(G(z)) as small as possible; while when training the generator G, it is to maximize the value of D(G(z)) as much as possible and minimize the value of log(1 - D(G(z))) at the same time; in other words, the discriminator D and the generator G are playing a two-player minimization game, and its objective function is V(G, D), expressed as:

[0034]

[0035] where, and denote the expectation operation, taken from the real data distribution p data (x) and the noise distribution p z (z) respectively; x ∼ p data (x) means sampling x from the real data distribution p data (x), where x comes from the samples of real data; denotes taking the expectation of logD(x) over p data (x), that is, taking the expectation of the logarithmic probability of the discriminator output for real data; z ∼ p z (z) means sampling z from the noise distribution p z (z), where z comes from the samples of the prior noise distribution; denotes taking the expectation of log(1 - D(G(z))) over p z (z), that is, taking the expectation of the logarithmic probability that the samples generated by the generator are judged to be false by the discriminator; these two expectation terms respectively constitute the objective function of the generative adversarial network, where the optimization objectives of the generator G and the discriminator D respectively involve the minimization and maximization of these two expectations.

[0036] In WGAN, the Wasserstein distance is introduced to measure the difference between the generated sample distribution and the real sample distribution, making the gradient signal smoother. The expression of the Wasserstein distance is as follows:

[0037]

[0038] where, ∏(P G ,Pdata ) represents that the edge is the generated sample distribution P G and the true sample distribution P data The set of all combined joint distributions γ (x,y) In each joint distribution, sample x and y, and calculate the expected value E of the distance between this pair of samples (x,y)~γ [||x - y||]. The minimum expected value obtained from all joint distributions is the Wasserstein distance;

[0039] In other words, the Wasserstein distance measures the minimum cost required to convert the distribution P G to the distribution P data This cost is calculated by the function D(x); the D(x) function is defined as the product of the distance and the mass for transferring unit mass from one position to another. In WGAN, weight clipping is used to implement the calculation of D(x). Lipschitz continuity is a property that requires the slope of the function not to exceed a fixed constant K throughout the domain. When K = 1, it is called 1-Lipschitz, that is, the derivative is always less than 1. In the discriminator, this means that for any pair of input samples, the difference between the discriminator outputs does not exceed a certain predetermined value. To implement weight clipping, after each update of the discriminator parameters in WGAN, all weight values are clipped to a predetermined range [-c, c]. When the weight is outside the predetermined range, we clip it to c or -c, so that D(x) satisfies the Lipschitz continuity condition. The objective function of WGAN is constructed using the Kantorovich-Rubinstein duality:

[0040]

[0041] Through a stable training process, WGAN can better optimize the Wasserstein distance cost function, making the generated sample distribution of the generator gradually approach the true sample distribution.

[0042] However, in order to satisfy the Lipschitz continuity constraint, the method of restricting the weight range of the discriminator by weight clipping in WGAN may still lead to training instability and gradient vanishing problems. Therefore, WGAN-GP is proposed to use gradient penalty instead of weight clipping in WGAN. A penalty factor λ is introduced to control the intensity of the gradient penalty, restricting the gradient magnitude of the discriminator to prevent gradient explosion or vanishing. The objective function of WGAN-GP can be expressed as:

[0043]

[0044] Among them, is a random interpolation between the real samples and the generated samples; is this interpolation distribution; λ is the weight of the gradient penalty term; by adding a gradient penalty term to the discriminator D during training:

[0045]

[0046] it can keep the gradient within a reasonable range, while avoiding the training instability and gradient sparsity problems that may be caused by weight clipping. WGAN-GP can also better stabilize the training and improve the quality of the generated samples.

[0047] In the specific implementation, in the pre-trained WGAN-GP network, the structural parameters of the generator model are shown in Table 1.

[0048] Table 1 Structural parameters of generator G model

[0049]

[0050] Input the protein sequence into the pre-trained WGAN model, save the pre-trained parameters of the generator, and form a pre-trained generator. The specific process is as follows:

[0051] First, by calling the WGAN model, input the random noise into the WGAN model. During the process of the generator G and the discriminator D playing against each other, the ability of the generator G to imitate the target protein sequence is continuously enhanced; after the training ends, the generator G in the WGAN model is already a generator that can generate synthetic proteins highly similar to the target protein sequence;

[0052] After that, use the torch.save function in the PyTorch library to save the weight parameters of the pre-trained generator to the specified file path.

[0053] Here, the PyTorch library is an open-source deep learning framework, which has become one of the first choices in research and practical applications with its dynamic computational graph and rich tool support. It is called "torch" in Chinese and is generally called the PyTorch library.

[0054] Step 2: Perform a Mask operation on specific amino acids (such as amino acid key sites) in the target protein sequence, use the generator pre-trained in Step 1 to imitate the target protein sequence to generate a batch of synthetic protein sequences, and number them;

[0055] In this step, first, use the PyTorch library to index and fill the tensor of the target protein sequence, set the elements at the specified index positions to 0, and adjust its shape;

[0056] Next, create a tensor of synthetic protein sequence clones, and also perform index filling on this tensor. By setting the information of specific dimensions to zero, perform Mask processing on specific amino acids;

[0057] Input the random noise vector and the target protein sequence into the generator that has been pre-trained in step 1. The generator imitates the target protein sequence to generate a batch of synthetic protein sequences that are different from the target protein sequence in structure and number them.

[0058] Step 3: After obtaining the synthetic protein sequence, first check whether the solubility of the synthetic protein sequence is greater than 0.70 and whether the number of special amino acids (C, N) exceeds the number of special amino acids in the target protein sequence plus 1;

[0059] Step 4: If the conditions of step 3 are met, enter step 5 for the rescreening operation of the synthetic protein sequence; if the conditions of step 3 are not met, return to step 2 for the generation operation of the protein sequence with Mask;

[0060] Step 5: In the rescreening operation stage, first calculate the backbone distance and the feature distance between the synthetic protein sequence and the target protein sequence. Sort these two distances in ascending order respectively, and select the corresponding synthetic protein as the final screening result according to the sorting results.

[0061] In this step, the backbone distance between the synthetic protein sequence and the target protein sequence is calculated by the AlphaFold2 tool to ensure their structural similarity; here, AlphaFold2 is an advanced protein structure prediction model developed by DeepMind, which can predict the three-dimensional structure of proteins through deep learning. PyMOL is an open-source software for molecular visualization and is widely used in the fields of biochemistry, structural biology, and computational biology.

[0062] First, submit the two protein sequences through the AlphaFold2 online prediction tool and obtain the corresponding structure predictions; then use the align command of PyMOL to perform the superposition alignment of the protein structures to calculate the backbone distance between the target protein sequence and the synthetic protein sequence; finally, determine their structural similarity based on the calculated backbone distance between the two protein sequences;

[0063] The characteristic distance between the synthetic protein sequence and the target protein sequence is calculated based on the pre-trained protein language model ESM2, and the consistency of the synthetic protein sequence and the target sequence in physicochemical properties is verified according to the characteristic distance. Here, the protein language model ESM2 (Evolutionary Scale Modeling 2) is based on the Transformer architecture and has been unsupervised pre-trained on a large protein sequence database. Through deep learning means, it explores the evolutionary laws of protein sequences and the complex relationships between sequence-structure-function. The calculation formula is as follows:

[0064]

[0065] Among them, ESM_RMSD represents the root mean square error, which is used to measure the difference between two vectors of the target protein sequence A and the synthetic protein sequence B; r i is the output feature value obtained by mapping the target protein sequence A and the synthetic protein sequence B through the protein language model ESM2; N is the embedding dimension, which is a parameter of the protein language model ESM2; represents the sum of all dimensions from i = 1 to N; ||A - B|| represents the Euclidean distance between vectors A and B, that is, the sum of the squares of the differences in each corresponding dimension;

[0066] This formula (6) is often used to compare the similarity or difference between two vectors, especially in the fields of structural alignment, molecular simulation, etc. to compare the similarity degree between different structures or conformations.

[0067] In the specific implementation, the specific process of selecting the corresponding synthetic protein as the final screening result according to the sorting result is as follows:

[0068] If the same numbered protein sequence is selected in the first position in the two sorting tables, the synthetic protein corresponding to this number is the synthetic protein closest to the target sequence finally screened out;

[0069] If different numbered protein sequences are selected in the first position, then further consider the solubility of these two numbered protein sequences, and select the synthetic protein corresponding to the number with solubility close to the target sequence as the final screening result.

[0070] It should be noted that the content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those skilled in the art.

[0071] The following is a specific example to illustrate the method described in the embodiments of the present invention. The Protein A sequence used in the example is derived from the immunoglobulin-binding B region of Protein A of Staphylococcus aureus. The protein structure of this Staphylococcus aureus Protein A sequence was obtained through solution nuclear magnetic resonance (Solution NMR) experiments, and this structure is labeled as Chain A, STAPHYLOCOCCUS AUREUS PROTEIN A, with an identification number of 1BDC_A in the PDB database.

[0072] Example 1: Dry screening of synthetic proteins mimicking the Protein A sequence

[0073] 1. Generate protein sequences

[0074] In this example, by inputting a protein sequence into a pre-trained WGAN model and saving the pre-trained parameters of the generator, a pre-trained generator model is formed. Using Protein A as the target protein sequence, 9 protein sequences were synthesized by mimicking Protein A using the pre-trained generator model, and each synthetic protein sequence was numbered. The one-dimensional amino acid sequences and numbers of the 9 synthetic proteins are shown in Table 2.

[0075] Table 2 Protein A and 9 synthetic protein sequences

[0076]

[0077] 2. Preliminary screening of synthetic protein sequences

[0078] After obtaining the synthetic protein sequences, the solubility of Protein A and its synthetic protein sequences was calculated using an online protein solubility prediction tool, and the number of special amino acids C (cysteine) and N (asparagine) contained in each sequence was calculated. The preliminary screening parameter table for synthetic protein sequences is shown in Table 3.

[0079] Table 3 Preliminary screening parameter table for synthetic protein sequences

[0080]

[0081]

[0082] Check whether the solubility of each synthetic protein sequence is greater than 0.70 according to the parameter table, and whether the number of special amino acids (C, N) exceeds the number of special amino acids in the target sequence plus 1. As can be seen from Table 3, the content of special amino acids in all synthetic protein sequences meets the condition of the number of special amino acids in the target sequence plus 1. However, the solubilities of the sequences numbered A_002, A_003, A_006, and A_007 are all less than 0.70. Therefore, these four are eliminated in the preliminary selection stage, and the remaining synthetic protein sequences enter the re-screening stage.

[0083] 3. Re-screening of synthetic protein sequences

[0084] After the preliminary screening, first calculate the backbone distance and feature distance between the synthetic protein sequence and the target sequence. The backbone distance between the synthetic protein sequence and the target sequence is calculated by the AlphaFold2 tool, and the feature distance is calculated based on ESM2. The backbone distance and feature distance between the synthetic protein and Protein A are shown in Table 4.

[0085] Table 4. Backbone distance and feature distance table of synthetic protein and Protein A

[0086]

[0087] Subsequently, sort these two distances in ascending order respectively and analyze the sorting results. The ascending order table of the backbone distance between the synthetic protein and Protein A is shown in Table 5, and the ascending order table of the feature distance is shown in Table 6.

[0088] Table 5. Ascending order table of the backbone distance between the synthetic protein and Protein A

[0089]

[0090] Table 6. Ascending order table of the feature distance between the synthetic protein and Protein A

[0091]

[0092]

[0093] For the convenience of analysis, integrate the sequence numbers of the ascending order tables of the backbone distance and the feature distance into one table, as shown in Table 7.

[0094] Table 7. Ascending order table of the backbone distance and feature distance between the synthetic protein and Protein A sequence

[0095]

[0096] As can be seen from the analysis of Table 7, the protein sequence with the same number (A_008) was selected in the first position. Then the synthetic protein corresponding to this number (A_008) was identified as the one closest to the target sequence characteristics, and A_008 was used as the final screening result.

[0097] To increase the reliability of the dry screening results, the ProtParam tool was used to predict the physicochemical properties of Protein A and A_008. The physicochemical properties of Protein A and A_008 are shown in Table 8.

[0098] Table 8. Physicochemical properties of Protein A and A_008

[0099]

[0100]

[0101] By comparing the two protein sequences (Protein A and A_008) in the physicochemical property Table 8, the following conclusions can be drawn:

[0102] (1) Number of amino acid compositions: The number of amino acid compositions of the two protein sequences is the same, both being 58.

[0103] (2) Molecular weight: The molecular weight of Protein A is 6598.26, while the molecular weight of A_008 is 6707.43. This indicates that the molecular weight of A_008 is slightly higher than that of Protein A.

[0104] (3) Theoretical isoelectric point: The theoretical isoelectric point of Protein A is 5.16, while the theoretical isoelectric point of A_008 is 5.23. This indicates that A_008 is slightly more alkaline in theory.

[0105] (4) Charged residues: They show similar performance in terms of charged residues, both having 9 negatively charged residues and 7 positively charged residues.

[0106] (5) Atomic composition: A_008 has slightly more carbon, hydrogen, nitrogen, and oxygen atoms than Protein A, but the total number of atoms of the two is 932 and 919 respectively.

[0107] (6) Extinction coefficient: The extinction coefficient of Protein A is 1490, while the extinction coefficient of A_008 is 13980. This difference is very large, indicating that under certain experimental conditions, the optical density of A_008 may be higher.

[0108] (7) Estimated half-life: The estimated half-lives of both are the same in different cellular environments, namely mammalian reticulocytes, in vitro (4.4 hours), yeast, in vivo (>20 hours), Escherichia coli, in vivo (>10 hours).

[0109] (8) Instability index: The instability index of Protein A is 68.12, while that of A_008 is 49.31. The lower instability index of A_008 indicates that it may be relatively more stable.

[0110] (9) Aliphatic index: The aliphatic index of Protein A is 72.59, while that of A_008 is 62.59. This indicates that Protein A is more prominent in terms of its aliphatic properties.

[0111] (10) Grand average of hydropathicity (GRAVY): The GRAVY of Protein A is -1.102, while that of A_008 is -0.876. A lower GRAVY value indicates that the protein is more hydrophilic. Here, A_008 is relatively more hydrophilic.

[0112] Generally speaking, A_008 shows more complex, basic, strong optical response, stable, prominent aliphatic properties and is more suitable for functioning in an aqueous environment compared to Protein A in terms of larger molecular weight, higher theoretical isoelectric point, higher extinction coefficient, lower instability index, lower aliphatic index and lower hydropathicity.

[0113] Example 2: Dry screening of synthetic proteins mimicking Protein A-like sequences

[0114] 1. Generate protein sequences

[0115] By inputting the protein sequence into a pre-trained WGAN model and retaining the pre-trained parameters of the generator, a pre-trained generator model is formed. Using Protein A-like as the target protein sequence, 9 protein sequences are successfully synthesized with the help of the pre-trained generator model, and each synthetic protein sequence is numbered. Table 9 shows the one-dimensional amino acid sequences of these 9 synthetic proteins and their corresponding numbers.

[0116] Table 9 Protein A-like and 9 synthetic protein sequences

[0117]

[0118]

[0119] 2. Preliminary screening of synthetic protein sequences

[0120] After obtaining the synthetic protein sequences, calculate the solubility of Protein A-like and its synthetic protein sequences using an online protein solubility prediction tool, and calculate the number of special amino acids C (cysteine) and N (asparagine) contained in each sequence. The preliminary screening parameter table for the synthetic protein sequences is shown in Table 10.

[0121] Table 10 Preliminary screening parameter table for synthetic protein sequences

[0122]

[0123] According to the parameter table, check whether the solubility of each synthetic protein sequence is greater than 0.70 and whether the number of special amino acids (C, N) exceeds the number of special amino acids in the target sequence plus 1. As can be seen from Table 10, the content of special amino acids in the synthetic protein sequences numbered A-like_002, A-like_006, and A-like_008 does not meet the condition of the number of special amino acids in the target sequence plus 1, and the solubility of the synthetic protein sequences numbered A-like_004, A-like_005, and A-like_007 is less than 0.70. Therefore, these six are eliminated in the primary selection stage, and the remaining synthetic protein sequences enter the re-screening stage.

[0124] 3. Re-screening of synthetic protein sequences

[0125] After preliminary screening, first calculate the backbone distance and feature distance between the synthetic protein sequence and the target sequence Protein A-like. The synthetic protein sequence and the target sequence are calculated using the AlphaFold2 tool, and the feature distance is calculated based on ESM2. The backbone distance and feature distance between the synthetic protein and Protein A-like are shown in Table 11.

[0126] Table 11 Table of backbone distance and feature distance between synthetic protein and Protein A-like

[0127]

[0128] Subsequently, sort these two distances in ascending order respectively and analyze the sorting results. The ascending order table of the backbone distance between the synthetic protein and ProteinA-like is shown in Table 12, and the ascending order table of the feature distance is shown in Table 13.

[0129] Table 12 Ascending order table of backbone distance between synthetic protein and Protein A-lke

[0130]

[0131] For ease of analysis, the sequence numbers of the backbone distance and the ascending list of characteristic distances are integrated into one table, as shown in Table 14.

[0132] Table 14 Ascending table of the backbone distance and characteristic distance of the synthetic protein and the Protein A-like sequence

[0133]

[0134] Analysis of Table 14 shows that at the first position, the protein sequences A-like_001 and A-like_009 with different numbers are selected. At this time, it should be determined which of the solubilities of these two protein sequences (i.e., A-like_001 and A-like_009) is closest to the solubility of the target sequence. From Table 10, it can be seen that the solubility of the synthetic protein sequence numbered A-like_001 is 0.739, and the solubility of the synthetic protein sequence numbered A-like_009 is 0.842. Therefore, the synthetic protein corresponding to the number A-like_009 is identified as the one with the closest characteristics to the target sequence, and A-like_009 is used as the final screening result.

[0135] To increase the reliability of the dry screening results, the ProtParam tool was used to predict the physicochemical properties of Protein A-like and A-like_009. The physicochemical properties of Protein A-like and A-like_009 are shown in Table 15.

[0136] Table 15 Physicochemical property table of Protein A-like and A-like_009

[0137]

[0138]

[0139] By comparing the two protein sequences (Protein A-like and A-like_009) in the physicochemical property table 15, the following conclusions can be drawn:

[0140] (1) Number of amino acid compositions: The number of amino acid compositions of the two protein sequences is the same, both being 58.

[0141] (2) Molecular weight: The molecular weight of A-like_009 is 6806.54, while the molecular weight of Protein A-like is 6627.35. This indicates that the molecular weight of A-like_009 is larger.

[0142] (3) Theoretical isoelectric point: The theoretical isoelectric point of A-like_009 is 4.94, while that of Protein A-like is 5.16. This indicates that A-like_009 is more acidic theoretically.

[0143] (4) Charged residues: A-like_009 has a total of 10 negatively charged residues and 7 positively charged residues in terms of charged residues, while Protein A-like has 9 negatively charged residues and 7 positively charged residues.

[0144] (5) Atomic composition: A-like_009 has slightly more atoms of carbon, hydrogen, nitrogen, oxygen, and sulfur than Protein A-like, with the total number of atoms being 945 and 919 respectively.

[0145] (6) Extinction coefficient: The extinction coefficient of A-like_009 is 2980, while that of Protein A-like is 1490. The higher extinction coefficient of A-like_009 indicates a stronger signal in optical experiments.

[0146] (7) Estimated half-life: The estimated half-life of A-like_009 in mammalian reticulocytes, in vitro environment is the same as that of Protein A-like, which is 100 hours.

[0147] (8) Instability index: The instability index of A-like_009 is lower, indicating that it is relatively more stable, while the instability index of Protein A-like is higher.

[0148] (9) Aliphatic index: The aliphatic index of A-like_009 is lower, indicating that it is slightly inferior to Protein A-like in terms of aliphatic properties.

[0149] (10) Grand average of hydropathicity (GRAVY): The GRAVY value of A-like_009 is lower, indicating that it is more hydrophilic, while the GRAVY value of Protein A-like is also negative but relatively closer to hydrophilicity.

[0150] Generally speaking, A-like_009 has a larger molecular weight, a lower theoretical isoelectric point, a higher extinction coefficient, and a lower instability index compared to Protein A-like. This series of advantages indicates that A-like_009 exhibits more superior characteristics in terms of structure, activity, and stability, making it more potential and valuable in specific biological or experimental conditions.

[0151] In summary, the method described in the embodiments of the present invention has the following advantages:

[0152] 1. Accuracy of the computational model: The application of advanced deep learning and neural network technologies in the pre-trained model enables it to provide more accurate and comprehensive protein sequence and structure information, providing more reliable guidance for the design of synthetic proteins;

[0153] 2. Data-driven design: By learning large-scale biological data, the pre-trained model can discover potential patterns and interrelationships therein. This data-driven design approach makes the design of synthetic proteins more traceable, fully considering the diversity and complexity in biological data;

[0154] 3. High-throughput analysis: The training of the pre-trained model on large-scale datasets endows it with the ability to process high-throughput data. This ability enables more rapid analysis during the design and optimization of synthetic proteins, accelerating the experimental process and improving efficiency;

[0155] 4. Structure prediction: Advanced pre-trained models are expected to provide more accurate protein structure predictions, reducing the dependence of experiments on protein structures, which is crucial for more precisely understanding the characteristics of protein structures during synthesis;

[0156] 5. Multi-modal data integration: The pre-trained model has the ability to integrate multiple data sources, including protein sequence, structure, function, etc. By fusing these multi-modal data, more comprehensive analysis and design can be provided, offering a more global perspective for the design of synthetic proteins;

[0157] 6. Automatic feature extraction: The automatic feature extraction ability of deep learning models, without the need to manually specify rules, enables them to better adapt to different types of proteins and reactions, providing an efficient and adaptive solution for feature extraction during the screening process;

[0158] 7. The dry screening of the experimental environment is completely carried out online, avoiding the problems of unreasonable or unoptimized conditions that may occur in wet experiments, which may lead to non-specific side reactions. It not only does not require a large amount of cost expenditure, but also has a simple and easy-to-operate process.

[0159] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of suggestion that this information constitutes the prior art known to those skilled in the art.

Claims

1. A method for in vitro screening of synthetic proteins based on a pre-trained model, characterized in that, The method includes the following steps: Step 1: First, input the target protein sequence into the Wasserstein Generative Adversarial Network (WGAN) model, save the pre-trained parameters of the generator, and form a pre-trained generator. In Step 1, specifically, input the target protein sequence into the WGAN model, save the pre-trained parameters of the generator, and form a pre-trained generator. The specific process is as follows: First, by calling the WGAN model, input random noise into the WGAN model. During the process of the generator G and the discriminator D playing against each other, the ability of the generator G to imitate the target protein sequence is continuously enhanced. After the training is completed, the generator G in the WGAN model is already a generator that can generate synthetic proteins highly similar to the target protein sequence. After that, use the torch.save function in the PyTorch library to save the weight parameters of the pre-trained generator to the specified file path. Step 2: Perform a Mask operation on specific amino acids in the target protein sequence. Use the pre-trained generator in Step 1 to imitate the target protein sequence to generate a batch of synthetic protein sequences and number them. Step 3: After obtaining the synthetic protein sequences, first check whether the solubility of the synthetic protein sequences is greater than 0.70 and whether the number of special amino acids exceeds the number of special amino acids in the target protein sequence plus 1. Step 4: If the conditions in Step 3 are met, enter Step 5 for the rescreening operation of the synthetic protein sequences; if the conditions in Step 3 are not met, return to Step 2 for the generation operation of the protein sequences with Mask. Step 5: In the rescreening operation stage, first calculate the backbone distance and the feature distance between the synthetic protein sequences and the target protein sequence, sort these two distances in ascending order respectively, and select the corresponding synthetic proteins as the final screening results according to the sorting results.

2. The method for in vitro screening of synthetic proteins based on a pre-trained model according to claim 1, wherein, In Step 1, the WGAN model is composed of a generator G and a discriminator D. The generator G receives a random noise vector z as input, which is sampled from a normal distribution. The output of the generator G will be normalized through an activation function. Its goal is to generate samples similar to the real data by learning the distribution characteristics of the training data, thereby deceiving the discriminator D. The discriminator D is a binary classifier used to accurately classify the given input samples into real samples from the training data and synthetic samples from the generator G. The distribution of the generator G over real samples is P G , which maps a random noise vector z to the sample space in the data space, denoted as D(z; θg); G is a tunable function represented by a multi-layer perceptron with parameters θg, and its output is denoted as G(z); the discriminator D maps the input x to the data space, denoted as D(z; θd), where x is the input of the discriminator D and θd represents the parameters of the discriminator D, and its output is a scalar; when x comes from real samples, the output of the discriminator D is D(x), and when x comes from the samples generated by the generator G, the output of the discriminator D is D(G(z)); when training the discriminator D, the goal is to make the value of D(x) as large as possible and the value of D(G(z)) as small as possible; when training the generator G, the goal is to maximize the value of D(G(z)) as much as possible and minimize the value of log(1 - D(G(z))) at the same time; its objective function is V(G, D), which is expressed as: Among them, and represent the desired operations, which are sampled from the true data distribution p data (x) and the noise distribution p z (z); x ~ p data (x) means sampling x from the true data distribution p data (x), where x is a sample from the true data; means taking the expectation of logD(x) over p data (x), that is, taking the expectation of the logarithmic probability of the discriminator output for the true data; z ~ p z (z) means sampling z from the noise distribution p z (z), where z is a sample from the prior noise distribution; means taking the expectation of log(1 - D(G(z))) over p z (z), that is, taking the expectation of the logarithmic probability that the samples generated by the generator are judged as false by the discriminator; these two expectation terms respectively constitute the objective function of the generative adversarial network, where the optimization objectives of the generator G and the discriminator D respectively involve the minimization and maximization of these two expectations.

3. The method for in vitro screening of synthetic proteins based on a pre-trained model according to claim 1, wherein The specific process of Step 2 is as follows: First, use the PyTorch library to perform index filling on the tensor of the target protein sequence, set the elements at the specified index positions to 0, and adjust its shape. Then create a tensor that clones the target protein sequence, also perform index filling on this tensor, and perform Mask processing on specific amino acids by setting the information of specific dimensions to zero. Input the random noise vector and the target protein sequence after Mask processing into the pre-trained generator in Step 1. The generator imitates the target protein sequence to generate a batch of synthetic protein sequences with structures similar to but different from the target protein sequence and number them.

4. The method for in vitro screening of synthetic proteins based on a pre-trained model according to claim 1, characterized in that, In step 5, the backbone distance between the synthetic protein sequence and the target protein sequence is calculated by the AlphaFold2 tool to ensure their structural similarity based on the backbone distance. Specifically: First, submit the two protein sequences through the AlphaFold2 online prediction tool and obtain the corresponding structure predictions; then use the align command of PyMOL to perform the superposition alignment of the protein structures to calculate the backbone distance between the target protein sequence and the synthetic protein sequence; finally, determine their structural similarity based on the calculated backbone distance between the two protein sequences. The feature distance between the synthetic protein sequence and the target sequence is calculated based on the pre-trained protein language model ESM2. The consistency of the synthetic protein sequence and the target sequence in physicochemical properties is verified according to the feature distance. The calculation formula is as follows: Among them, ESM_RMSD represents the root mean square deviation, which is used to measure the difference between two vectors of the target protein sequence A and the synthetic protein sequence B; r i is the output eigenvalue obtained by mapping the target protein sequence A and the synthetic protein sequence B through the protein language model ESM2; N is the embedding dimension and is a parameter of the protein language model ESM2; It means summing over all dimensions from i = 1 to N; ||A - B|| represents the Euclidean distance between vectors A and B, that is, the sum of the squares of the differences in each corresponding dimension.

5. The method for in vitro screening of synthetic proteins based on a pre-trained model according to claim 1, wherein In step 5, the specific process of selecting the corresponding synthetic protein as the final screening result according to the sorting result is as follows: If the protein sequences with the same number are selected in the first position in the two sorting tables, the synthetic protein corresponding to this number is the synthetic protein with the closest characteristics to the target sequence finally screened out; If the protein sequences with different numbers are selected in the first position, the solubility of these two numbered protein sequences is further considered, and the synthetic protein corresponding to the number with solubility close to the target sequence is selected as the final screening result.

Citation Information

Patent Citations

  • Engineered chimeric fusion protein compositions and methods of use thereof

    CN116916940A

  • Method and device for training protein prediction model based on graph neural network

    CN116935952A