Methods of polypeptide design using combined masked language modeling and yeast surface display and sequences
By combining mask language modeling and yeast surface display-based peptide design methods, the problems of long development cycles and high costs in peptide drug development are solved. This approach enables efficient generation of peptide candidates, improves the success rate and efficiency of peptide design, and is applicable to drug development and biomaterial design.
Patent Information
- Application Number
- CN202311074037.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-24
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2043-08-24
AI Technical Summary
Traditional peptide drug development is characterized by long development cycles, high costs, and low success rates, making it difficult to quickly respond to sudden outbreaks of pandemics such as pneumonia. Existing technologies are also unable to efficiently generate peptide candidates with specific properties or functions.
A peptide design method combining joint mask language modeling and yeast surface display was adopted. The pre-trained model was reconstructed by self-supervised masking, and peptide candidates were generated by combining downstream task fine-tuning and molecular dynamics simulation. Protein expression and affinity determination were then performed using yeast display technology.
It enables the efficient generation of a large number of peptide candidates, shortens the research and development cycle, and improves the success rate. It can quickly evaluate the properties and functions of peptides through a dry-wet design process, and is suitable for drug development and biomaterial design.
Smart Images

Figure CN117133358B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of peptide design technology, specifically to a peptide design method and sequence based on joint masking language modeling and yeast surface display. Background Technology
[0002] The virus has had a huge impact on public health systems worldwide. New variants are constantly emerging due to evolutionary mutations, which greatly challenges the traditional pharmaceutical industry and places higher demands on the timeliness of drug design.
[0003] Compared to small molecule drugs, peptide drugs have many advantages, including a shorter evaluation period for their medicinal value, and a unique advantage in combating pandemic infectious diseases such as pneumonia. Therefore, developing efficient and high-success-rate peptide design processes is of great significance.
[0004] With the convergence of artificial intelligence and natural sciences, natural language processing (NLP) technology has been proven applicable to protein sequence processing. In recent years, language models have been transferred to the protein domain, enabling effective characterization of protein sequences and demonstrating excellent performance in important downstream tasks such as secondary structure prediction. Therefore, this invention designs a peptide design workflow based on a protein language model, achieving efficient and low-cost peptide design and wet experimental characterization.
[0005] However, new drug development has always faced challenges such as lengthy cycles, high costs, and low success rates. Traditional drug development speeds are insufficient to cope with sudden outbreaks like those of pandemics like pneumonia. Therefore, developing an efficient drug design process is essential. Given the advantages of peptide drugs, this study applies natural language processing technology to the protein field, constructing a complete AI-based peptide design process and combining it with wet experimental characterization, aiming to accelerate the peptide drug development process and provide insights for rapid response to pandemics. Summary of the Invention
[0006] The present invention proposes a peptide design method and sequence that combines mask language modeling and yeast surface display, which can at least solve one of the technical problems in the background art.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] A peptide design method combining mask language modeling and yeast surface display is implemented through the following steps:
[0009] The publicly available protein sequence database is cleaned, and protein sequences that meet the requirements are selected as the training set for the language model. Mask language modeling is then performed on the protein sequences contained in the training set, i.e., a pre-trained model is established for mask reconstruction pre-training.
[0010] Based on the pre-trained model, a downstream task is designed and set. The parameters of the pre-trained model are updated through the training of the downstream task. The set property information is stored in the model parameters, which is to perform fine-tuning of the downstream task.
[0011] A new polypeptide sequence is obtained by randomly masking residues in a selected reference sequence and predicting the masked residues.
[0012] Virtual screening of peptide candidates generated by the model is performed using manually set rules and molecular dynamics simulations.
[0013] The protein expression level and affinity of the screened peptides were determined using yeast display technology.
[0014] Furthermore, the mask reconstruction pre-training step specifically includes:
[0015] First, the publicly available protein sequence database is cleaned according to the length of the target polypeptide;
[0016] Next, a self-supervised mask reconstruction task is performed on the selected training set. Specifically, each protein sequence is segmented into individual residues, which have a chance of being masked. The model's task is to predict the masked residues using the remaining context.
[0017] Furthermore, the specific steps for fine-tuning the downstream task include:
[0018] The downstream task's objective provides the pre-trained model with information about the set properties or functions, enabling the pre-trained model to update its parameters and store the set properties or functions during the training process.
[0019] Furthermore, the predicted masked residues specifically include,
[0020] First, a reference polypeptide sequence is selected, and then the residues in the sequence are randomly masked, thus obtaining an incomplete polypeptide sequence. The incomplete sequence is then completed using a fine-tuned mask reconstruction model, and the masked residues are predicted, thereby obtaining a large number of polypeptide sequences.
[0021] Furthermore, the virtual screening steps are as follows:
[0022] First, screening is conducted based on manually set rules. Second, calculations are performed based on molecular dynamics simulations to further reduce the number of peptides to be used in wet experiments.
[0023] Furthermore, the determination of protein expression levels and affinity of the screened peptides using yeast display technology specifically includes,
[0024] First, by demonstrating the integration of highly efficient fluorescent proteins into the system, direct qualitative and quantitative studies of peptide expression levels were conducted using laser confocal microscopy and flow cytometry.
[0025] Secondly, the affinity between the designed peptide and the target protein was directly determined using a flow cytometry-activated cell sorting method.
[0026] On the other hand, the present invention also discloses a polypeptide sequence generated using the aforementioned joint masking language modeling and yeast surface display polypeptide design method.
[0027] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0028] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0029] As can be seen from the above technical solutions, the peptide design method of the present invention, which combines mask language modeling and yeast surface display, can generate a large number of peptide candidates that may have specific properties or functions by transferring natural language processing technology to the field of peptide generation. By combining artificial intelligence generation, virtual screening and wet experimental characterization, a relatively complete "dry and wet" design process is established.
[0030] Specifically, the advantages and benefits of this invention are mainly reflected in the following aspects:
[0031] 1. Innovative Peptide Design Method: This method innovatively utilizes protein language models to accurately predict and generate peptide sequences with specific properties. This approach holds promise for significant breakthroughs in drug development and other biotechnology fields.
[0032] 2. Large-scale generation of peptide candidates: Traditional peptide design methods may be limited by computational resources and time, making it difficult to efficiently generate a large number of diverse candidates. The method of this invention, however, can rapidly generate a large number of peptide sequences, providing more options for subsequent screening and optimization.
[0033] 3. Combined Dry and Wet Design Process: Another feature of this invention is the design process that combines artificial intelligence generation with wet experimental characterization. This combined dry and wet approach creates a good feedback loop between computational simulation and experimental verification. By combining virtual screening and wet experimental characterization, the properties and functions of peptide candidates can be evaluated efficiently, thereby identifying potential high-quality peptides more quickly.
[0034] 4. Improved peptide design success rate and significantly shortened research cycle: Peptide design is a complex task involving numerous variables and possible combinations. Using the method of this invention, peptide sequences with specific functions can be found more quickly, thus greatly improving design efficiency. More accurate prediction and screening can reduce invalid experiments and increase the success rate of peptide design.
[0035] 5. Wide Range of Applications: The technical aspects of this invention are expected to be applied in multiple fields, including drug development and biomaterial design. By generating polypeptide sequences with specific properties, more efficient and precise drugs can be created, and superior biomaterials can be developed, which will have significant application value in the fields of medicine and biotechnology.
[0036] In summary, this invention, as a peptide design method based on joint mask language modeling and yeast surface display, has advantages such as innovation, high efficiency, and wide applicability. It can greatly improve the efficiency and success rate of peptide design, and bring important impetus and contributions to the development of the biotechnology field. Attached Figure Description
[0037] Figure 1 This is an overall flowchart of an embodiment of the present invention;
[0038] Figure 2 This is a schematic diagram of the mask reconstruction pre-training process according to an embodiment of the present invention;
[0039] Figure 3 This is a schematic diagram of protein sequence masking and the generation of embedding features in an embodiment of the present invention;
[0040] Figure 4 This is a schematic diagram of masked residue prediction according to an embodiment of the present invention;
[0041] Figure 5 This is a schematic diagram of downstream task fine-tuning according to an embodiment of the present invention;
[0042] Figure 6 This is a schematic diagram of sequence generation according to an embodiment of the present invention;
[0043] Figure 7 This is a flowchart of the virtual screening process according to an embodiment of the present invention;
[0044] Figure 8 This is a flowchart of the wet screening experiment according to an embodiment of the present invention;
[0045] Figure 9 This is a schematic diagram illustrating the determination of the surface display efficiency of each sample according to an embodiment of the present invention;
[0046] Figure 10 This is a schematic diagram illustrating the results of affinity determination between candidate peptides and target proteins using laser confocal scanning microscopy and flow cytometry experiments in an embodiment. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0048] like Figure 1 As shown in this embodiment, the peptide design method combining masked language modeling and yeast surface display is implemented through the following steps: First, a publicly available protein sequence database is cleaned according to the characteristics of the designed peptide, and protein sequences that meet the requirements are selected as the training set for the language model. Second, masked language modeling is performed on the protein sequences included in the training set. This process can be summarized as randomly masking some input residues with a certain probability (usually 0.15) and predicting these masked residues, i.e., the mask reconstruction task. In this way, a deep bidirectional representation can be obtained, which is called the "pre-training" of the model. Third, a specific downstream task is designed based on the pre-trained model. The parameters of the pre-trained model are updated through training on the downstream task, and specific property information is stored in the model parameters, which is called the "fine-tuning" of the model. Fourth, random residue masking is performed on a selected reference sequence, and the masked residues are predicted to obtain a new peptide sequence. Fifth, the peptide candidates generated by the model are virtually screened through some manually set rules and molecular dynamics simulations. The sixth step involves using yeast display technology to determine the protein expression level and affinity of the screened peptides.
[0049] The following are detailed explanations:
[0050] 1. Mask Reconstruction Pre-training
[0051] like Figure 2 As shown, firstly, the publicly available protein sequence database is cleaned according to the length of the target peptide. If the target is a peptide of approximately 50 units in length, only protein sequences shorter than 60 units in the database can be used as the training set for the model. Next, a self-supervised mask reconstruction task is performed on the selected training set. Specifically, as shown... Figure 3 As shown, each protein sequence is segmented into individual residues, which have a probability (usually taken as 0.15) of being masked. The model's task is to predict the masked residues (e.g., ...) using the remaining context. Figure 4 As shown in the figure, this process enables the model to learn the distribution probability of amino acids at each position in the natural protein sequence, thus giving the model a basic understanding of natural proteins.
[0052] 2. Downstream task fine-tuning
[0053] like Figure 5 As shown, the pre-trained model has a basic understanding of natural protein sequences, but still lacks the ability to learn specific functional information. Therefore, specific downstream tasks can be designed to further train the pre-trained model. Specifically, the objectives of the downstream tasks provide the pre-trained model with specific property or functional information, allowing it to update its parameters and store this information during training. For example, if the goal is for the model to generate peptides with ideal solubility, a solubility-related downstream task can be designed for "fine-tuning," providing the solubility property to the pre-trained model and guiding it to generate peptide candidates with ideal solubility.
[0054] 3. Masked residue prediction
[0055] like Figure 6 As shown, the fine-tuned model has the ability to generate sequences with specific properties or functional information. Based on this model, sequences are "filled in" to generate new polypeptide sequences. Specifically, a reference polypeptide sequence is first selected, such as a polypeptide sequence with good solubility that can specifically bind to the original viral strain. Then, residues in this sequence are randomly masked, resulting in an incomplete polypeptide sequence. The fine-tuned mask reconstruction model is then used to complete the incomplete sequence and predict the masked residues, thereby obtaining a large number of polypeptide sequences.
[0056] 4. Virtual Filtering
[0057] like Figure 7 As shown, after generating a large number of peptide candidates, a large-scale virtual screening was performed on these peptides to reduce the number of peptides to be tested in wet experiments. First, screening was conducted based on manually defined rules, which were flexible and could be set according to the knowledge and experience of experts in specific fields. Second, calculations were performed based on molecular dynamics simulations to further reduce the number of peptides to be tested in wet experiments.
[0058] 5. Yeast Display
[0059] like Figure 8 As shown, this section describes the soluble expression and affinity screening of designed peptides using a modified Saccharomyces cerevisiae surface display system. First, this display system incorporates a highly efficient fluorescent protein, allowing for direct qualitative and quantitative studies of peptide expression levels via laser confocal microscopy and flow cytometry, eliminating the need for expensive antibody labeling methods. Second, flow cytometry-activated cell sorting (…) Fluorescence-Activated Cell Sorting,The FACS method can directly determine the affinity between designed peptides and target proteins without going through complex and time-consuming peptide synthesis / expression, purification and affinity determination processes, which greatly improves the throughput of wet experimental validation.
[0060] The following examples illustrate the experimental verification of the solubility and affinity of peptides displayed on the yeast surface in this embodiment of the invention;
[0061] To test these peptides designed using language models, this invention uses a yeast surface display method to determine the protein expression level of the designed peptides and their affinity for target proteins. Taking 15 generated and screened peptide sequences as an example, this invention first measures the surface display efficiency of each sample, such as... Figure 9 As shown, the results indicate that two candidates have similar display efficiencies to the positive control, nanobody TY1.
[0062] Furthermore, this invention uses laser confocal scanning microscopy and flow cytometry to determine the affinity between candidate peptides and target proteins. For example... Figure 10 As shown, the results indicate that among the four candidate peptides, SN193 exhibited a relatively high affinity (approximately 10 μM) for the viral Omicron BA.5RBD, while the other three showed weak binding. This invention also determined the binding strength of these peptides to the original viral RBD, finding no interaction. This result demonstrates that the designed peptides are highly targeted and efficient; this invention requires only the synthesis of a single-digit number of candidates to obtain peptide binders with relatively ideal affinity.
[0063] Among them, the four candidate polypeptide sequences in the above examples are as follows:
[0064] SN193:
[0065] ELEEQVVKIIEQVDELVHEALHKASKEDAVHIIQLEHIVNELALEVARIDDEREIRELEEEVRRLLEKVQRILN;
[0066] SN19587:DLEEQVLKIIEQMQEIMKKALKKASEDDLDGLLQFVWWAQEAAIGAMKGDGERGMRGIVAEMQRTYPYAQQAAG; SN10267:
[0067] DLEEQLLKLMEQLQEIAKELLKKLGGDGAEGIIQLVWLAQEMALGAIKSEDERGAKEAEEAMEWLYKQVLQLMKMM; SN3646:
[0068] DLEEYLFKLLRQIERILKELLRRASERELERLLLIFNLAYELILGALMSEDERVLRFLIKLLEQYLLAVLSYKMS.
[0069] In summary, the embodiments of this application have designed an efficient peptide design process based on artificial intelligence technology, which transfers natural language processing technology to the field of peptide generation. It can generate a large number of peptide candidates that may have specific properties or functions. By combining artificial intelligence generation, virtual screening and wet experimental characterization, a relatively complete "dry and wet" design process has been established.
[0070] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.
[0071] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.
[0072] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the joint masking language modeling and yeast surface display peptide design methods described in the above embodiments.
[0073] It is understood that the system provided in the embodiments of the present invention corresponds to the method provided in the embodiments of the present invention, and the explanation, examples and beneficial effects of the relevant content can be referred to the corresponding parts of the above methods.
[0074] This application also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, communication interface, and memory communicate with each other via the communication bus.
[0075] Memory, used to store computer programs;
[0076] The processor, when executing programs stored in memory, implements the aforementioned peptide design method based on joint masking language modeling and yeast surface display.
[0077] The communication bus mentioned in the aforementioned electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.
[0078] The communication interface is used for communication between the aforementioned electronic devices and other devices.
[0079] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0080] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0081] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0082] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0083] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0084] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A peptide design method combining mask language modeling and yeast surface display, characterized in that, Through the following steps, The publicly available protein sequence database is cleaned, and protein sequences that meet the requirements are selected as the training set for the language model. Mask language modeling is then performed on the protein sequences contained in the training set, i.e., a pre-trained model is established for mask reconstruction pre-training. Based on the pre-trained model, downstream tasks are designed and set. The parameters of the pre-trained model are updated through the training of the downstream tasks. The set properties or functional information are stored in the model parameters, which is to perform fine-tuning of the downstream tasks. A new polypeptide sequence is obtained by randomly masking residues in a selected reference sequence and predicting the masked residues. Virtual screening of peptide candidates generated by the model is performed using manually set rules and molecular dynamics simulations. The protein expression level and affinity of the screened peptides were determined using yeast display technology.
2. The peptide design method based on joint masking language modeling and yeast surface display according to claim 1, characterized in that: The mask reconstruction pre-training steps specifically include: First, the publicly available protein sequence database is cleaned according to the length of the target polypeptide; Next, a self-supervised mask reconstruction task is performed on the selected training set. Specifically, each protein sequence is segmented into individual residues, which have a chance of being masked. The model's task is to predict the masked residues using the remaining context.
3. The peptide design method based on joint masking language modeling and yeast surface display according to claim 2, characterized in that: The specific steps for fine-tuning the downstream task include: The downstream task's objective provides the pre-trained model with information about the set properties or functions, enabling the pre-trained model to update its parameters and store the set properties or functions during the training process.
4. The peptide design method based on joint masking language modeling and yeast surface display according to claim 3, characterized in that: The predicted masked residues specifically include, First, a reference polypeptide sequence is selected, and then the residues in the sequence are randomly masked, thus obtaining an incomplete polypeptide sequence. The incomplete sequence is then completed using a fine-tuned mask reconstruction model, and the masked residues are predicted, thereby obtaining a large number of polypeptide sequences.
5. The peptide design method based on joint masking language modeling and yeast surface display according to claim 4, characterized in that: The steps of the virtual filtering are as follows. First, screening is conducted based on manually set rules. Second, calculations are performed based on molecular dynamics simulations to further reduce the number of peptides to be used in wet experiments.
6. The peptide design method based on joint masking language modeling and yeast surface display according to claim 5, characterized in that: The determination of protein expression levels and affinity of screened peptides using yeast display technology specifically includes, First, by demonstrating the integration of highly efficient fluorescent proteins into the system, direct qualitative and quantitative studies of peptide expression levels were conducted using laser confocal microscopy and flow cytometry. Secondly, the affinity between the designed peptide and the target protein was directly determined using a flow cytometry-activated cell sorting method.
7. A computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the method as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Antibacterial peptide prediction method and device based on protein pre-training representation learning
CN112614538A
Generation method of polypeptide sequence, and training method and device of polypeptide generation model
CN115512763A