Method, apparatus, device, and storage medium for generating active peptide segments
By training the peptide generation model corresponding to specific activities, the problem of insufficient diversity of peptides with specific activities in the prior art is solved, and the generation of peptides with rich diversity with specific activities is achieved.
Patent Information
- Application Number
- CN202210278322.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-03-21
AI Technical Summary
The prior art cannot directly generate peptides with specific activities, and the diversity of peptides is insufficient.
By using the general peptide data set and the preset active peptide data set corresponding to specific activities, the peptide generation model corresponding to specific activities is obtained to generate peptides with rich diversity with specific activities.
The generation of peptides with different specific activities and rich diversity based on specific activity needs is achieved, which improves the diversity and efficiency of peptides.
Smart Images

Figure CN114783521B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of biomedicine, and particularly relates to methods, devices, equipment, and storage media for generating bioactive peptide segments. Background Art
[0002] In recent years, many bioactive peptides have shown effective therapeutic effects against various complex diseases, such as antiviral, antibacterial, and anticancer peptides. Currently, a variety of peptide drugs have been approved for marketing to treat various diseases. For example, effective results have been achieved in treating diabetes, cancer, osteoporosis, multiple sclerosis, human immunodeficiency virus (HIV) infection, and chronic pain. Therefore, it is very necessary to study bioactive peptides.
[0003] In the prior art, new bioactive peptides are mainly searched by randomly mutating existing peptides, and then new bioactive peptide segments are searched in the mutated existing peptides. This method is too blind, unable to directly generate peptide segments with specific activities, and the diversity of peptide segments is insufficient. Summary of the Invention
[0004] In view of this, the embodiments of this application provide methods, devices, equipment, and storage media for generating bioactive peptide segments to solve the problems in the prior art that peptide segments with specific activities cannot be directly generated and the diversity of peptide segments is insufficient.
[0005] The first aspect of the embodiments of this application provides a method for generating bioactive peptide segments, and the method includes:
[0006] Obtain a peptide segment generation model corresponding to a specific activity, where the peptide segment generation model is obtained by training an LSTM network using a general peptide data set and a preset active peptide data set corresponding to the specific activity, and the general peptide data set includes multiple general peptide segments;
[0007] Use the peptide segment generation model to generate multiple peptide segments with specific activities.
[0008] In the above solution, the peptide segment generation model is obtained by training an LSTM network using a general peptide data set and a preset active peptide data set corresponding to the specific activity. During the training process, it can learn the potential features in the general peptide segments and the potential activity rules in the known active peptide segments. Furthermore, during the actual use of the peptide segment generation model, peptide segments with different specific activities and rich diversity can be generated according to different specific activity requirements.
[0009] Optionally, after using the peptide segment generation model to generate multiple peptide segments with the specific activity, the method further includes:
[0010] Perform clustering analysis on multiple of the peptide segments, and generate a peptide segment set according to the analysis result, where the similarity between each peptide segment in the peptide segment set is less than a preset similarity.
[0011] In this embodiment, clustering analysis is performed on multiple peptide segments, and peptide segments with high similarity are removed according to the analysis result, so that the similarity between the remaining peptide segments is less than the preset similarity. It can be understood that the remaining peptide segments are all different from each other, thus enriching the diversity of peptide segments.
[0012] Optionally, the method further includes:
[0013] Obtain a de novo peptide library, where the de novo peptide library includes multiple peptide segments corresponding to different specific activities;
[0014] Determine a target peptide segment that acts on a specific target in the de novo peptide library.
[0015] In this embodiment, since the peptide segments included in the de novo peptide library are a large number of peptide segments with different specific activities and the activity of each peptide segment has been determined, there is no need to spend time verifying the activity of each peptide segment. On this basis, a target peptide segment that can act on a specific target can be quickly found. Compared with the prior art of finding a peptide segment that can act on a specific target among random peptide segments, this embodiment improves the search rate, takes less time, saves a large amount of resources, and reduces the economic cost.
[0016] Optionally, the determining a target peptide segment that acts on a specific target in the de novo peptide library includes:
[0017] Construct a 3D structure model corresponding to each peptide segment in the de novo peptide library;
[0018] Perform protein docking based on each 3D structure model and the specific target to obtain a docking result;
[0019] Perform molecular dynamics simulation based on the docking result to obtain the binding free energy corresponding to each 3D structure model;
[0020] Determine the target peptide segment according to each binding free energy.
[0021] In this embodiment, the target peptide segment is determined in the de novo peptide library through processes such as 3D modeling, protein docking, and molecular dynamics simulation. On the one hand, the peptide segments in the de novo peptide library have been screened, greatly narrowing the candidate range for determining the target peptide segment that acts on a specific target and improving the accuracy; on the other hand, the peptide segments in the de novo peptide library are all known for their corresponding specific activities, clarifying the action mechanism of each peptide segment, and improving the accuracy and efficiency of determining the target peptide segment. It is beneficial for further research on the screened target peptide segment, greatly saving the research cost.
[0022] Optionally, before obtaining the peptide generation model corresponding to a specific activity, the method further includes:
[0023] Training the LSTM network using the general peptide dataset to obtain a general model;
[0024] Fine-tuning the general model using the active peptide dataset and a preset loss function to obtain the peptide generation model.
[0025] In this embodiment, the LSTM network is first trained using the general peptide dataset to obtain a general model, and then the general model is fine-tuned using the active peptide dataset of a certain specific activity to obtain the peptide generation model corresponding to this specific activity. Since the fine-tuning is based on the general model, the speed of training the peptide generation model is greatly improved. And using the known active peptide segments as the training set enables the peptide generation model to learn the potential activity rules in the known active peptide segments during the training process. Thus, during the actual use of the peptide generation model, peptides with different specific activities and rich diversity can be generated according to different specific activity requirements.
[0026] Moreover, training the peptide generation model corresponding to each specific activity is beneficial to generating accurate new peptides containing this specific activity according to the peptide generation model corresponding to each specific activity.
[0027] Optionally, the active peptide dataset includes multiple sample peptides. The step of fine-tuning the general model using the active peptide dataset and a preset loss function to obtain the peptide generation model includes:
[0028] Obtaining the target matrix corresponding to each sample peptide;
[0029] Inputting each target matrix into the general model for processing to obtain the sample active peptide corresponding to each target matrix;
[0030] Calculating the loss value between each sample active peptide and the sample peptide corresponding to each sample active peptide based on the loss function;
[0031] When it is detected that the loss value is greater than a preset threshold, adjusting the model parameters of the general model during training, and continuing to train the general model during training using the active peptide dataset;
[0032] When it is detected that the loss value is less than or equal to the preset threshold, stopping training the general model during training, and determining the trained general model as the peptide generation model.
[0033] In this embodiment, fine-tuning is performed on the basis of a general model to obtain a trained peptide generation model, which greatly improves the speed of training the peptide generation model. Moreover, by using known active peptides as the training set, the peptide generation model learns the potential activity rules in the known active peptides during the training process. As a result, during the actual use of the peptide generation model, peptides with different specific activities and rich diversity can be generated according to different specific activity requirements.
[0034] The second aspect of the embodiments of the present application provides a device for generating active peptides, including:
[0035] An acquisition unit, configured to acquire a peptide generation model corresponding to a specific activity, where the peptide generation model is obtained by training an LSTM network using a general peptide data set and a preset active peptide data set corresponding to the specific activity, and the general peptide data set includes a plurality of general peptides;
[0036] A generation unit, configured to generate a plurality of peptides with the specific activity by using the peptide generation model.
[0037] The third aspect of the embodiments of the present application provides a device for generating active peptides, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is characterized in that when the processor executes the computer program, the steps of the method for generating active peptides described in the first aspect above are implemented.
[0038] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, the steps of the method for generating active peptides described in the first aspect above are implemented.
[0039] The fifth aspect of the embodiments of the present application provides a computer program product, which when running on a device for generating active peptides, causes the device for generating active peptides to execute the steps of the method for generating active peptides described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 It is a schematic flowchart of a method for generating active peptides provided by an exemplary embodiment of the present application;
[0042] Figure 2 It is a schematic diagram of peptide segment data corresponding to different specific activities shown in this application;
[0043] Figure 3 It is a schematic diagram of the lengths of peptide segments corresponding to different specific activities shown in this application;
[0044] Figure 4 It is a schematic flowchart of a method for generating bioactive peptide segments shown in another exemplary embodiment of this application;
[0045] Figure 5 It is a schematic flowchart of a method for generating bioactive peptide segments shown in still another exemplary embodiment of this application;
[0046] Figure 6 It is a schematic diagram of different 3D structure models of peptide segments shown in this application;
[0047] Figure 7 It is a specific flowchart of a method for training a peptide segment generation model shown in an exemplary embodiment of this application;
[0048] Figure 8 It is a schematic diagram of the model structure of a general model shown in an exemplary embodiment of this application;
[0049] Figure 9 It is a schematic diagram of a device for generating bioactive peptide segments provided in an embodiment of this application;
[0050] Figure 10 It is a schematic diagram of a device for generating bioactive peptide segments provided in another embodiment of this application. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0052] Bioactive peptides (BAP) are a general term for different peptides composed of 20 natural amino acids in proteins with different compositions and arrangements, ranging from dipeptides to complex linear and cyclic structures, and are multifunctional compounds derived from proteins.
[0053] In recent years, many bioactive peptides have shown effective therapeutic effects against a variety of complex diseases, such as antiviral, antibacterial, and anticancer peptides. And currently, a variety of peptide drugs have been approved for marketing to treat various diseases. For example, diabetes, cancer, osteoporosis, multiple sclerosis, HIV infection, and chronic pain, etc., have all achieved effective results.
[0054] Peptides are easier to synthesize and have lower costs compared to compounds. Therefore, in order to promote the development of peptide drugs, it is very necessary to study active peptides.
[0055] In the prior art, the generation of peptides mainly relies on randomly mutating existing peptides, and then searching for new active peptide segments among the mutated existing peptides. This method does not utilize the known active peptide data, but only performs blind random mutations, and cannot learn the potential active peptide rules in the known active peptide data, so it is impossible to automatically generate peptide segments with specific activities, and this random mutation method cannot guarantee the diversity of peptide segments.
[0056] In view of this, the embodiments of the present application provide a method for generating active peptide segments. First, a peptide segment generation model corresponding to a specific activity is obtained, and then the peptide segment generation model is used to generate multiple peptide segments with specific activities. Since this peptide segment generation model is obtained by training an LSTM network using a general peptide data set and a preset active peptide data set corresponding to a specific activity, it can learn the potential features in the general peptide segments and the potential active peptide rules in the known active peptide segments during the training process. Furthermore, during the actual use of this peptide segment generation model, peptide segments with different specific activities and rich diversity can be generated according to different specific activity requirements.
[0057] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the method for generating active peptide segments provided by an exemplary embodiment of the present application. The execution subject of the method for generating active peptide segments provided by the present application is a device for generating active peptide segments. Among them, the device includes, but is not limited to, in-vehicle computers, tablet computers, computers, personal digital assistants (PDAs), etc., and can also include various types of servers. For example, the server can be an independent server or a cloud service that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0058] As Figure 1 shown, the method for generating active peptide segments may include: S101 - S102, specifically as follows:
[0059] S101: Obtain a peptide segment generation model corresponding to a specific activity.
[0060] Specific activity is also known as designated activity. In this embodiment, specific activity is used to represent the activity designated by the user. For example, the user designates a certain activity as the specific activity, which facilitates the subsequent generation of peptide segments with this specific activity according to the specific activity.
[0061] Such activities can include Antifungal, Antibacterial, Antibiofilm, Anticancer, Anti-diabetic, AntiHIV, Antimalarial, Anti-MRSA, Antioxidant, Anti_parasite, Anti_TB, Anti-toxin, Antiviral, Insecticidal, Ion_channel, Protease inhibitors, Spermicidal, Surface_immobilized, Wound_healing and other various activities.
[0062] Exemplarily, the number of specific activities can be one or more. That is to say, the user can designate any one or more of the above activities as specific activities according to actual needs.
[0063] For example, the user designates the activity of antifungal as the specific activity according to actual needs. Specifically, when peptide segments with the activity of antifungal are needed, the user designates the activity of antifungal as the specific activity, which facilitates the subsequent generation of multiple peptide segments with the activity of antifungal according to this specific activity of antifungal.
[0064] Also for example, the user designates anti-HIV virus and protease inhibition as specific activities according to actual needs. Specifically, when peptide segments with the two activities of anti-HIV virus and protease inhibition are needed, the user designates the activity of anti-HIV virus as the specific activity, which facilitates the subsequent generation of multiple peptide segments with the activity of anti-HIV virus according to this specific activity of anti-HIV virus.
[0065] At the same time, the user can designate the activity of protease inhibition as the specific activity, which facilitates the subsequent generation of multiple peptide segments with the activity of protease inhibition according to this specific activity of protease inhibition.
[0066] For another example, the user specifies each of the above activities as a specific activity according to actual needs, and then generates multiple peptide segments with specific activities corresponding to each specific activity. This is only an exemplary illustration and is not limited thereto.
[0067] Different specific activities correspond to different peptide segment generation models. For example, if the specific activity is antifungal, the corresponding peptide segment generation model can be an antifungal peptide segment generation model; if the specific activity is antiviral, the corresponding peptide segment generation model can be an antiviral peptide segment generation model. This is only an exemplary illustration and is not limited thereto.
[0068] In this embodiment, the peptide segment generation model is obtained by training the LSTM network using a general peptide dataset and a preset active peptide dataset corresponding to a specific activity. Among them, the general peptide dataset includes multiple general peptide segments. The Long Short-Term Memory (LSTM) network is a special recurrent neural network (RNN) that can learn long-term dependence information. Since the network structure of the LSTM network can control which information is passed to the next node through the hidden state, important information can pass through consecutive nodes unchanged, and in this way, the potential features in the general peptide segments and the potential activity rules in the known active peptide segments can be effectively learned.
[0069] Pre-train different peptide segment generation models according to different specific activities, and each trained peptide segment generation model corresponds to a specific activity. The trained different peptide segment generation models can be stored in the database of the device for generating active peptide segments, or stored in other devices.
[0070] After determining the specific activity, search for the peptide segment generation model corresponding to this specific activity in the database of this device. Or, after determining the specific activity, send a query instruction to other devices. The query instruction is used to search for the peptide segment generation model corresponding to this specific activity in other devices, and the other devices send the found peptide segment generation model to this device, and this device receives the peptide segment generation model.
[0071] S102: Generate multiple peptide segments with specific activities using the peptide segment generation model.
[0072] For example, when the peptide segment generation model is an antitoxin peptide segment generation model, the multiple peptide segments generated using the antitoxin peptide segment generation model all have the specific activity of antitoxin; when the peptide segment generation model is an antiparasitic peptide segment generation model, the multiple peptide segments generated using the antiparasitic peptide segment generation model all have the specific activity of antiparasitic. This is only an exemplary illustration and is not limited thereto.
[0073] Since the peptide generation model is trained by using a general peptide dataset and an active peptide dataset corresponding to a specific activity preset for an LSTM network, the network structure of the trained peptide generation model is basically the same as that of the LSTM network. That is, the peptide generation model may include two layers of LSTM networks and one layer of fully connected layer (such as a dense layer Dense), and use the Softmax function as the activation function of the output node (that is, the final data is output through the Softmax function).
[0074] Exemplarily, when the peptide generation model is started, it can generate multiple random amino acid sequences, and use these amino acid sequences as the input of the peptide generation model. Specifically, it is input into the first layer of LSTM network for processing. For each amino acid sequence, the first layer of LSTM network grows multiple amino acid sequences according to the amino acid sequence, and uses the input amino acid sequence and the grown multiple amino acid sequences as the output result of the first layer of LSTM network, and inputs it into the second layer of LSTM network for processing.
[0075] The processing process of the second layer of LSTM network is similar to that of the first layer of LSTM network, which will not be elaborated here. The second layer of LSTM network inputs the processing result into the dense layer. The dense layer performs matrix operations on the content input by the second layer of LSTM network, and outputs the next amino acid of the peptide after classification through the Softmax function. As the amino acid sequence extends, a sequence with a specified length is finally output.
[0076] It should be noted that among the multiple generated peptides with specific activities, the length of each peptide can be the same or different. A termination symbol can be preset in advance. During the peptide generation process, when the termination symbol is first detected, all the amino acid sequences before the termination symbol are output as the final peptide; or the termination symbol is removed, and the part after removing the termination symbol is output.
[0077] For example, set Z in the sequence as the termination symbol, and set the specified length of the peptide to 50.
[0078] During the peptide generation process, when the termination symbol Z is detected, the termination symbol Z is removed, and different peptide sequences within the specified length are intercepted and output from the part after removing the termination symbol Z.
[0079] For example, when the peptide generation model is started, it uses the symbol U as the starting sequence. A column of starting U in one batch forms a column of UUU…UUU. The peptide generation model generates the following peptides according to the starting column of UUU…UUU:
[0080] UNCYNGFFCCRPCNKPGCCNTGCCGYNCGAGKCVVLKZZZZZZZZZZZZZZZ
[0081] UYLFGGISSVLGKVVGHLVSHIVPHIVPHIVKLZZZZZZZZZZZZZZZZZZZ
[0082] UDTERCSSCCGKNCVLYATCASTMCSSDYKLGLAGHVGQGIGVSVFIPKNPZ
[0083] …
[0084] UIGGIVLTCLGTMLGGVLKKVFQKVKEAYRNZZZZZZZZZZZZZZZZZZZZZZ
[0085] URIGAKVCYCKCTFCVGVCTNNGPCCYTDVVCGLCKNZZZZZZZZZZZZZZZZ
[0086] UGLGATVRSVLGSVAPHVLPHVVPVIAEHLZZZZZZZZZZZZZZZZZZZZZZZ
[0087] Here is only an exemplary illustration. In the actual usage process, the amino acid sequences are far more than this, and correspondingly, the number of generated peptide segments is also far more than this. Specifically, the number of peptide segments generated by each peptide segment generation model can be set by oneself. For example, it can be set to generate 5,000, 50,000 peptide segments, etc., and there is no limitation on this.
[0088] In the above implementation method, first obtain the peptide segment generation model corresponding to a specific activity, and then use the peptide segment generation model to generate multiple peptide segments with specific activities. Since this peptide segment generation model is trained on the LSTM network using a general peptide dataset and a preset active peptide dataset corresponding to a specific activity, it can learn the potential features in the general peptide segments and the potential activity rules in the known active peptide segments during the training process. Furthermore, in the actual process of using this peptide segment generation model, peptide segments with different specific activities and rich diversity can be generated according to different specific activity requirements.
[0089] Exemplarily, please refer to Figure 2 , Figure 2 which is a schematic diagram of peptide segment data corresponding to different specific activities shown in this application. That is Figure 2 what is shown in
[0090] such as Figure 2As shown, the horizontal axis represents specific activities such as Antifungal, Antibiofilm, Anticancer, Anti - diabetic, AntiHIV, Antibacterial, Anti - MRSA, Antioxidant, Anti - parasite, Anti - TB, Anti - toxin, Antiviral, Insecticidal, Ion_channel, Protease inhibitors, Spermicidal, Surface_immobilized, Wound_healing, etc. The vertical axis represents the number of peptide segments corresponding to each specific activity.
[0091] As can be clearly seen from Figure 2 it, through the peptide segment generation models corresponding to different specific activities, a large number of peptide segments corresponding to each specific activity are generated, ensuring the diversity of peptide segments.
[0092] Exemplarily, please refer to Figure 3 , Figure 3 which is a schematic diagram of the lengths of peptide segments corresponding to different specific activities shown in this application.
[0093] As Figure 3 shown, the horizontal axis represents specific activities such as Antifungal, Antibiofilm, AntiHIV, Antibacterial, Antioxidant, Anti - parasite, Anti - TB, Anti - toxin, Antiviral, Insecticidal, Ion_channel, Protease inhibitors, Spermicidal, Surface_immobilized, Wound_healing, Anticancer, Anti - diabetic, Anti - MRSA, etc. The vertical axis represents the mean length of peptide segments corresponding to each specific activity, and the black thin lines on each bar represent the standard deviation of the lengths of peptide segments corresponding to each specific activity.
[0094] It can be known from the Figure 3 data that the lengths of the peptide segments corresponding to different specific activities conform to the lengths of the peptide segments that these specific activities should correspond to, that is, Figure 3 the mean value of the lengths of the peptide segments corresponding to each specific activity in
[0095] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a method for generating active peptide segments shown in another exemplary embodiment of the present application; as Figure 4 shown, the method for generating active peptide segments may include: S201 to S203, where S201 and S202 are exactly the same as S101 and S102 in the Figure 1 corresponding embodiment, and for the specific reference, please refer to the description of S101 and S102 in the Figure 1 corresponding embodiment. S203 is specifically as follows:
[0096] S203: Perform clustering analysis on multiple peptide segments and generate a peptide segment set according to the analysis result.
[0097] The similarity between each peptide segment in the peptide segment set is less than the preset similarity.
[0098] Each peptide segment includes multiple amino acid sequences, and the similarity between two peptide segments refers to the similarity between the amino acid sequences included in the two peptide segments. The preset similarity can be set by the user according to the actual situation. For example, the preset similarity can be set to 90%, 80%, 60%, 45%, etc. This is only for illustrative purposes and is not limited herein.
[0099] Perform clustering analysis on multiple peptide segments corresponding to a certain specific activity by using a clustering algorithm to obtain an analysis result. This analysis result includes multiple groups, and the similarity between each peptide segment in each group is greater than or equal to the preset similarity.
[0100] For example, use methods such as the K-means algorithm and the hierarchical clustering method to perform clustering analysis on multiple peptide segments corresponding to a certain specific activity, and aggregate the similar peptide segments among these peptide segments to obtain multiple groups. The specific clustering analysis process can refer to the prior art and will not be elaborated here.
[0101] Retain one peptide segment in each group, remove the remaining peptide segments, and generate a peptide segment set according to the retained peptide segments. Since most of the similar peptide segments have been removed, the similarity between each peptide segment in the peptide segment set is less than the preset similarity.
[0102] In this embodiment, clustering analysis is performed on multiple peptide segments, and peptide segments with high similarity are removed according to the analysis results, so that the similarity between the remaining peptide segments is less than a preset similarity. It can be understood that the remaining peptide segments are all different from each other, thus enriching the diversity of peptide segments.
[0103] Optionally, in a possible implementation manner, for multiple peptide segments generated by a peptide segment generation model corresponding to a specific activity, peptide segments that are exactly the same among the multiple peptide segments can be deleted first, and / or peptide segments that are exactly the same as those in the active peptide dataset used for training the peptide segment generation model can be deleted. Then, clustering analysis is performed on the remaining peptide segments after deleting the exactly same peptide segments, and a peptide segment set is generated according to the analysis results.
[0104] In this embodiment, peptide segments that are exactly the same among the multiple peptide segments are deleted first, that is, redundant peptide segments among the multiple peptide segments are filtered out, which helps to improve the efficiency of clustering analysis. Subsequently, clustering analysis is performed on the basis of removing redundant peptide segments, and peptide segments with high similarity are removed according to the analysis results, so that the similarity between the remaining peptide segments is less than a preset similarity. It can be understood that the remaining peptide segments are all different from each other, thus enriching the diversity of peptide segments.
[0105] To further demonstrate the diversity of peptide segments, please refer to Table 1.
[0106] Table 1 shows the number of peptide segments in the active peptide dataset for fine-tuning training corresponding to different specific activities, the number of remaining peptide segments after removing redundancy, and the number of non-similar peptide segments obtained after removing peptide segments with 60% similarity.
[0107] Table 1
[0108]
[0109] It can be clearly seen from Table 1 that the number of peptide segments in the active peptide dataset for fine-tuning training is small, while the number of remaining peptide segments after removing redundancy and the number of non-similar peptide segments obtained after removing peptide segments with 60% similarity are large, which proves that the diversity of peptide segments generated by the method provided in this application is very rich.
[0110] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a method for generating active peptide segments shown in another exemplary embodiment of this application; as Figure 5 shown, the method for generating active peptide segments may include: S301 to S304, where S301 and S302 are exactly the same as S101 and S102 in the Figure 1 corresponding embodiment, and for specific reference, please refer to Figure 1Regarding the descriptions of S101 and S102 in the corresponding embodiments, S303 to S304 are specifically as follows:
[0111] S303: Obtain a brand-new peptide library.
[0112] The brand-new peptide library includes multiple peptides corresponding to different specific activities respectively.
[0113] In S101 and S102, according to the peptides corresponding to different specific activities to generate models, multiple peptides corresponding to each specific activity are generated, and all these peptides are collected and stored in a database to obtain a brand-new peptide library.
[0114] Optionally, in order to improve the speed of subsequently determining the target peptides acting on a specific target in the brand-new peptide library, it is also possible to obtain the peptide set corresponding to each specific activity after S203, collect the peptides in each peptide set and store them in a database to obtain a brand-new peptide library.
[0115] Since highly similar peptides have been removed when generating the peptide set, building a brand-new peptide library based on this peptide set, without the interference of highly similar peptides, greatly improves the speed when subsequently determining the target peptides.
[0116] S304: Determine the target peptides acting on a specific target in the brand-new peptide library.
[0117] Target: In medicine, during certain radiotherapy, the radiation irradiates from different directions and converges on the lesion site, and this lesion site is called the target. A specific target is a designated target. For example, a specific target can be a designated XX virus.
[0118] The number of specific targets can be one or more. That is to say, the user can specify one or more as specific targets according to actual needs.
[0119] It is possible to determine the target peptides acting on a specific target in the brand-new peptide library by performing 3D modeling, protein docking, molecular dynamics simulation, molecular simulation, etc. on each peptide in the brand-new peptide library. For a certain specific target, the number of target peptides acting on this specific target determined in the brand-new peptide library depends on the actual situation and can be one or more.
[0120] In this embodiment, since the peptides included in the brand-new peptide library are a large number of peptides with different specific activities and the activity of each peptide has been clarified, there is no need to spend time verifying the activity of each peptide. On this basis, the target peptides that can act on specific targets can be quickly found. Compared with the prior art of finding peptides that can act on specific targets in random peptides, this embodiment improves the search rate, takes less time, saves a large amount of resources, and reduces the economic cost.
[0121] Optionally, in some possible implementation manners of the present application, the above S304 may include S3041 to S3044, specifically as follows:
[0122] S3041: Construct a 3D structure model corresponding to each peptide in the de novo peptide library.
[0123] Exemplarily, a modeling tool can be used to perform 3D modeling on each peptide in the de novo peptide library to obtain a 3D structure model of each peptide. Among them, the modeling tool can adopt trRosetta, Alphafold2, etc.
[0124] For example, install the modeling tool trRosetta in advance, input each peptide in the de novo peptide library into the modeling tool for processing, and output a 3D structure model corresponding to each peptide. For the specific modeling process, please refer to the prior art and will not be elaborated here. Since trRosetta utilizes rich structural data, the 3D structure models of each peptide output have diversity.
[0125] To further demonstrate the diversity of the 3D structure models of peptides, take the peptide with antiviral activity as an example for illustration. Please refer to Figure 6 , Figure 6 which is a schematic diagram of different 3D structure models of peptides shown in the present application.
[0126] As Figure 6 shown, 3D modeling is performed on the peptide with antiviral activity, and multiple different 3D structure models are obtained, which fully demonstrates that for the peptides generated by the peptide generation model, the corresponding 3D structure models have diversity. It can be understood that only for display in the figure, and the actual types of 3D structure models are far more than this.
[0127] S3042: Based on each 3D structure model and a specific target, perform protein docking to obtain a docking result.
[0128] Protein docking is an operation used to verify whether a 3D structure model and a specific target can interact. Protein docking can include rigid docking and flexible docking.
[0129] Perform rigid docking on the 3D structure model corresponding to each peptide and a specific target to obtain a rigid docking result; this rigid docking result is used to provide an initial complex conformation for flexible docking. Based on the rigid docking result, perform flexible docking to obtain a complex model corresponding to the 3D structure model of each peptide.
[0130] For example, the rigid docking of each 3D structural model with a specific target can be achieved through rigid docking software (such as zdock). Specifically, a docking pocket is set on the specific target, and the parameters of the docking pocket are set in zdock. For example, the docking pocket is defined as the amino acids within 1 nanometer around the ligand. Then, the rigid docking objects (each 3D structural model and the specific target) are uploaded to zdock, and zdock performs rigid docking on each 3D structural model and the specific target respectively, and outputs the rigid docking results of each 3D structural model and the specific target. It can be understood that zdock constructs a relatively reasonable starting structure for the binding of each peptide segment and the specific target.
[0131] After that, flexible docking is achieved through flexible docking software (such as rosetta). Specifically, the rigid docking results are input into rosetta, and rosetta can find a suitable binding mode by performing small-scale perturbations, thereby outputting the complex models of the 3D structural models corresponding to each peptide segment.
[0132] S3043: Perform molecular dynamics simulations based on the docking results to obtain the binding free energies corresponding to each 3D structural model.
[0133] Perform molecular dynamics simulations on each complex model to obtain simulation results; then perform metadynamics simulations on the simulation results through the metadynamics model to output the free energy surfaces corresponding to the 3D structural models of each peptide segment. Calculate the binding free energy corresponding to each peptide segment through the free energy surface. The specific calculation method can refer to the prior art and will not be elaborated here.
[0134] S3044: Determine the target peptide segment according to each binding free energy.
[0135] Convert the binding free energies corresponding to the 3D structural models of each peptide segment into affinities, and determine the target peptide segment according to the affinities corresponding to the 3D structural models of each peptide segment. The greater the affinity, the easier it is for the peptide segment to react with the specific target; the smaller the affinity, the less likely the peptide segment is to react with the specific target.
[0136] Specifically, a conversion software can be obtained in the network, and each binding free energy is input into the conversion software, and the conversion software outputs the affinity corresponding to each binding free energy.
[0137] Sort each peptide segment according to the magnitude of the affinity, and determine the target peptide segment in the sorting result. For example, sort multiple peptide segments in descending order of affinity, and select the peptide segment at the top of the sorting and determine it as the target peptide segment.
[0138] For another example, multiple peptide segments are sorted in ascending order of affinity, and the last peptide segment in the sorting is selected and determined as the target peptide segment. This is only an exemplary illustration and is not limited thereto.
[0139] In this embodiment, the target peptide segment is determined in a brand-new peptide segment library through processes such as 3D modeling, protein docking, molecular dynamics simulation, and molecular simulation. On the one hand, the peptide segments in the brand-new peptide segment library have been screened, greatly narrowing the candidate range for determining the target peptide segment acting on a specific target and improving the accuracy; on the other hand, the peptide segments in the brand-new peptide segment library all have their corresponding specific activities known, clarifying the action mechanism of each peptide segment and enhancing the accuracy and efficiency of determining the target peptide segment. It is beneficial for subsequent further research on the screened target peptide segment and greatly saves the research cost.
[0140] Please refer to Figure 7 , Figure 7 which is a specific flowchart of the method for training a peptide segment generation model shown in an exemplary embodiment of the present application; optionally, in some possible implementation manners of the present application, before performing the method as shown in Figure 1 , a method for training a peptide segment generation model may further be included. The method for training a peptide segment generation model may include: S401 - S402, specifically as follows:
[0141] S401: Train the LSTM network using a general peptide data set to obtain a general model.
[0142] The general peptide data set includes multiple general peptide segments.
[0143] A large number of peptides are obtained from a peptide database (such as the peptideatlas database), and the peptide segments corresponding to the obtained peptides are extracted. To improve the efficiency and accuracy of training the general model, the duplicate peptide segments and / or peptide segments containing non-standard amino acids in these peptide segments can be deleted, and a general peptide data set is generated based on the remaining large number (such as 3,274,675) of peptide segments after deletion.
[0144] The general model is trained based on the LSTM network as the basic model. Therefore, the network structure of the general model is similar to that of the LSTM network. That is, the general model may include two layers of LSTM networks and one layer of fully connected layer (such as a dense layer Dense), and the Softmax function is used as the activation function of the output node (that is, the final data is output through the Softmax function).
[0145] Among them, both the first-layer LSTM network and the second-layer LSTM network contain nodes, and each node is used to input or output the potential features in the learned peptide segments. The number of nodes can be set by the user according to the actual situation. For example, in this embodiment, the first-layer LSTM network contains 256 nodes, and the second-layer LSTM network also contains 256 nodes.
[0146] To effectively alleviate the overfitting problem of the peptide generation model, a regularization network (Dropout) is set for the first-layer LSTM network and the second-layer LSTM network. The network layer without adding Dropout needs to learn each node in the network, while the network layer after adding Dropout only needs to train the nodes in the network layer that are not masked.
[0147] The value corresponding to Dropout can be set by the user according to the actual situation. For example, in this embodiment, the value corresponding to Dropout in the first-layer LSTM network can be set to 0.5, and the value corresponding to Dropout in the second-layer LSTM network can be set to 0.3.
[0148] The loss function used to train the general model can be the multi-classification loss function (categorical_crossentropy) loss1. The number of iterations (epoch) can be set by the user according to the actual situation. For example, in this embodiment, the number of iterations can be set to 22.
[0149] In the actual training process, each general peptide is first converted to obtain a matrix corresponding to each general peptide, and the amino acids in the matrix are represented in one-hot form. Each matrix is input into the LSTM network for processing, and the LSTM network outputs the active peptide corresponding to each matrix.
[0150] The loss value corresponding to the LSTM network is calculated through the multi-classification loss function loss1, and the loss value corresponding to this LSTM network is the loss value between the general peptide and the active peptide corresponding to this general peptide.
[0151] Specifically, when generating the active peptide corresponding to the general peptide, it is generated by sequentially generating the amino acid sequence. The LSTM network in the training will predict the next amino acid sequence, and the loss value between the currently actually generated amino acid sequence and the predicted amino acid sequence is calculated through the multi-classification loss function to obtain the loss value corresponding to each amino acid sequence. The loss values corresponding to each amino acid sequence are superimposed to obtain the loss value between the general peptide and the active peptide corresponding to this general peptide.
[0152] When it is detected that the loss value is greater than the preset loss threshold, it is determined that the LSTM network currently under training has not met the requirements. At this time, adjust the network parameters (such as weight values) of the LSTM network under training, and continue to train the LSTM network under training using the general peptide dataset.
[0153] When it is detected that the loss value is less than or equal to the preset loss threshold, it is determined that the LSTM network currently under training meets the requirements. At this time, fix the network parameters in the LSTM network, and determine the LSTM network with fixed network parameters as the trained general model.
[0154] This general model can be used to generate brand-new bioactive peptide segments, providing a basis for subsequently training a peptide segment generation model using this general model.
[0155] For ease of understanding, the process of training the general model is described in conjunction with the accompanying drawings. Please refer to Figure 8 , Figure 8 which is a schematic diagram of the model structure of the general model shown in an exemplary embodiment of the present application.
[0156] Figure 8 Amino acid 1, amino acid 2, and amino acid 3 in correspond to 3 general peptide segments. For any one of the general peptide segments, input the amino acid sequence corresponding to the general peptide segment (such as amino acid 1) into the first-layer LSTM network for processing. For each amino acid sequence, the first-layer LSTM network learns the potential features in the amino acid sequence and transmits the learned potential features to the second-layer LSTM network. The second-layer LSTM network further learns the deep potential features and inputs the processing result into the dense layer. The dense layer performs matrix operations on the content input by the second-layer LSTM network and outputs the generated peptide sequence (such as amino acid X, amino acid Y, amino acid Z) after classification by the Softmax function, that is, outputs the bioactive peptide segment corresponding to the general peptide segment.
[0157] Among them, each of the first-layer LSTM network and the second-layer LSTM network contains 256 nodes. The value corresponding to Dropout in the first-layer LSTM network can be set to 0.5, and the value corresponding to Dropout in the second-layer LSTM network can be set to 0.3.
[0158] S402: Fine-tune the general model using the bioactive peptide dataset and the preset loss function to obtain a peptide segment generation model.
[0159] Different specific activities correspond to different bioactive peptide datasets. Each bioactive peptide dataset includes multiple sample peptide segments with the same specific activity. Furthermore, the general model can be fine-tuned according to each bioactive peptide dataset to obtain a peptide segment generation model corresponding to the specific activity.
[0160] A large number of peptides with known activities are obtained from an active peptide database (such as the Antimicrobial Peptide Database). The peptides in this database are classified into multiple peptide data sets according to different activities, such as the Antifungal peptide data set, the Antibacterial peptide data set, the Antibiofilm peptide data set, the Anticancer peptide data set, the Anti-diabetic peptide data set, the AntiHIV peptide data set, the Antimalarial peptide data set, the Anti-MRSA peptide data set, the Antioxidant peptide data set, the Anti_parasite peptide data set, the Anti_TB peptide data set, the Anti-toxin peptide data set, the Antiviral peptide data set, the Insecticidal peptide data set, the Ion_channel peptide data set, the Protease inhibitors peptide data set, the Spermicidal peptide data set, the Surface_immobilized peptide data set, the Wound_healing peptide data set, etc.
[0161] When a peptide segment generation model with a specific activity needs to be trained, the peptide data set corresponding to this specific activity is obtained from the active peptide database, and the sample peptide segments corresponding to each peptide in this peptide data set are extracted.
[0162] The peptide segment generation model is trained based on a general model. Therefore, the network structure of the peptide segment generation model is similar to that of the LSTM network. That is, the peptide segment generation model can include two layers of LSTM networks and one layer of fully connected layer (such as the Dense layer), and the Softmax function is used as the activation function of the output node (that is, the final data is output through the Softmax function).
[0163] The process of training the general model with the active peptide data set to obtain the peptide segment generation model is similar to the process of training the LSTM network with the general peptide data set to obtain the general model, and can refer to the description in S401.
[0164] By using the active peptide data set corresponding to each specific activity through the above method to fine-tune the general model, multiple peptide segment generation models with different specific activities can be obtained, and each peptide segment generation model can be used to generate multiple peptide segments with its corresponding specific activity.
[0165] In this embodiment, the LSTM network is first trained using a general peptide dataset to obtain a general model, and then the general model is fine-tuned using an active peptide dataset with a certain specific activity to obtain a peptide segment generation model for this specific activity. Since the fine-tuning is performed on the basis of the general model, the speed of training the peptide segment generation model is greatly improved. Moreover, by using known active peptide segments as the training set, the peptide segment generation model learns the potential activity rules in the known active peptide segments during the training process. As a result, during the actual use of the peptide segment generation model, peptide segments with different specific activities and rich diversity can be generated according to different specific activity requirements.
[0166] Moreover, training the corresponding peptide segment generation model for each specific activity is beneficial to generating accurate new peptide segments containing this specific activity according to the peptide segment generation model corresponding to each specific activity.
[0167] Optionally, in some possible implementation manners of the present application, the above S402 may include S4021 to S4044, which are specifically as follows:
[0168] S4021: Obtain the target matrix corresponding to each sample peptide segment.
[0169] Convert each sample peptide segment to obtain the target matrix corresponding to each sample peptide segment. Each sample peptide segment is composed of several amino acids, and the amino acids in the converted target matrix are represented in one-hot form.
[0170] S4022: Input each target matrix into the general model for processing to obtain the sample active peptide segment corresponding to each target matrix.
[0171] The specific processing process of the general model for the target matrix is the same as the processing process of the LSTM network for the matrix corresponding to the general peptide segment. For specific reference, please refer to the description in S401 and will not be elaborated here.
[0172] S4023: Calculate the loss value between the sample active peptide segment and the sample peptide segment corresponding to the sample active peptide segment based on the loss function.
[0173] The loss function used for training the peptide segment generation model is the same as the loss function used for training the general model. The loss function used for training the peptide segment generation model can also be a multi-classification loss function (categorical_crossentropy) loss2.
[0174] Specifically, when generating the sample active peptide corresponding to the sample peptide segment, it is generated by sequentially generating the amino acid sequence. The general model in training will predict the next amino acid sequence, calculate the loss value between the currently actually generated amino acid sequence and the predicted amino acid sequence through the multi-classification loss function loss2, obtain the loss value corresponding to each amino acid sequence, and superimpose the loss values corresponding to each amino acid sequence to obtain the loss value between the sample peptide segment and the sample active peptide corresponding to the sample peptide segment.
[0175] Compare the size of this loss value with a preset threshold. When the loss value is greater than the preset threshold, execute S4024; when the loss value is less than or equal to the preset threshold, execute S4025.
[0176] S4024: When it is detected that the loss value is greater than the preset threshold, adjust the model parameters of the general model in training, and continue to train the general model in training using the active peptide data set.
[0177] When it is detected that the loss value is greater than the preset threshold, it is determined that the general model currently being trained has not met the requirements. At this time, adjust the model parameters (such as weight values) of the general model in training, and continue to train the general model in training using the active peptide data set. That is, return to execute S4021~S4043 until it is detected in S4024 that the loss value is less than or equal to the preset threshold, and then execute S4025.
[0178] S4025: When it is detected that the loss value is less than or equal to the preset threshold, stop training the general model in training, and determine the trained general model as the peptide segment generation model.
[0179] When it is detected that the loss value is less than or equal to the preset threshold, it is determined that the general model currently being trained meets the requirements. At this time, fix the model parameters in the general model, and determine the general model with fixed model parameters as the trained peptide segment generation model.
[0180] It can be understood that the general model and the peptide segment generation model can be pre-trained by the device for generating active peptide segments, or can be pre-trained by other devices and then transplant the files corresponding to the general model and the peptide segment generation model to the device for generating active peptide segments.
[0181] In the above embodiments, fine-tuning is performed on the basis of the general model to obtain the trained peptide segment generation model, which greatly improves the training speed of the peptide segment generation model. And using the known active peptide segments as the training set enables the peptide segment generation model to learn the potential activity rules in the known active peptide segments during the training process. Furthermore, in the actual process of using the peptide segment generation model, peptide segments with different specific activities and rich diversity can be generated according to different specific activity requirements.
[0182] Optionally, in a possible implementation, when training a peptide generation model corresponding to a specific activity, a batch of peptides with the specific activity can be generated by the peptide generation model, and this batch of peptides can be used as a training set to continue training the peptide generation model, and this process can be repeated multiple times. That is, the peptides generated in one iteration are used as the input for fine-tuning and training the peptide generation model, so that more and more peptides with specific activities can be generated finally, which is beneficial to obtaining peptides with high affinity acting on specific targets in the follow-up.
[0183] In summary, in the technical solution provided by this application, the peptide generation model is obtained by training an LSTM network using a general peptide data set and a preset active peptide data set corresponding to a specific activity. During the training process, it can learn the potential features in the general peptides and the potential activity rules in the known active peptides. Furthermore, in the actual process of using this peptide generation model, peptides with different specific activities and rich diversity can be generated according to different specific activity requirements.
[0184] Optionally, cluster analysis can be performed on the generated multiple peptides, and peptides with high similarity can be removed according to the analysis results, so that the similarity between the remaining peptides is less than the preset similarity. It can be understood that the remaining peptides are all different from each other, thus enriching the diversity of the peptides.
[0185] Optionally, a brand-new peptide library is generated based on the peptides with different specific activities, and the target peptides acting on specific targets are determined in the brand-new peptide library. Since the peptides included in the brand-new peptide library are a large number of peptides with different specific activities, and the activity of each peptide has been determined, there is no need to spend time verifying the activity of each peptide. On this basis, the target peptides that can act on specific targets can be quickly found. Compared with the prior art of finding peptides that can act on specific targets in random peptides, this embodiment improves the search rate, consumes less time, saves a large amount of resources, and reduces the economic cost.
[0186] Optionally, through processes such as 3D modeling, protein docking, molecular dynamics simulation, and molecular simulation, the target peptides are determined in the brand-new peptide library. On the one hand, the peptides in the brand-new peptide library have been screened, greatly narrowing the candidate range for determining the target peptides acting on specific targets and improving the accuracy; on the other hand, the peptides in the brand-new peptide library all have their corresponding specific activities known, and the action mechanism of each peptide is clear, improving the accuracy and efficiency of determining the target peptides. It is beneficial to further study the screened target peptides, greatly saving the research cost.
[0187] In the method for training a peptide segment generation model provided by this application, first, a general peptide dataset is used to train an LSTM network to obtain a general model, and then the general model is fine-tuned through an active peptide dataset with a certain specific activity to obtain a peptide segment generation model with this specific activity. Since it is fine-tuned based on the general model, the speed of training the peptide segment generation model is greatly improved. Moreover, by using known active peptide segments as the training set, the peptide segment generation model learns the potential activity rules in the known active peptide segments during the training process. Therefore, during the actual use of the peptide segment generation model, peptide segments with different specific activities and rich diversity can be generated according to different specific activity requirements.
[0188] Moreover, training a corresponding peptide segment generation model for each specific activity is beneficial to generating accurate new peptide segments containing this specific activity according to the peptide segment generation model corresponding to each specific activity.
[0189] Please refer to Figure 9 , Figure 9 which is a schematic diagram of a device for generating active peptide segments provided in an embodiment of this application. Each unit included in the device for generating active peptide segments is used to execute Figure 1 , Figure 4 , Figure 5 , Figure 7 the respective steps in the corresponding embodiments. Specifically, please refer to Figure 1 , Figure 4 , Figure 5 , Figure 7 the relevant descriptions in their respective corresponding embodiments. For the sake of convenience, only the parts related to this embodiment are shown. Refer to Figure 9 , including:
[0190] An acquisition unit 510, configured to acquire a peptide segment generation model corresponding to a specific activity, where the peptide segment generation model is obtained by training an LSTM network using a general peptide dataset and a preset active peptide dataset corresponding to the specific activity, and the general peptide dataset includes a plurality of general peptide segments;
[0191] A generation unit 520, configured to generate a plurality of peptide segments with the specific activity by using the peptide segment generation model.
[0192] Optionally, the device further includes:
[0193] A clustering unit, configured to perform clustering analysis on a plurality of the peptide segments and generate a peptide segment set according to the analysis result, where the similarity between each peptide segment in the peptide segment set is less than a preset similarity.
[0194] Optionally, the device further includes:
[0195] A peptide library acquisition unit for acquiring a brand-new peptide library, where the brand-new peptide library includes multiple peptides corresponding to different specific activities respectively;
[0196] A determination unit for determining a target peptide acting on a specific target in the brand-new peptide library.
[0197] Optionally, the determination unit is specifically configured to:
[0198] Construct a 3D structure model corresponding to each peptide in the brand-new peptide library;
[0199] Perform protein docking based on each 3D structure model and the specific target to obtain a docking result;
[0200] Perform molecular dynamics simulation based on the docking result to obtain the binding free energy corresponding to each 3D structure model;
[0201] Determine the target peptide according to each binding free energy.
[0202] Optionally, the device further includes:
[0203] A first training unit for training the LSTM network using the general peptide dataset to obtain a general model;
[0204] A second training unit for fine-tuning the general model using the active peptide dataset and a preset loss function to obtain the peptide generation model.
[0205] Optionally, the second training unit is specifically configured to:
[0206] Obtain a target matrix corresponding to each sample peptide;
[0207] Input each target matrix into the general model for processing to obtain a sample active peptide corresponding to each target matrix;
[0208] Calculate the loss value between each sample active peptide and the sample peptide corresponding to each sample active peptide based on the loss function;
[0209] When it is detected that the loss value is greater than a preset threshold, adjust the model parameters of the general model in training and continue to train the general model in training using the active peptide dataset;
[0210] When it is detected that the loss value is less than or equal to the preset threshold, stop training the general model in training and determine the trained general model as the peptide generation model.
[0211] Please refer to Figure 10 , Figure 10The figure is a schematic diagram of a device for generating active peptide segments provided by another embodiment of the present application. As Figure 10 shown, the device 6 for generating active peptide segments in this embodiment includes: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60. When the processor 60 executes the computer program 62, it implements the steps in the above method embodiments for generating active peptide segments, such as Figure 1 S101 to S102 shown. Alternatively, when the processor 60 executes the computer program 62, it implements the functions of each unit in the above embodiments, such as Figure 9 the functions of the units 510 to 520 shown.
[0212] Exemplarily, the computer program 62 can be divided into one or more units. The one or more units are stored in the memory 61 and executed by the processor 60 to complete the present application. The one or more units can be a series of computer instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 62 in the device 6 for generating active peptide segments. For example, the computer program 62 can be divided into an acquisition unit and a generation unit, and the specific functions of each unit are as described above.
[0213] The device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand that Figure 10 this is only an example of the device 6 for generating active peptide segments and does not constitute a limitation on the device. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the device may further include input / output devices, network access devices, buses, etc.
[0214] The so-called processor 60 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0215] The memory 61 may be an internal storage unit of the device, such as the hard disk or memory of the device. The memory 61 may also be an external storage terminal of the device, such as a plug-in hard disk equipped on the device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 61 may also include both the internal storage unit of the device and the external storage terminal. The memory 61 is used to store the computer instructions and other programs and data required by the terminal. The memory 61 may also be used to temporarily store the data that has been output or will be output.
[0216] An embodiment of the present application also provides a computer storage medium. The computer storage medium may be non-volatile or volatile. The computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments for generating active peptide segments are implemented.
[0217] The present application also provides a computer program product. When the computer program product runs on a device, the device is caused to execute the steps in the above-mentioned method embodiments for generating active peptide segments.
[0218] An embodiment of the present application also provides a chip or integrated circuit. The chip or integrated circuit includes: a processor, configured to call and run a computer program from a memory, so that a device equipped with the chip or integrated circuit executes the steps in the above-mentioned method embodiments for generating active peptide segments.
[0219] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In practical applications, the above-mentioned functions may be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system may refer to the corresponding process in the foregoing method embodiments and will not be described herein again.
[0220] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0221] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0222] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for generating active peptide segments, characterized in that, Including: Obtain a peptide generation model corresponding to a specific activity, where the peptide generation model is obtained by training an LSTM network using a general peptide dataset and a preset active peptide dataset corresponding to the specific activity. The general peptide dataset includes multiple general peptide segments; during the training process, the LSTM network learns the latent features in the multiple general peptide segments and the latent activity rules in the active peptide dataset corresponding to the specific activity; Use the peptide generation model to generate multiple peptides with the specific activity; Perform clustering analysis on the multiple peptides to obtain an analysis result; the analysis result includes multiple groups, and the similarity between each peptide segment in each group is greater than or equal to a preset similarity; retain one peptide segment in each group and remove the remaining peptide segments; generate a peptide segment set according to all the retained peptide segments; the similarity between each peptide segment in the peptide segment set is less than the preset similarity.
2. The method according to claim 1, characterized in that, The method further includes: Obtain a brand-new peptide library, where the brand-new peptide library includes multiple peptide segments corresponding to different specific activities respectively; Determine a target peptide acting on a specific target in the brand-new peptide library.
3. The method according to claim 2, characterized in that, The determining a target peptide acting on a specific target in the brand-new peptide library includes: Construct a 3D structure model corresponding to each peptide segment in the brand-new peptide library; Perform protein docking based on each 3D structure model and the specific target to obtain a docking result; Perform molecular dynamics simulation based on the docking result to obtain the binding free energy corresponding to each 3D structure model; Determine the target peptide according to each binding free energy.
4. The method according to any one of claims 1 to 3, characterized in that, Before obtaining the peptide generation model corresponding to the specific activity, the method further includes: Train the LSTM network using the general peptide dataset to obtain a general model; Fine-tune the general model using the active peptide dataset and a preset loss function to obtain the peptide generation model.
5. The method according to claim 4, characterized in that, The active peptide dataset includes multiple sample peptide segments. The fine-tuning the general model using the active peptide dataset and a preset loss function to obtain the peptide generation model includes: Obtain a target matrix corresponding to each sample peptide segment; Input each target matrix into the general model for processing to obtain a sample active peptide corresponding to each target matrix; Calculate the loss value between each sample active peptide and the sample peptide segment corresponding to each sample active peptide based on the loss function; When it is detected that the loss value is greater than a preset threshold, adjust the model parameters of the general model during training and continue to train the general model during training using the active peptide dataset; When it is detected that the loss value is less than or equal to the preset threshold, stop training the general model during training and determine the trained general model as the peptide generation model.
6. An apparatus for generating active peptide segments, characterized in that, Including: An acquisition unit for acquiring a peptide generation model corresponding to a specific activity, where the peptide generation model is obtained by training an LSTM network using a general peptide dataset and a preset active peptide dataset corresponding to the specific activity, and the general peptide dataset includes a plurality of general peptide segments; the LSTM network learns the latent features in the plurality of general peptide segments and the latent activity rules in the active peptide dataset corresponding to the specific activity during the training process; A generation unit for generating a plurality of peptides with the specific activity by using the peptide generation model; Performing clustering analysis on the plurality of peptides to obtain an analysis result; the analysis result includes a plurality of groups, and the similarity between each peptide in each group is greater than or equal to a preset similarity; retaining one peptide in each group and removing the remaining peptides; generating a peptide set according to all the retained peptides; the similarity between each peptide in the peptide set is less than the preset similarity.
7. An equipment for generating active peptide segments, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 5 is implemented.
9. A chip, characterized in that, The chip includes a processor, and the processor executes a computer program stored in a memory to implement the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Functional peptide recommendation method and device and computing equipment
CN112786141A
Peptide fragment detectability prediction method
CN114093415A