Artificial intelligence-based polypeptide sequence modification method, device, equipment and medium

Through the artificial intelligence-based peptide sequence modification method, the polypeptide sequence is transformed using occlusion lexical and deep learning models to generate high-activity and low-toxicity polypeptide sequences, solving the problem of low efficiency of polypeptide drug design and achieving improvements in the quality and speed of polypeptide design.

CN119314546BActive Publication Date: 2025-08-29BEIJING YUEKANGKECHUANG PHARM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411823421.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-08-29
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

The existing polypeptide drugs have low efficiency and low quality design, making it difficult to effectively improve the activity of polypeptide drugs and reduce their toxicity.

Method used

Using artificial intelligence-based polypeptide sequence modification method, the reference polypeptide sequence is modified through occlusion meta technology and deep learning model to generate new polypeptide sequences, and the occlusion meta is used to replace some amino acid residues in the reference sequence, and combined with a multimodal deep learning model to predict the activity and toxicity of the modified polypeptide sequence.

Benefits of technology

The quality and speed of peptide design have been improved, high-activity and low-toxicity polypeptide sequences have been generated, and a peptide modification library has been established, which has improved the efficiency of peptide drug design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119314546B_ABST
    Figure CN119314546B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of polypeptide design technology, and discloses a method, device, equipment, and medium for polypeptide sequence modification based on artificial intelligence. The method comprises: obtaining a reference polypeptide sequence; using preset masking tokens to mask some tokens in the reference polypeptide sequence to obtain a masked polypeptide sequence; inputting the masked polypeptide sequence into a pre-trained polypeptide sequence modification model; determining predicted tokens of the masked portion of the masked polypeptide sequence based on the prediction results output by the polypeptide sequence modification model; replacing some of the masked tokens in the reference polypeptide sequence with the predicted tokens to obtain a predicted polypeptide sequence; if the predicted polypeptide sequence is a polypeptide sequence different from any existing polypeptide sequence, using the predicted polypeptide sequence as the modified polypeptide sequence. The present invention obtains a new polypeptide sequence based on modification of an existing polypeptide sequence, which can improve the quality and speed of polypeptide design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of polypeptide design, and in particular to a method, device, equipment and medium for polypeptide sequence modification based on artificial intelligence. Background Art

[0002] Peptide drugs are short-chain molecules composed of amino acids, lying between small-molecule drugs and large-molecule biopharmaceuticals. They show great potential in treating a variety of diseases, such as infectious diseases, cancer, and diabetes. However, current peptide drug design suffers from low efficiency and quality. Summary of the Invention

[0003] In view of this, the present invention provides an artificial intelligence-based polypeptide sequence modification method, device, equipment and medium to solve the problems of low efficiency and low quality in polypeptide drug design.

[0004] In a first aspect, the present invention provides a method for polypeptide sequence modification based on artificial intelligence, the method comprising:

[0005] Obtain reference peptide sequences;

[0006] Using a preset masking word, masking part of the word in the reference polypeptide sequence to obtain a masked polypeptide sequence;

[0007] Inputting the masked polypeptide sequence into a pre-trained polypeptide sequence modification model;

[0008] Determining predicted tokens of the masked portion of the masked polypeptide sequence based on the prediction results output by the polypeptide sequence modification model;

[0009] Replacing the masked word elements in the reference polypeptide sequence with the predicted word elements to obtain a predicted polypeptide sequence;

[0010] If the predicted polypeptide sequence is a polypeptide sequence that is different from any existing polypeptide sequence, the predicted polypeptide sequence is used as the modified polypeptide sequence.

[0011] In an optional embodiment, the word element of the reference polypeptide sequence is an amino acid residue word element;

[0012] and / or,

[0013] The masked word is a set non-amino acid residue word, or a designated amino acid residue word.

[0014] In an optional embodiment, the use of preset masking words to mask some words in the reference polypeptide sequence to obtain a masked polypeptide sequence includes:

[0015] Acquire a set word-unit covering parameter, wherein the word-unit covering parameter includes at least one of a word-unit covering ratio, a word-unit covering number, and a word-unit covering position;

[0016] According to the set word-metaphor masking parameters, some words in the reference polypeptide sequence are masked to obtain the masked polypeptide sequence.

[0017] In an optional embodiment, the reference polypeptide sequence is multiple, and the corresponding covering polypeptide sequence is multiple; the polypeptide sequence modification model is based on the multiple covering polypeptide sequences, and the predicted polypeptide sequence is multiple;

[0018] and / or,

[0019] The polypeptide sequence modification model is based on one masked polypeptide sequence, and the predicted polypeptide sequences obtained are multiple.

[0020] In an optional embodiment, if the predicted polypeptide sequence is a polypeptide sequence that is different from any existing polypeptide sequence, after using the predicted polypeptide sequence as the modified polypeptide sequence, the method further comprises:

[0021] Inputting the sequence information of the modified polypeptide sequence and the sequence information of the target receptor, as well as the structural image of the modified polypeptide sequence and the structural image of the target receptor into a first multimodal deep learning model to predict the activity information of the modified polypeptide sequence;

[0022] and / or,

[0023] The sequence information of the modified polypeptide sequence and the sequence information of the target receptor, as well as the structural image of the modified polypeptide sequence and the structural image of the target receptor are input into the second multimodal deep learning model to predict the toxicity information of the modified polypeptide sequence.

[0024] In an optional embodiment, the first multimodal deep learning model and the second multimodal deep learning model both include a Star-Transformer encoder, a convolutional neural network, a feature fusion layer, and a fully connected layer;

[0025] The Star-Transformer encoder is used to output first feature information based on the sequence information of the transformed polypeptide sequence and the sequence information of the target receptor;

[0026] The convolutional neural network is used to output second feature information based on the structural image of the modified polypeptide sequence and the structural image of the target receptor;

[0027] The feature fusion layer is used to fuse the first feature information and the second feature information to obtain fused feature information;

[0028] The fully connected layer of the first multimodal deep learning model predicts the activity information of the modified polypeptide sequence based on the fused feature information, and the fully connected layer of the second multimodal deep learning model predicts the toxicity information of the modified polypeptide sequence based on the fused feature information.

[0029] In an optional embodiment, the training process of the polypeptide sequence modification model is:

[0030] Obtaining sample peptide sequences;

[0031] Using the masking word to mask part of the word in the sample polypeptide sequence to obtain a masked sample polypeptide sequence;

[0032] The masked sample polypeptide sequence is used as the input of the polypeptide sequence transformation model to be trained, and the masked part of the word in the sample polypeptide sequence is used as the training label to train the polypeptide sequence transformation model.

[0033] In an optional embodiment, the word-meta coverage parameters of at least part of the sample polypeptide sequence are different from the word-meta coverage parameters of the reference polypeptide sequence; wherein the word-meta coverage parameters include: at least one of a word-meta coverage ratio, a word-meta coverage number, and a word-meta coverage position;

[0034] and / or,

[0035] The length of at least a portion of the sample polypeptide sequence is different from the length of the reference polypeptide sequence.

[0036] In a second aspect, the present invention provides an artificial intelligence-based polypeptide sequence modification device, comprising:

[0037] A reference polypeptide sequence acquisition module is used to obtain a reference polypeptide sequence;

[0038] A masking module, configured to mask some of the words in the reference polypeptide sequence using a preset masking word to obtain a masked polypeptide sequence;

[0039] A prediction module, configured to input the masked polypeptide sequence into a pre-trained polypeptide sequence modification model;

[0040] A predicted word unit determination module, configured to determine the predicted word unit of the masked portion of the masked polypeptide sequence based on the prediction result output by the polypeptide sequence modification model;

[0041] A predicted polypeptide sequence acquisition module is used to replace the masked word elements in the reference polypeptide sequence with the predicted word elements to obtain a predicted polypeptide sequence;

[0042] The modified polypeptide sequence determination module is used to use the predicted polypeptide sequence as the modified polypeptide sequence if the predicted polypeptide sequence is a polypeptide sequence different from any existing polypeptide sequence.

[0043] In a third aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to thereby execute the artificial intelligence-based polypeptide sequence modification method of the above-mentioned first aspect or any corresponding embodiment thereof.

[0044] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the artificial intelligence-based polypeptide sequence modification method of the above-mentioned first aspect or any corresponding embodiment thereof.

[0045] In a fifth aspect, the present invention provides a computer program product comprising computer instructions for causing a computer to execute the artificial intelligence-based polypeptide sequence modification method of the first aspect or any corresponding embodiment thereof.

[0046] The artificial intelligence-based polypeptide sequence modification methods, devices, equipment, and media provided in the embodiments of the present invention utilize polypeptide backbone modification based on an existing polypeptide sequence (i.e., a reference polypeptide sequence) to generate a new polypeptide sequence (i.e., a modified polypeptide sequence), explore the possibility of replacing different amino acid residues, and thus improve the quality and speed of polypeptide design. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in related technologies, the following briefly introduces the drawings required for use in the specific embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0048] Figure 1 1 is a schematic diagram of a process for modifying a polypeptide sequence based on artificial intelligence according to an embodiment of the present invention;

[0049] Figure 2 is a schematic diagram of a word unit table (including masked word units) according to an embodiment of the present invention;

[0050] Figure 3 is a schematic diagram of a polypeptide sequence modification model and a polypeptide sequence modification prediction process according to an embodiment of the present invention;

[0051] Figure 4is a schematic diagram of the architecture of the first multimodal deep learning model and the second multimodal deep learning model according to an embodiment of the present invention;

[0052] Figure 5 is a structural block diagram of a polypeptide sequence modification device based on artificial intelligence according to an embodiment of the present invention;

[0053] Figure 6 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0054] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0055] Artificial Intelligence (AI) refers to the intelligent behavior exhibited by computer systems. It is a discipline that uses interdisciplinary techniques such as computer science, mathematics, and neuroscience to create systems capable of performing tasks that typically require human intelligence.

[0056] Peptide design is an important field in biochemistry and molecular biology. It involves designing amino acid sequences based on specific functional requirements to form peptides with predetermined structures and functions. Peptides can be used as drugs, enzyme inhibitors, vaccine components, or for other biomedical applications. With the development of artificial intelligence (AI) technology, AI has played an increasingly important role in peptide design. The advantages of using AI for peptide design lie in its ability to process complex datasets, rapidly generate numerous hypotheses, and continuously refine design strategies through iterative learning.

[0057] According to an embodiment of the present invention, an embodiment of a polypeptide sequence modification method based on artificial intelligence is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of executable computer instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0058] This embodiment provides a method for modifying polypeptide sequences based on artificial intelligence, which can be used in various computer devices. Figure 1 Flowchart of the method for modifying a polypeptide sequence based on artificial intelligence according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0059] Step S101: Obtain a reference polypeptide sequence.

[0060] In the embodiments of the present invention, the reference polypeptide sequence may be a polypeptide sequence with high activity and low toxicity, so as to increase the possibility that the modified polypeptide sequence obtained by modification has high activity and low toxicity.

[0061] Specifically, the process of obtaining the reference polypeptide sequence can be:

[0062] Peptide datasets, including peptide sequences and corresponding activity-toxicity data, were collected and organized. Peptide sequences with high target binding activity and low toxicity were selected as reference sequences. For example, RSV (Respiratory Syncytial Virus) peptide data were collected from databases such as the DRAVP database (http: / / dravp.cpu-bioinfor.org / ) and the AVPdb database (http: / / crdd.osdd.net / servers / avpdb / ), as well as from other sources. Active and non-toxic peptide sequences were then screened based on a set threshold. An IC50 / EC50 of 30 μM was used as the toxicity threshold. A CC50 > 30 μM was assigned a value of 0, indicating non-toxicity, and a CC50 ≤ 30 μM was assigned a value of 1, indicating toxicity. Using an IC50 / EC50 ratio of 1 μM as the activity threshold, an IC50 / EC50 value greater than 1 μM was assigned a value of 0, indicating inactivity, and an IC50 / EC50 value ≤ 1 μM was assigned a value of 1, indicating activity. Finally, 119 active and non-toxic RSV peptide reference sequences were identified. All 119 sequences were derived from the RSV HR2 sequence (Reference: Sun, Z.; Pan, Y.; Jiang, S.; Lu, L. Respiratorysyncytial virus entry inhibitors targeting the F protein. Viruses 2013, 5, 211–225.). The RSV receptor sequence was obtained from the NCBI database (identifier FJ614815).

[0063] Step S102: Using preset masking words, masking some words in the reference polypeptide sequence to obtain a masked polypeptide sequence.

[0064] Specifically, the word element of the reference polypeptide sequence can be an amino acid residue word element. The mask word element can be a set non-amino acid residue word element or a specified amino acid residue word element. For example, each amino acid residue is regarded as a separate word element, and the mask word element is regarded as a special word element. Figure 2The word table shown, where <mask>To cover the word, <pad>To complete the word (used to complete the peptide sequence to the set maximum length of the model input), Y, A, T, V, L, D, E, G, R, H, I, W, Q, K, M, F, N, S, P, C are 20 amino acid residue words.

[0065] A word in the reference polypeptide sequence can be masked by a masking word, or multiple words in the reference polypeptide sequence can be masked by multiple masking words respectively. The multiple masked words in the reference polypeptide sequence can be non-adjacent, partially adjacent, or completely adjacent (i.e., multiple continuous words).

[0066] In an embodiment of the present invention, using a designated amino acid residue word as a mask word can improve the accuracy of predicting the masked word using a subsequent polypeptide sequence transformation model. When the masked polypeptide sequence is subsequently input into the pre-trained polypeptide sequence transformation model, the position information of the masked word in the masked polypeptide sequence is also input into the polypeptide sequence transformation model. In the case where multiple words are replaced with mask words in a masked polypeptide sequence, that is, in the case where there are multiple mask words in a masked polypeptide sequence, these mask words can be the same mask word or different mask words, for example, some are <mask>, some are specified amino acid residue words.

[0067] In some optional embodiments, step S102, i.e., using a preset masking word to mask some words in the reference polypeptide sequence to obtain a masked polypeptide sequence, includes:

[0068] Step S1021, obtain the set word-meta masking parameters, which include: at least one of the word-meta masking ratio, the number of word-meta masking, and the word-meta masking position. The word-meta masking ratio and the number of word-meta masking are generally selected from one of the two, that is, the word-meta masking parameter only includes one of them. Setting the word-meta masking ratio or the number of word-meta masking can control the number of masked word-meta. In addition, when the word-meta masking position parameter can uniquely determine the number of word-meta masking, the word-meta masking parameter can only include the word-meta masking position.

[0069] In an embodiment of the present invention, the word-blocking positions of the reference polypeptide sequence may be designated. In other embodiments, the word-blocking positions of the reference polypeptide sequence may be random.

[0070] Step S1022 : masking some of the words in the reference polypeptide sequence according to the set word masking parameters to obtain the masked polypeptide sequence.

[0071] For example, Figure 3 As shown, the HR2 sequence of RSV (i.e., NFYDPLVFPSDEFDASISQVNEKINQSLAFIRKSDELLHNVNAGKSTTN) is divided into multiple single amino acid residue tokens, wherein the token Y at the 3rd position and the token L at the 6th position are masked (the masked tokens are marked with <mask>indicates, ... indicates the omission of other tokens).

[0072] In addition, the set token masking parameter may be a range of token masking numbers, such as 1-5. Then, when masking some tokens in the reference polypeptide sequence according to the set token masking parameter, no more than 5 tokens may be randomly selected for masking. In other words, when masking tokens in the reference polypeptide sequence, the number of masked tokens is a random number between 1 and 5.

[0073] Step S103: input the masked polypeptide sequence into a pre-trained polypeptide sequence modification model.

[0074] Specifically, the masked polypeptide sequence does not need to be directly input into the polypeptide sequence modification model, but word segmentation is first performed, and the word units obtained by word segmentation are replaced with corresponding numerical values ​​to obtain a digitized masked polypeptide sequence, and then the digitized masked polypeptide sequence is input into the polypeptide sequence modification model.

[0075] The peptide sequence modification model can be a deep learning model. For example, the peptide sequence modification model can be a model built based on the Transformer (i.e., transformer) encoder architecture. The Transformer encoder architecture can be an architecture built from scratch. For example, Figure 3 As shown in Figure 1, the HR2 sequence of RSV with some words masked is input into the Transformer encoder to predict the masked words.

[0076] Step S104: determining the predicted word element of the masked portion in the masked polypeptide sequence based on the prediction result output by the polypeptide sequence modification model.

[0077] When there is only one masked word in the masked polypeptide sequence, the output of the polypeptide sequence modification model is N probability values, where N is the total number of amino acid residue words (this total number does not include the number of masked words and other special words (such as filler words, etc.)). For example, Figure 2 The word list shown is all the words, and the total number of words other than the masked words and the completed words is 20. The N probability values ​​output by the polypeptide sequence modification model correspond to N words, and each probability value indicates the probability value of the masked word in the reference polypeptide sequence is the corresponding word. The word corresponding to the maximum probability value can be used as the predicted word of the polypeptide sequence modification model, that is, the word corresponding to the maximum probability value is the masked word predicted by the polypeptide sequence modification model. It is also possible not to use the word corresponding to the maximum probability value as the predicted word, but to use the word corresponding to other probability values ​​as the predicted word, for example, the word corresponding to the second largest probability value as the predicted word.

[0078] When there are multiple masked words in the masked polypeptide word sequence, the number of probability values ​​output by the polypeptide sequence transformation model is M×N, where M is the number of masked words in the masked polypeptide word sequence.

[0079] Step S105 : replacing the masked word elements in the reference polypeptide sequence with the predicted word elements to obtain a predicted polypeptide sequence.

[0080] In the embodiment of the present invention, the predicted word may be the same as or different from the masked original word of the reference polypeptide sequence.

[0081] In step S106, if the predicted polypeptide sequence is different from any existing polypeptide sequence, the predicted polypeptide sequence is used as the modified polypeptide sequence. Specifically, existing polypeptide sequences can be collected in advance, and then the predicted polypeptide sequence can be compared one by one with the collected existing polypeptide sequences to determine whether the predicted polypeptide sequence is identical to the existing polypeptide sequence. Predicted polypeptide sequences that are identical to existing polypeptide sequences are discarded to ensure the novelty of the modified polypeptide sequence.

[0082] If the predicted token predicted by the peptide sequence modification model is the original token at the masked position in the reference peptide sequence, then the resulting predicted peptide sequence is the existing peptide sequence. Only when the predicted token predicted by the peptide sequence modification model is not the original token at the masked position in the reference peptide sequence can a modified peptide sequence be obtained. Of course, if multiple tokens are masked in the reference peptide sequence, then even if multiple predicted tokens are identical to the corresponding original tokens, as long as one predicted token is different from the corresponding original token, the predicted peptide sequence will be different from the reference peptide sequence and may not be the existing peptide sequence.

[0083] In addition, if the word corresponding to the maximum probability value in the probability value output by the polypeptide sequence modification model is the same as the original word at the masked position in the reference polypeptide sequence, then the word corresponding to the second largest probability value can be used as the predicted word, and then the masked part of the word in the reference polypeptide sequence is replaced with the predicted word to obtain the predicted polypeptide sequence.

[0084] In some optional specific embodiments, there are multiple reference polypeptide sequences, and there are multiple corresponding covering polypeptide sequences; the polypeptide sequence modification model is based on multiple covering polypeptide sequences, and the predicted polypeptide sequences are multiple.

[0085] That is, multiple predicted polypeptide sequences can be obtained based on multiple reference polypeptide sequences. If the multiple predicted polypeptide sequences are all different from the existing polypeptide sequences, multiple modified polypeptide sequences can be obtained.

[0086] When there are multiple reference polypeptide sequences, first use masking tokens to mask one or more tokens in the reference polypeptide sequences to obtain multiple masked polypeptide sequences (one masked polypeptide sequence is obtained from one reference polypeptide sequence). Then, sequentially input the multiple masked polypeptide sequences into the polypeptide sequence modification model to predict the masked tokens. Finally, use the predicted tokens (i.e., predicted tokens) to replace the masking tokens in the corresponding masked polypeptide sequences, and multiple predicted polypeptide sequences can be obtained.

[0087] In addition, based on one of the masked polypeptide sequences, the polypeptide sequence modification model predicts multiple predicted polypeptide sequences. That is, the polypeptide sequence modification model can predict multiple predicted polypeptide sequences based on one of the masked polypeptide sequences.

[0088] For example, if the polypeptide sequence modification model outputs N probability values corresponding to all tokens, then the N probability values can be sorted in descending order, and then the first M (M < N) probability values are taken. The M tokens corresponding to these M probability values are all used as predicted tokens. Finally, use these M tokens to replace the masked tokens in the reference polypeptide sequence respectively to obtain M predicted polypeptide sequences. The predicted polypeptide sequences that are different from the existing polypeptide sequences among these M predicted polypeptide sequences are the modified polypeptide sequences. That is, multiple modified polypeptide sequences can be obtained based on one reference polypeptide sequence.

[0089] In the embodiments of the present invention, the relationship between the input of the polypeptide sequence modification model and the modified polypeptide sequence can be one-to-one, that is, one input sequence (i.e., the reference polypeptide sequence) can be modified to obtain one polypeptide sequence. The relationship between the input of the polypeptide sequence modification model and the modified polypeptide sequence can also be one-to-many, that is, one input sequence (i.e., the reference polypeptide sequence) can be modified to obtain multiple polypeptide sequences. The relationship between the input of the polypeptide sequence modification model and the modified polypeptide sequence can also be many-to-many, that is, multiple input sequences (i.e., the reference polypeptide sequences) can be modified to obtain multiple polypeptide sequences.

[0090] In some optional specific embodiments, after the step of using the predicted polypeptide sequence as the modified polypeptide sequence if the predicted polypeptide sequence is a polypeptide sequence that is different from all existing polypeptide sequences, the method further includes:

[0091] Inputting the sequence information of the modified polypeptide sequence, the sequence information of the target receptor, as well as the structural image of the modified polypeptide sequence and the structural image of the target receptor into the first multi-modal deep learning model to predict the activity information of the modified polypeptide sequence;

[0092] and / or,

[0093] The sequence information of the modified polypeptide sequence and the sequence information of the target receptor, as well as the structural image of the modified polypeptide sequence and the structural image of the target receptor are input into the second multimodal deep learning model to predict the toxicity information of the modified polypeptide sequence.

[0094] In some specific implementations, such as Figure 4 As shown, both the first multimodal deep learning model and the second multimodal deep learning model include a Star-Transformer encoder, a convolutional neural network, a feature fusion layer, and a fully connected layer;

[0095] The Star-Transformer encoder is used to output first feature information based on the sequence information of the transformed polypeptide sequence and the sequence information of the target receptor;

[0096] The convolutional neural network is used to output second feature information based on the structural image of the modified polypeptide sequence and the structural image of the target receptor;

[0097] The feature fusion layer is used to fuse the first feature information and the second feature information to obtain fused feature information;

[0098] The fully connected layer of the first multimodal deep learning model predicts the activity information of the modified polypeptide sequence based on the fused feature information, and the fully connected layer of the second multimodal deep learning model predicts the toxicity information of the modified polypeptide sequence based on the fused feature information.

[0099] In this embodiment of the present invention, the architecture of the first and second multimodal deep learning models can be identical, both including a Star-Transformer encoder and a convolutional neural network (CNN). The first multimodal deep learning model is used to predict the activity of the modified polypeptide sequence, while the second multimodal deep learning model is used to predict the toxicity of the modified polypeptide sequence.

[0100] Specifically, when predicting the activity information of the modified polypeptide sequence, the sequence information of the modified polypeptide sequence and the sequence information of the target receptor are input into the Star-Transformer encoder of the first multimodal deep learning model, and the structural image of the modified polypeptide sequence and the structural image of the target receptor are input into the convolutional neural network of the first multimodal deep learning model. The output of the Star-Transformer encoder and the output of the convolutional neural network are merged and input into the fully connected layer to predict the activity information of the modified polypeptide sequence.

[0101] When predicting the toxicity information of the modified polypeptide sequence, the sequence information of the modified polypeptide sequence and the sequence information of the target receptor are input into the Star-Transformer encoder of the second multimodal deep learning model, and the structural image of the modified polypeptide sequence and the structural image of the target receptor are input into the convolutional neural network of the second multimodal deep learning model. The output of the Star-Transformer encoder and the output of the convolutional neural network are merged and input into the fully connected layer to predict the toxicity information of the modified polypeptide sequence.

[0102] The structural image of the modified polypeptide sequence can be predicted using relevant technologies.

[0103] While the architectures of the first and second multimodal deep learning models can be identical, they need to be trained and stored separately because they predict different information. During training, multiple data augmentation methods can be used on the peptide and receptor structural image samples for the convolutional neural network to improve model generalization, including but not limited to image rotation, random cropping of receptor images, and image color changes.

[0104] In an embodiment of the present invention, after the reference polypeptide sequence is modified using the polypeptide sequence modification model to obtain a modified polypeptide sequence, the modified polypeptide sequence can also be screened for activity and / or toxicity using the first multimodal deep learning model and the second multimodal deep learning model to screen out modified polypeptide sequences with high activity and low toxicity.

[0105] In some optional embodiments, the training process of the polypeptide sequence modification model is:

[0106] Obtaining sample peptide sequences;

[0107] Using the masking word to mask part of the word in the sample polypeptide sequence to obtain a masked sample polypeptide sequence;

[0108] The masked sample polypeptide sequence is used as the input of the polypeptide sequence transformation model to be trained, and the masked part of the word in the sample polypeptide sequence is used as the training label to train the polypeptide sequence transformation model.

[0109] In embodiments of the present invention, a large number of sample polypeptide sequences are used to train a polypeptide sequence modification model. These multiple sample polypeptide sequences can be divided into a training set, a test set, and a validation set. The multiple sample polypeptide sequences may or may not include the aforementioned reference polypeptide sequence.

[0110] For example, all of the 119 active and non-toxic RSV polypeptide reference sequences described above can be used as sample polypeptide sequences, some can be used as sample polypeptide sequences, or none can be used as sample polypeptide sequences.

[0111] In an embodiment of the present invention, the masked sample polypeptide sequence is used as the input of the polypeptide sequence transformation model to be trained. Of course, the masked sample polypeptide sequence can also be divided into word units first, and then the word units obtained by the division are replaced with corresponding numerical values ​​to obtain a digitized masked sample polypeptide sequence as the input of the polypeptide sequence transformation model to be trained. Then, based on the output of the model, the word units of the masked part predicted by the model, that is, the predicted word units, are obtained. Finally, the original word units of the masked position in the sample polypeptide sequence are used as labels to calculate the loss of the model. The training process such as the loss calculation method and the adjustment process of the model parameters can be implemented by selecting appropriate related technologies as needed, and will not be described in detail here.

[0112] In some optional embodiments, the word-meta coverage parameters of at least some of the sample polypeptide sequences are different from the word-meta coverage parameters of the reference polypeptide sequence; wherein the word-meta coverage parameters include at least one of: word-meta coverage ratio, word-meta coverage number, and word-meta coverage position;

[0113] and / or,

[0114] The length of at least a portion of the sample polypeptide sequence is different from the length of the reference polypeptide sequence.

[0115] The polypeptide sequence modification model in the embodiment of the present invention has strong generalization capabilities. Therefore, the masking method (including quantity, position, etc.) of the reference polypeptide sequence input during the polypeptide sequence modification process can be different from the masking method of the training sample during model training, which facilitates the adjustment of the masking method of the reference polypeptide sequence during the polypeptide sequence modification process.

[0116] The artificial intelligence-based polypeptide sequence modification method provided in this embodiment utilizes polypeptide backbone modification based on an existing polypeptide sequence (i.e., a reference polypeptide sequence) to generate a new polypeptide sequence (i.e., a modified polypeptide sequence), thereby exploring the possibility of replacing different amino acid residues, thereby improving the quality and speed of polypeptide design.

[0117] The present invention also uses deep learning to screen the activity and toxicity of modified peptide sequences. This can help to create peptide sequences with high activity and low toxicity, further improving the quality and speed of peptide drug design.

[0118] In the embodiments of the present invention, existing high-activity, low-toxicity polypeptide sequences are modified and upgraded based on a deep learning model, which can increase polypeptide activity and reduce toxicity, thereby improving the quality of polypeptide design.

[0119] In the embodiments of the present invention, a reference polypeptide sequence can be modified to obtain many new polypeptide sequences by changing the position and number of masked words, replacing them with different predicted words, and other methods. In addition, various types of polypeptide sequences can be used as reference polypeptide sequences. Therefore, a large number of new polypeptide sequences can be modified to establish a polypeptide modification library. Subsequently, the desired new polypeptide sequences can be screened from the polypeptide modification library, greatly improving the efficiency of polypeptide drug design.

[0120] This embodiment also provides an artificial intelligence-based polypeptide sequence modification device, which is used to implement the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented using software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0121] This embodiment provides a polypeptide sequence modification device based on artificial intelligence, such as Figure 5 As shown, including:

[0122] Reference polypeptide sequence acquisition module 501, used to acquire a reference polypeptide sequence;

[0123] A masking module 502 is configured to mask some of the words in the reference polypeptide sequence using a preset masking word to obtain a masked polypeptide sequence;

[0124] Prediction module 503, for inputting the masked polypeptide sequence into a pre-trained polypeptide sequence modification model;

[0125] A predicted word unit determination module 504 is used to determine the predicted word unit of the masked portion of the masked polypeptide sequence based on the prediction result output by the polypeptide sequence modification model;

[0126] The predicted polypeptide sequence acquisition module 505 is configured to replace the masked word elements in the reference polypeptide sequence with the predicted word elements to obtain a predicted polypeptide sequence;

[0127] The remodeled polypeptide sequence determination module 506 is configured to use the predicted polypeptide sequence as a remodeled polypeptide sequence if the predicted polypeptide sequence is a polypeptide sequence that is different from any existing polypeptide sequence.

[0128] In some optional embodiments, the word element of the reference polypeptide sequence is an amino acid residue word element;

[0129] and / or,

[0130] The masked word is a set non-amino acid residue word, or a designated amino acid residue word.

[0131] In some optional embodiments, the covering module 502 includes:

[0132] a word-unit covering setting parameter acquiring unit, configured to acquire set word-unit covering parameters, wherein the word-unit covering parameters include at least one of a word-unit covering ratio, a word-unit covering number, and a word-unit covering position;

[0133] The word-meta masking unit is used to mask some words in the reference polypeptide sequence according to the set word-meta masking parameters to obtain the masked polypeptide sequence.

[0134] In some optional embodiments, the reference polypeptide sequence is multiple, and the corresponding covering polypeptide sequence is multiple; the polypeptide sequence modification model is based on the multiple covering polypeptide sequences, and the predicted polypeptide sequence is multiple;

[0135] and / or,

[0136] The polypeptide sequence modification model is based on one masked polypeptide sequence, and the predicted polypeptide sequences obtained are multiple.

[0137] In some optional embodiments, the artificial intelligence-based polypeptide sequence modification device further comprises:

[0138] an activity prediction module, configured to input the sequence information of the modified polypeptide sequence and the sequence information of the target receptor, as well as the structural image of the modified polypeptide sequence and the structural image of the target receptor into a first multimodal deep learning model to predict the activity information of the modified polypeptide sequence;

[0139] and / or,

[0140] The toxicity prediction module is used to input the sequence information of the modified polypeptide sequence and the sequence information of the target receptor, as well as the structural image of the modified polypeptide sequence and the structural image of the target receptor into the second multimodal deep learning model to predict the toxicity information of the modified polypeptide sequence.

[0141] In some optional embodiments, the first multimodal deep learning model and the second multimodal deep learning model both include a Star-Transformer encoder, a convolutional neural network, a feature fusion layer, and a fully connected layer;

[0142] The Star-Transformer encoder is used to output first feature information based on the sequence information of the transformed polypeptide sequence and the sequence information of the target receptor;

[0143] The convolutional neural network is used to output second feature information based on the structural image of the modified polypeptide sequence and the structural image of the target receptor;

[0144] The feature fusion layer is used to fuse the first feature information and the second feature information to obtain fused feature information;

[0145] The fully connected layer of the first multimodal deep learning model predicts the activity information of the modified polypeptide sequence based on the fused feature information, and the fully connected layer of the second multimodal deep learning model predicts the toxicity information of the modified polypeptide sequence based on the fused feature information.

[0146] In some optional embodiments, the above-mentioned artificial intelligence-based polypeptide sequence modification device further includes:

[0147] A sample acquisition module is used to obtain sample polypeptide sequences;

[0148] A sample masking module, configured to mask some of the word elements in the sample polypeptide sequence using the masking word element to obtain a masked sample polypeptide sequence;

[0149] The training module is used to take the masked sample polypeptide sequence as the input of the polypeptide sequence transformation model to be trained, and to train the polypeptide sequence transformation model using the masked part of the word in the sample polypeptide sequence as the training label.

[0150] In some optional embodiments, the word-meta coverage parameters of at least some of the sample polypeptide sequences are different from the word-meta coverage parameters of the reference polypeptide sequence; wherein the word-meta coverage parameters include: at least one of a word-meta coverage ratio, a word-meta coverage number, and a word-meta coverage position;

[0151] and / or,

[0152] The length of at least a portion of the sample polypeptide sequence is different from the length of the reference polypeptide sequence.

[0153] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0154] The artificial intelligence-based polypeptide sequence modification device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0155] The embodiment of the present invention also provides a computer device having the above Figure 5 The artificial intelligence-based polypeptide sequence modification device shown.

[0156] See also Figure 6 , Figure 6 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.

[0157] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0158] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0159] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0160] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0161] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 6 The bus connection is taken as an example.

[0162] The input device 30 can receive input digital or character information and generate key signal input related to user settings and function control of the computer device. Examples include a touch screen, keypad, mouse, trackpad, touchpad, pointing stick, one or more mouse buttons, trackball, joystick, etc. The output device 40 may include a display device, auxiliary lighting devices (e.g., LEDs), and tactile feedback devices (e.g., vibration motors). Such display devices include, but are not limited to, liquid crystal displays, light emitting diodes, monitors, and plasma displays. In some optional embodiments, the display device may be a touch screen.

[0163] The computer device further includes a communication interface for the computer device to communicate with other devices or a communication network.

[0164] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0165] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0166] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.< / mask> < / mask> < / pad> < / mask>

Claims

1. A method for polypeptide sequence modification based on artificial intelligence, characterized in that: The method comprises: Obtain reference peptide sequences; Using a preset masking word, masking part of the word in the reference polypeptide sequence to obtain a masked polypeptide sequence; the masking word is a specified amino acid residue word; Inputting the masked polypeptide sequence into a pre-trained polypeptide sequence modification model; Based on the prediction results output by the polypeptide sequence modification model, a predicted word element of the masked portion in the masked polypeptide sequence is determined; if the polypeptide sequence modification model outputs N probability values ​​corresponding to all word elements, the N probability values ​​are arranged in descending order, and then the first M probability values ​​are taken, and the M word elements corresponding to these M probability values ​​are all used as predicted word elements; wherein the probability value corresponding to the word element indicates the probability value of the masked word element in the reference polypeptide sequence being the corresponding word element; Replacing the masked word elements in the reference polypeptide sequence with the predicted word elements to obtain a predicted polypeptide sequence; If the predicted polypeptide sequence is a polypeptide sequence that is different from any existing polypeptide sequence, the predicted polypeptide sequence is used as the modified polypeptide sequence; If the predicted polypeptide sequence is a polypeptide sequence that is different from any existing polypeptide sequence, then after using the predicted polypeptide sequence as the modified polypeptide sequence, the method further includes: screening the modified polypeptide sequence for activity and / or toxicity; specifically, the method includes: Inputting the sequence information of the modified polypeptide sequence and the sequence information of the target receptor, as well as the structural image of the modified polypeptide sequence and the structural image of the target receptor into a first multimodal deep learning model to predict the activity information of the modified polypeptide sequence; and / or, Inputting the sequence information of the modified polypeptide sequence and the sequence information of the target receptor, as well as the structural image of the modified polypeptide sequence and the structural image of the target receptor into a second multimodal deep learning model to predict toxicity information of the modified polypeptide sequence; Wherein, the first multimodal deep learning model and the second multimodal deep learning model both include a Star-Transformer encoder, a convolutional neural network, a feature fusion layer, and a fully connected layer; The Star-Transformer encoder is used to output first feature information based on the sequence information of the transformed polypeptide sequence and the sequence information of the target receptor; The convolutional neural network is used to output second feature information based on the structural image of the modified polypeptide sequence and the structural image of the target receptor; The feature fusion layer is used to fuse the first feature information and the second feature information to obtain fused feature information; The fully connected layer of the first multimodal deep learning model predicts the activity information of the modified polypeptide sequence based on the fused feature information, and the fully connected layer of the second multimodal deep learning model predicts the toxicity information of the modified polypeptide sequence based on the fused feature information; During the training process, multiple data augmentation methods are used for the structural image samples of peptides and receptors of the convolutional neural network, including but not limited to image rotation, random cropping of receptor images, and changing image color.

2. The method according to claim 1, characterized in that The terms of the reference polypeptide sequence are amino acid residue terms.

3. The method according to claim 1, characterized in that Using the preset masking token to mask some tokens in the reference polypeptide sequence to obtain a masked polypeptide sequence, including: Obtaining the set token masking parameters, where the token masking parameters include at least one of: token masking ratio, number of masked tokens, and token masking position; Masking some tokens in the reference polypeptide sequence according to the set token masking parameters to obtain the masked polypeptide sequence.

4. The method according to claim 1, wherein There are multiple reference polypeptide sequences, and the corresponding masked polypeptide sequences are multiple; the multiple predicted polypeptide sequences are predicted by the polypeptide sequence transformation model based on the multiple masked polypeptide sequences; and / or The polypeptide sequence transformation model predicts multiple predicted polypeptide sequences based on one masked polypeptide sequence.

5. The method according to claim 1, wherein The training process of the polypeptide sequence transformation model is as follows: Obtaining a sample polypeptide sequence; Using the masking token to mask some tokens in the sample polypeptide sequence to obtain a masked sample polypeptide sequence; Taking the masked sample polypeptide sequence as the input of the polypeptide sequence transformation model to be trained, and using the masked part of the tokens in the sample polypeptide sequence as the training label to train the polypeptide sequence transformation model.

6. The method according to claim 5, characterized in that At least some of the token masking parameters of the sample polypeptide sequence are different from those of the reference polypeptide sequence; where the token masking parameters include at least one of: token masking ratio, number of masked tokens, and token masking position; and / or At least some of the lengths of the sample polypeptide sequences are different from the lengths of the reference polypeptide sequences.

7. A polypeptide sequence modification device based on artificial intelligence, characterized in that: The device includes: A reference polypeptide sequence acquisition module for acquiring a reference polypeptide sequence; A masking module for using a preset masking token to mask some tokens in the reference polypeptide sequence to obtain a masked polypeptide sequence; the masking token is a set non-amino acid residue token; A prediction module for inputting the masked polypeptide sequence into a pre-trained polypeptide sequence transformation model; A predicted token determination module for determining the predicted tokens of the masked part in the masked polypeptide sequence based on the prediction result output by the polypeptide sequence transformation model; if the polypeptide sequence transformation model outputs N probability values corresponding to all tokens, then arrange the N probability values in descending order, and then take the first M (M < N) probability values, and regard the M tokens corresponding to these M probability values as predicted tokens; where the probability value corresponding to a token indicates the probability that the masked token in the reference polypeptide sequence is the corresponding token; A predicted polypeptide sequence acquisition module for replacing the masked part of the tokens in the reference polypeptide sequence with the predicted tokens to obtain a predicted polypeptide sequence; A transformed polypeptide sequence determination module for, if the predicted polypeptide sequence is a polypeptide sequence different from all existing polypeptide sequences, regarding the predicted polypeptide sequence as the transformed polypeptide sequence; The artificial intelligence-based polypeptide sequence transformation device further includes: a module for screening the activity and / or toxicity of the transformed polypeptide sequence; specifically including: an activity prediction module, configured to input the sequence information of the modified polypeptide sequence and the sequence information of the target receptor, as well as the structural image of the modified polypeptide sequence and the structural image of the target receptor into a first multimodal deep learning model to predict the activity information of the modified polypeptide sequence; and / or, a toxicity prediction module, configured to input the sequence information of the modified polypeptide sequence and the sequence information of the target receptor, as well as the structural image of the modified polypeptide sequence and the structural image of the target receptor into a second multimodal deep learning model to predict toxicity information of the modified polypeptide sequence; Wherein, the first multimodal deep learning model and the second multimodal deep learning model both include a Star-Transformer encoder, a convolutional neural network, a feature fusion layer, and a fully connected layer; The Star-Transformer encoder is used to output first feature information based on the sequence information of the transformed polypeptide sequence and the sequence information of the target receptor; The convolutional neural network is used to output second feature information based on the structural image of the modified polypeptide sequence and the structural image of the target receptor; The feature fusion layer is used to fuse the first feature information and the second feature information to obtain fused feature information; The fully connected layer of the first multimodal deep learning model predicts the activity information of the modified polypeptide sequence based on the fused feature information, and the fully connected layer of the second multimodal deep learning model predicts the toxicity information of the modified polypeptide sequence based on the fused feature information; During the training process, multiple data augmentation methods are used for the structural image samples of peptides and receptors of the convolutional neural network, including but not limited to image rotation, random cropping of receptor images, and changing image color.

8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the artificial intelligence-based polypeptide sequence modification method according to any one of claims 1 to 6 by executing the computer instructions.

Citation Information

Patent Citations

  • System for polypeptide design based on protein large language model

    CN118658514A

  • Polypeptide generation method and device and computer equipment

    CN118982999A