Protein design method and system based on selection state space model

By conducting protein sequence design and mutation prediction based on the Mamba model based on the selected state space model, the problems of scarcity of data, complexity of sequence space and uncertainty in the prior art are solved, and efficient and accurate protein sequence generation and mutation prediction are achieved.

CN119943131APending Publication Date: 2025-05-06ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510416145.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art faces challenges such as data scarcity and deviation, sequence space complexity, and sequence-function mapping uncertainty in protein sequence design and mutation prediction, resulting in insufficient generation accuracy and mutation prediction accuracy.

Method used

The protein design method based on the selective state space model (Mamba model) is adopted to transform the protein sequence generation task into the next marker prediction task. By training the Mamba model to learn the biological laws of protein sequences, capture the local and global dependencies between amino acids, and generate complete protein sequences or predict the impact of mutations.

Benefits of technology

Effectively capture long-term dependencies in protein sequences, improve generation accuracy, reduce dependence on large amounts of labeled data, improve the generalization ability of the model, and achieve efficient and accurate protein sequence generation and mutation prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943131A_ABST
    Figure CN119943131A_ABST
Patent Text Reader

Abstract

The invention discloses a protein design method and system based on a selection state space model, and belongs to the field of deep learning and computational biology. The method comprises the following steps: firstly, obtaining a protein sequence sample through random sampling, forming an embedded representation of the protein sequence sample according to a predefined vocabulary, and obtaining a training data set by constructing a training sample pair; regarding a protein sequence generation task as a next mark prediction task, and training a Mama model by using the training data set; and finally, executing a protein sequence generation task or a protein sequence mutation prediction task based on the trained Mamba model. According to the method, a traditional protein sequence design task is converted into a sequence generation task performed through the Mama model, the long-term dependency relationship in the protein sequence can be effectively captured, and the generation precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning and computational biology, and specifically relates to a protein design method and system based on a selection state space model. Background Art

[0002] Protein is the core substance of life activities, and its function is closely related to its three-dimensional structure. The function of protein is mainly determined by its amino acid sequence, and mutations in protein sequences often directly affect its three-dimensional structure and functional performance. Therefore, the design of protein sequences and mutation prediction have become key tasks in biological research and drug development. Traditional protein design methods usually rely on physical and chemical principles or existing structural information, and use computational methods to optimize protein sequences and screen mutations. However, traditional structure-based protein design methods have certain limitations, which are mainly manifested in high-dimensional and complex problems. In addition, these methods usually require a large amount of experimental data and computing resources, which significantly restricts their application in large-scale protein sequence design and mutation prediction tasks.

[0003] In recent years, deep learning technology has made significant progress in the field of protein sequence analysis and design. By training deep neural networks, researchers can mine potential patterns from a large number of protein sequences, and then perform tasks such as sequence function prediction and sequence-function mapping. For example, protein language models based on models such as convolutional neural networks (CNN) and recurrent neural networks (RNN) have been successfully applied to protein sequence classification, functional annotation and other fields. Although these methods can learn certain features from protein sequence data, they still face some challenges, including:

[0004] Data scarcity and bias: Although a large amount of protein sequence data is available, sequence samples with specific functions or mutations are still relatively scarce. This makes existing models susceptible to data scarcity and sample bias, which in turn affects their accuracy and generalization ability in specific design tasks.

[0005] Complexity of sequence space: The spatial dimensions involved in protein sequence design are extremely large. The length and composition changes of amino acid sequences can generate a diverse sequence space. How to effectively search and optimize sequences in this high-dimensional space to achieve the expected function or mutation prediction remains a technical challenge. Current methods often have difficulty in achieving precise control and efficient exploration in this complex space.

[0006] Uncertainty in sequence-function mapping: The relationship between protein sequence and its function is extremely complex, and many functional changes depend only on small changes or mutations in the sequence. Existing models are usually unable to accurately capture this fine-grained mapping when dealing with the subtle relationship between protein mutations and functional changes, especially in the case of multiple mutations. How to predict the comprehensive impact of these mutations on protein function is still a major problem in protein design.

[0007] In order to meet the above challenges, protein language models based on self-supervised learning (such as Tranception, ESM series models, etc.) have emerged in recent years. By pre-training on a large amount of protein sequence data, these models can learn the underlying rules in protein sequences and apply them to tasks such as protein function prediction and mutation effect analysis. Although these models have made significant progress in sequence generation and sequence-function relationship prediction, they still face certain limitations in terms of efficiency in exploring complex sequence spaces and accuracy in mutation prediction. Summary of the invention

[0008] The purpose of the present invention is to solve the above problems in the prior art and to provide a protein design method and system based on a selection state space model.

[0009] The specific technical solutions adopted by the present invention are as follows:

[0010] In a first aspect, the present invention provides a protein design method based on a selection state space model for generating a protein sequence, comprising:

[0011] S1. Obtain protein sequence samples by random sampling in a protein sequence database. According to a predefined vocabulary, map each amino acid symbol and sequence tag symbol in the protein sequence sample to their corresponding unique identifiers to form an embedded representation of the protein sequence sample. Obtain a training data set by constructing training sample pairs.

[0012] S2. Considering the protein sequence generation task as the next tag prediction task, the Mamba model is trained using the training data set, so that the Mamba model learns the biological laws of protein sequences and can capture the local and global dependencies between amino acids in the protein sequence; the trained Mamba model is used as a protein sequence generation network;

[0013] S3. According to the length of the protein sequence to be generated, an input sequence that meets the sequence length requirement is constructed through a padding operation based on the known amino acid part in the protein sequence. The input sequence mapping is converted into an embedded representation according to a predefined vocabulary and input into the protein sequence generation network. The amino acid corresponding to each padding mark symbol is predicted to generate a complete protein sequence.

[0014] As a preferred embodiment of the first aspect, during the random sampling process, the length of a single protein sequence sample cannot exceed 1024 amino acids. If the length of the sampled protein sequence exceeds 1024 amino acids, a starting site is randomly selected to intercept 1024 amino acids as a protein sequence sample.

[0015] As a preferred embodiment of the first aspect, when the Mamba model is trained using the training data set, the Mamba model generates each site in the protein sequence respectively, and the loss function adopts the cross entropy loss between the generated protein sequence and the original input protein sequence.

[0016] In a second aspect, the present invention provides a protein design method based on a selection state space model, which is used to predict mutations in a protein sequence, comprising:

[0017] S1. Obtain protein sequence samples by random sampling in a protein sequence database. According to a predefined vocabulary, map each amino acid symbol and sequence tag symbol in the protein sequence sample to their corresponding unique identifiers to form an embedded representation of the protein sequence sample. Obtain a training data set by constructing training sample pairs.

[0018] S2. Considering the protein sequence generation task as the next tag prediction task, the Mamba model is trained using the training data set, so that the Mamba model learns the biological laws of protein sequences and can capture the local and global dependencies between amino acids in the protein sequence; the trained Mamba model is used as a protein sequence generation network;

[0019] S3. For the protein sequence to be predicted to be mutated, its mapping is converted into an embedded representation according to a predefined vocabulary and input into the protein sequence generation network. The protein generation task is performed for each mutation site in the protein sequence to obtain the probability distribution of the amino acid categories that may be generated at each amino acid site. The probability value of each amino acid category in the probability distribution is used to characterize the possibility that the protein sequence mutates to this amino acid at the mutation site.

[0020] As a preferred embodiment of the second aspect, during the random sampling process, the length of a single protein sequence sample cannot exceed 1024 amino acids. If the length of the sampled protein sequence exceeds 1024 amino acids, a starting site is randomly selected to intercept 1024 amino acids as a protein sequence sample.

[0021] As a preferred embodiment of the second aspect, when the Mamba model is trained using the training data set, the Mamba model generates each site in the protein sequence respectively, and the loss function adopts the cross entropy loss between the generated protein sequence and the original input protein sequence.

[0022] In a third aspect, the present invention provides a protein design system based on a selection state space model, comprising:

[0023] A design mode selection module, used for allowing a user to select a design mode to be executed, wherein the design mode includes a protein sequence generation mode and a protein sequence mutation prediction mode;

[0024] A protein design module, used to read the design mode currently selected by the user to execute in the design mode selection module. If the user chooses to execute the protein sequence generation mode, the protein design method based on the selection state space model as described in any one of the schemes of the first aspect is used to call the trained protein sequence generation network to perform the protein sequence generation task; if the user chooses to execute the protein sequence mutation prediction mode, the protein design method based on the selection state space model as described in any one of the schemes of the second aspect is used to call the trained protein sequence generation network to perform the protein sequence mutation prediction task.

[0025] In a fourth aspect, the present invention provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, can implement the protein design method based on the selection state space model as described in any one of the first and second aspects above.

[0026] In a fifth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the protein design method based on the selection state space model as described in any one of the first and second aspects above can be implemented.

[0027] In a sixth aspect, the present invention provides a computer electronic device comprising a memory and a processor;

[0028] The memory is used to store computer programs;

[0029] The processor is used to implement the protein design method based on the selection state space model as described in any one of the first and second aspects above when executing the computer program.

[0030] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0031] The present invention designs a protein sequence generation method framework based on the Mamba model, which transforms the traditional protein sequence design task into a sequence generation task performed through the Mamba model. This design can effectively capture the long-term dependencies in the protein sequence and improve the generation accuracy. At the same time, the present invention can use large protein sequence data sets for pre-training, which not only reduces the dependence on a large amount of annotated data, but also improves the generalization ability of the model by optimizing the training process, and finally achieves efficient and accurate protein sequence generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 A schematic diagram of the steps of the protein design method based on the selection state space model;

[0033] Figure 2 Schematic diagram of the processing of the Mamba model in the training phase and the prediction phase;

[0034] Figure 3 This is an exemplary generation result when the Mamba model is used to generate an amino acid sequence;

[0035] Figure 4 A schematic diagram of the module composition of the protein design system based on the selection state space model;

[0036] Figure 5 It is a schematic diagram of the structure of computer electronic equipment. DETAILED DESCRIPTION

[0037] In order to make the above-mentioned purpose, features and advantages of the present invention more obvious and easy to understand, the specific implementation mode of the present invention is described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined accordingly without conflicting with each other.

[0038] There is usually a complex nonlinear and high-dimensional mapping relationship between protein sequences and their functions, which brings significant challenges to protein design and mutation prediction tasks. The present invention designs a protein sequence design and mutation prediction method based on the selection state space model (Mamba), which converts the traditional protein sequence design task into a sequence generation task performed through the Mamba model, so that the long-term dependencies in the protein sequence can be effectively captured and the generation accuracy can be improved. The core idea of ​​the Mamba model is to dynamically optimize the sequence generation and mutation prediction process by selectively exploring different states in the protein sequence space and combining deep learning technology. In the Mamba model, the state space represents all possibilities of protein sequences. By accurately selecting different states, optimization and adjustment can be effectively performed in the protein sequence design and mutation prediction tasks. Unlike traditional structure-based protein design methods, the Mamba model does not rely on complex structural information, but focuses on the sequence data itself, deeply explores the potential laws in the sequence, and accurately predicts the impact of specific mutations on protein function, thereby realizing the design of new sequences.

[0039] The protein design method based on the selection state space model of the present invention can be used to generate protein sequences or to predict protein sequence mutations. The basic concepts of the two are the same. Both require the use of deep learning technology to train the Mamba model to selectively explore the protein sequence space, extract potential laws and patterns from a large amount of protein sequence data, and learn the biological laws of protein sequences, so as to efficiently explore the protein sequence space, optimize the sequence design and mutation prediction process, and accurately evaluate the impact of mutations on protein function. The specific implementation of the present invention is described in detail below.

[0040] like Figure 1 As shown, in a preferred embodiment of the present invention, a protein design method based on a selection state space model is provided, which includes the following steps S1 to S3.

[0041] S1. Obtain protein sequence samples by random sampling in the protein sequence database. According to the predefined vocabulary, map each amino acid symbol and sequence tag symbol in the protein sequence sample to their corresponding unique identifiers to form an embedded representation of the protein sequence sample. Obtain a training data set by constructing training sample pairs.

[0042] It should be noted that the above protein sequence database is a database containing a large number of protein sequences, but it does not limit the specific database. In an embodiment of the present invention, the protein sequence database can use the UniRef50 data set, which contains protein sequences from different species. In this random sampling process, in order to cope with the graphics card memory limitation, it is necessary to trim the overly long protein sequences. The length of a single protein sequence sample cannot exceed 1024 amino acids. If the length of the sampled protein sequence exceeds 1024 amino acids, the starting position is randomly selected to intercept 1024 amino acids as a protein sequence sample. Specifically, if the length of the sampled protein sequence is greater than 1024 amino acids, a starting position (denoted as position a) is randomly selected, and 1024 amino acids are extracted from position a to form a subsequence with a length of 1024; if the length of the sampled protein sequence does not exceed 1024, the entire sequence can be directly used as a sample. This random truncation strategy not only avoids memory overflow caused by overly long sequences, but also ensures that each sequence can be fully utilized during the training process.

[0043] In addition, the sampled protein sequence samples are represented by amino acid symbols and some special sequence tags. The sequence tags include <cls>(indicates the start of a sequence), <eos>(marking the end of a sequence) and <pad>(indicates fill marker). All protein sequences will have a fill marker added at the beginning and end. <cls>and <eos>At the same time, since the length of the sequence obtained by random sampling is uncertain, in order to ensure that the sequence length of the input model is consistent, a padding strategy can be adopted to fill a series of <pad>Symbol, adjust sequences of different lengths to the same fixed length.

[0044] It should be noted that the original protein sequence sample cannot be directly used as the model input, but needs to be combined with a predefined vocabulary to perform certain vectorization operations. In an embodiment of the present invention, all possible amino acid symbols (such as A, M, etc.) need to be mapped to the corresponding token ID in the predefined vocabulary through an encoder. Token ID is a unique identifier for each amino acid symbol or special sequence tag in the vocabulary, which is used to convert the amino acid symbol or special tag into a digital form for model processing and calculation. The specific form of the Token ID in the vocabulary can be adjusted according to actual conditions. For example, all amino acid symbols and special sequence tags can be numbered in ascending order starting from 1 as the Token ID. The vocabulary not only contains the standard 20 amino acid symbols, but also includes the above-mentioned special sequence tags, such as <cls>(indicates the start of the sequence), <eos>(indicating the end of the sequence) and <pad>(used for padding). Therefore, in the present invention, after a series of protein sequence samples are obtained, it is necessary to map each amino acid symbol and sequence tag symbol in the protein sequence sample to their respective corresponding Token IDs according to a predefined vocabulary, and the Token IDs of all symbols in the protein sequence sample sequentially constitute the embedded representation of the protein sequence sample.

[0045] In addition, according to the general training principles of the model, each protein sequence needs to be converted into a data format suitable for model input, that is, to construct a training data pair consisting of a sample input and a sample label pair. As mentioned above, the embedded representation of the protein sequence sample is used as the sample input for model training, and the sample label needs to be processed according to the training requirements of the Mamba model. In the present invention, the protein sequence generation task needs to be regarded as the next token prediction task, so the sample label also needs to be constructed in the same way as the next token prediction task. In this task, if a sequence is generally represented as ABCD, then its sample input is <cls>ABCD <eos>, the sample label is the amino acid sequence ABCD after removing the sequence start and end markers.

[0046] S2. The protein sequence generation task is regarded as the next tag prediction task, and the Mamba model is trained using the training data set, so that the Mamba model learns the biological laws of protein sequences and can capture the local and global dependencies between amino acids in the protein sequence; the trained Mamba model is used as the protein sequence generation network.

[0047] It should be noted that the Mamba model is a neural network architecture based on the state space model (SSM), which aims to solve the efficiency and long-term dependency problems of traditional Transformer and RNN when processing long sequence data, and combines hardware optimization design. The Mamba model itself belongs to the prior art, and the focus of the present invention is not to optimize the internal model structure and framework of the Mamba model. Therefore, the existing Mamba model can be directly called as the initial model for training to construct a neural network for protein sequence generation, which is called a protein sequence generation network in the present invention.

[0048] Mamba introduced in the present invention adopts a structured state space model, which can efficiently capture local and global dependencies in protein sequences. Since protein sequences have very strong biological dependencies and regularities, the traditional Transformer architecture, although it performs well in language modeling, often faces bottlenecks in computational and memory efficiency when processing protein sequences. By introducing a structured state space method, Mamba can reduce computational complexity while ensuring the expressiveness of the model. In the architectural design, the Mamba model utilizes hardware optimization strategies so that it can run efficiently on hardware platforms such as GPUs. When processing protein sequences, Mamba can not only capture long-range dependencies in the sequence, but also reduce redundant calculations while ensuring computational efficiency.

[0049] Although the Mamba model itself belongs to the existing technology, in order to facilitate understanding of the principles and advantages of the Mamba model in performing protein sequence generation tasks, the specific method of constructing a neural network structure for protein sequence generation based on the Mamba model is introduced below.

[0050] In the embodiment of the present invention, in the implementation of the Mamba model, the network structure includes the following key components:

[0051] 1. Mamba Block: Mamba Block is the basic building block of the Mamba model. Multiple Mamba Blocks are stacked together to form a complete Mamba network. The design goal of each Mamba Block is to model protein sequences through deep learning methods and extract their local and global features. The core functions of the Mamba Block include feature extraction, pattern recognition, and information propagation.

[0052] The calculation process of each Mamba block can be described by the following steps:

[0053] 1. Input: Assume that the input is the embedded representation of the protein sequence , the dimension is .

[0054] 2. Feature extraction: First, the Mamba block performs a convolution operation on the input to extract local features:

[0055]

[0056] in is the convolution kernel, is the bias term.

[0057] 3. Activation function: Next, use the nonlinear activation function to process the extracted features:

[0058]

[0059] 4. Information transfer: The processed feature information is transferred to the next layer through residual connection to enhance the expressiveness of the model. By stacking multiple Mamba blocks, the model can learn the complex patterns of protein sequences layer by layer.

[0060] 5. Selective State Space Model (S6):

[0061] The selective state space model (S6) is the core component of the Mamba model. It borrows the structure of the traditional state space model (SSM), but introduces a selection mechanism that enables the model to dynamically adjust its parameters according to the input data. The traditional SSM has the property of being time-invariant, that is, the state transfer matrix and observation matrix of the system are fixed. The selective state space model introduces a selectivity mechanism that can selectively adjust the state transfer matrix according to different input data. and the observation matrix , to adapt to the complex patterns in protein sequences. The specific implementation of the selective state space model (S6) is as follows:

[0062] State Space Model: Setting Initial States and the observation matrix , define the state transfer and observation equations:

[0063]

[0064] in, is the system status, is the state update function, is the observation matrix, and are process noise and observation noise, respectively. is the state of the input sequence at time t.

[0065] Selection mechanism: S6's selectivity mechanism can dynamically determine how to update the parameters at the current time step. The selection mechanism can be based on the characteristics of the input sequence, such as the local pattern or global structure of the amino acid sequence, and dynamically adjust the learned weights. and

[0066] Update rule: Assume that at time step t, the selection mechanism determines the weights of the state transfer matrix based on the input features ,but: =

[0067] in, and There are two possible state transition matrices, are coefficients computed dynamically from the input sequence.

[0068] 6. Output module of Mamba model:

[0069] The output module of the Mamba model is responsible for mapping the final state of the model to the prediction results of the protein sequence. To achieve this goal, the output module uses a fully connected layer to output the probability distribution or amino acid category of the protein sequence. Assume that the final hidden state of the model is The calculation steps of the output module are as follows:

[0070] Linear transformation: The hidden state is transformed through a fully connected layer Mapping to the output space:

[0071]

[0072] in, is the output weight matrix, is the bias term.

[0073] Softmax activation: To get the predicted probability of each amino acid, the output is normalized using the Softmax function:

[0074]

[0075] in, It is The output of amino acid class is the total number of amino acid classes.

[0076] Through this process, the output module of the Mamba model is able to generate prediction results for each time step, ultimately forming the prediction of the protein sequence.

[0077] It should be emphasized that the above description of the internal modules and data processing flow of the Mamba model is mainly to facilitate understanding of the principles of the Mamba model in the task of generating protein sequences, but it is not an improvement or optimization of the Mamba model. The specific Mamba model can still be implemented by referring to the existing technology or directly called.

[0078] After the Mamba model is constructed, the training data set consisting of a series of training sample pairs can be used to train the Mamba model. During training, the Mamba model can generate each site in the protein sequence separately, and the loss function uses the cross entropy loss between the generated protein sequence and the original input protein sequence.

[0079] In the embodiment of the present invention, the specific method of training the Mamba model is shown in step 1) and step 2).

[0080] Step 1) In this step, the Mamba model will be initially trained using a training dataset containing protein sequences and their corresponding target labels. The training dataset Uniref50 includes a large number of protein sequences, which is designed to reflect the diversity of protein sequences and their inherent regularities. By inputting these data into the Mamba model, the initial training goal of the model is to enable it to capture the structural and functional relationship of protein sequences and then learn the basic rules for generating protein sequences. At the beginning of training, the parameters of the model (such as weights and biases) are initialized to random values ​​or initial estimates based on prior knowledge. At this point, the model has not yet made accurate predictions about the patterns of the input data, and the quality of its output sequences is low. Therefore, the goal of the initial training is to gradually adjust the parameters of the model so that the model can effectively understand and generate outputs that conform to the universal laws of protein sequences.

[0081] Step 2), optimize the Mamba model parameters, adjust the weights by calculating the loss function, and gradually improve the generation accuracy. Specifically, during the model training process, the core task is to adjust the parameters of the Mamba model through the optimization algorithm so that the protein sequence it generates gradually approaches the target sequence. To this end, the cross entropy loss function is used to quantify the gap between the sequence generated by the model and the target sequence. The cross entropy loss function can effectively measure the similarity between the generated sequence and the true sequence, reflecting the prediction accuracy. Specifically, the cross entropy loss function compares each predicted site in the generated sequence with the true label of the corresponding site in the target sequence, and calculates the overall error:

[0082]

[0083]

[0084] In the formula is the amino acid category output for protein sequence site i The predicted probability on , N is the total number of protein sequence samples.

[0085] By minimizing this loss function , the Mamba model can continuously update its parameters to reduce the error between the generated sequence and the target sequence, thereby improving the accuracy and quality of sequence generation. The optimization process is usually carried out using a gradient descent algorithm, which calculates the gradient of the loss function with respect to the model parameters through a back-propagation mechanism, and uses this gradient information to gradually adjust the model parameters in the direction of minimizing the loss function.

[0086]

[0087] in and represent the model parameters before and after updating, respectively. represents the gradient of the loss function with respect to the model parameters, Represents the learning rate of training.

[0088] The above process of calculating loss and backpropagation will continue for multiple rounds of iterations to continuously improve the quality of the generated protein sequence. In the early stages of training, the model will focus on learning the basic patterns and rules in sequence generation, such as the basic arrangement of amino acids, common sequence structures, etc. In this process, the model may generate some sequences that do not conform to actual biological laws. However, by continuously optimizing parameters, the model can gradually capture more complex and detailed laws, such as interactions between amino acids, sequence stability and functionality. As the training progresses, the Mamba model will be able to more accurately simulate the protein sequence generation process and produce high-quality output sequences that meet given goals. Ultimately, through continuous optimization and adjustment, the Mamba model will be able to generate highly accurate protein sequences that conform to biological laws. After training, the Mamba model can be used as a protein sequence generation network for actual protein sequence generation or protein mutation prediction. The processing process of the Mamba model in the training and prediction stages is as follows: Figure 2 shown.

[0089] S3. Generate protein sequences based on the protein sequence generation network, or predict mutations of protein sequences.

[0090] Based on the trained protein sequence generation network, it can be used to generate protein sequences or predict mutations in protein sequences according to task requirements. The specific practices of the two modes are introduced below.

[0091] S31: The method for generating protein sequences is as follows:

[0092] According to the length of the protein sequence to be generated, an input sequence that meets the sequence length requirement is constructed through a padding operation based on the known amino acid part in the protein sequence. The input sequence mapping is converted into an embedded representation according to a predefined vocabulary and input into the protein sequence generation network. The amino acid corresponding to each padding mark symbol is predicted to generate a complete protein sequence.

[0093] It should be noted that in the task of generating protein sequences, the length of the protein sequence to be generated can be specified according to actual needs. Moreover, this task allows the amino acid types at some amino acid sites to be known. These known amino acids need to be constructed in the input sequence according to their respective sites, while other unknown sites are filled in through the filling operation. <pad>And like the protein sequence samples, for protein sequences consisting of known amino acids and filler markers, they should also be preceded and followed by <cls>and <eos>Indicate the start and end of the sequence respectively, thus obtaining a complete model input sequence. In addition, it should be noted that in this task, although the input sequence that meets the sequence length requirement can be constructed by filling in the known amino acid part of the protein sequence, it does not necessarily require the known amino acid part. In fact, there may not be any known amino acid. <cls>and <eos>All the fillings between <pad>.

[0094] Thus, after constructing the input sequence, the input sequence mapping can be converted into an embedded representation according to the predefined vocabulary. The specific method of this mapping conversion is consistent with the samples in the aforementioned training stage, and each amino acid symbol and sequence marker symbol in the input sequence can be mapped to their corresponding unique identifiers. On this basis, the trained Mamba model can be used to generate a predicted complete protein sequence by inputting part of the known protein sequence information. In this task, the input part of the information can be a part of the protein sequence (for example, the first half of the sequence) or an amino acid residue at a specific position. Based on this known information, the Mamba model will use the biological laws and sequence generation rules it has learned to predict the part of the sequence that has not yet been determined. Figure 3 The example sequence contains 19 sites in total. The amino acid category and probability generated at each site can be intuitively visualized.

[0095] In addition, the complete protein sequence generated by the Mamba model should comply with a series of biological rules and constraints. Specifically, the generated sequence must not only comply with the chemical properties of amino acids (such as hydrophilicity, hydrophobicity, acidity and alkalinity, etc.), but also follow the folding characteristics, stability and functional requirements of proteins. For example, the Mamba model will consider the interactions between amino acid residues when generating sequences to ensure that the generated sequence has a reasonable spatial conformation so that it can form an effective functional structure in the body.

[0096] This function is of great significance for the design of new drugs, protein engineering and synthetic biology. By providing a complete protein sequence based on partial sequence information, the Mamba model can not only accelerate the discovery of drug targets, but also provide an important theoretical basis for the design of proteins with specific functions.

[0097] S32: The method for predicting mutations in protein sequences is as follows:

[0098] For the protein sequence to be predicted to be mutated, its mapping is converted into an embedded representation according to a predefined vocabulary and input into the protein sequence generation network. The protein generation task is performed for each mutation site in the protein sequence to obtain the probability distribution of amino acid categories that may be generated at each amino acid site. The probability value of each amino acid category in the probability distribution is used to characterize the possibility that the protein sequence mutates to this amino acid at the mutation site.

[0099] It should be noted that the above protein sequence to be predicted for mutation is a known protein sequence. After the sequence is input into the Mamba model through mapping transformation, the Mamba model can perform the protein generation task for each mutation site in the protein sequence. Since the output of the Mamba model in the protein generation task is actually a K-dimensional amino acid category probability distribution, K is the number of all amino acid category labels. In the protein generation task of S31, the amino acid category with the highest probability can be selected from this probability distribution as the final amino acid generated at this site through the argmax operation. However, in the protein sequence mutation prediction task, this amino acid category probability distribution is equivalent to the mutation probability, and the probability value of each amino acid category in the probability distribution is used to represent the possibility of the protein sequence mutating to this amino acid at the mutation site. The larger the probability value, the greater the possibility of mutation.

[0100] It should be noted that the mutation site in the protein sequence can be specified according to actual needs. It can be a single site or a series of sites. In addition, all amino acid sites in the sequence can be regarded as mutation sites for overall mutation prediction. When the Mamba model is applied to protein mutation prediction, a known protein sequence can be input. The model can predict the sequence changes that may occur after mutation based on the amino acid relationships learned during the training process. This process can predict the impact of mutations (such as substitutions, deletions, or insertions) at one or more amino acid sites on the protein sequence and its structure.

[0101] The Mamba model also has broad application prospects in the field of drug development. The Mamba model can predict the impact of mutations on protein targets by working with structural prediction models such as Alphafold2. The model can assist in drug design, especially in the development of targeted drugs for mutation-related diseases (such as cancer, genetic diseases, etc.), providing an important reference for the design of drug molecules.

[0102] In order to demonstrate the technical effect of the present invention, in the embodiment of the present invention, two Mamba models with different parameter amounts are trained, which are respectively denoted as Mamba M and Mamba S. The number of stacked Mamba blocks in Mamba M is 48 layers, the dimension is 1024, and the parameter amount is 240M, while the number of stacked Mamba blocks in Mamba S is 24 layers, the dimension is 768, and the parameter amount is 85M. At the same time, in order to better demonstrate the technical effect of the present invention, a plurality of comparison models in the prior art are set on the same data set, namely Progen2 M, Progen2 Base, Tranception L, Tranception M, and Tranception S. After all Mamba models and comparison models are trained, the protein sequence generation network performs mutation prediction, and the final result comparison is shown in Table 1:

[0103] Table 1 Comparison of mutation prediction results of each model

[0104]

[0105] It can be seen that the accuracy of Mamba M with 240M parameters and Mamba S with 85M parameters can reach 0.384 and 0.317 respectively. The accuracy of Mamba M with 240M parameters has exceeded all other comparison models, but its parameter count has been greatly reduced relative to the comparison model with the closest accuracy. Therefore, the protein sequence generation network trained based on the Mamba model of the present invention can obtain a higher mutation prediction accuracy with a smaller parameter count.

[0106] It should be noted that although the above effect verification experiment was tested on the mutation prediction task, since the basic principles of the mutation prediction task and the protein sequence generation task are the same, the above results can also be extended to the protein sequence generation task.

[0107] In summary, traditional protein design methods rely on physical and chemical principles or known structural information, and have limitations such as low efficiency in high-dimensional sequence space search and insufficient mutation prediction accuracy. The present invention avoids dependence on complex structural information by introducing a selection state space model, and can search flexibly and efficiently in high-dimensional sequence space, and accurately predict the fine-grained effects of protein mutations on function. In addition, through the application of deep learning technology, the present invention effectively overcomes the problems of wide sample space and sample bias, and significantly improves the accuracy and efficiency of protein sequence optimization and mutation prediction. This method has broad application prospects in protein engineering, drug design and related biological research fields.

[0108] In addition, based on the same inventive concept, Figure 4 As shown, the present invention also provides a protein design system based on a selection state space model corresponding to the protein design method based on a selection state space model provided in the above embodiment, which includes:

[0109] A design mode selection module, used for allowing a user to select a design mode to be executed, wherein the design mode includes a protein sequence generation mode and a protein sequence mutation prediction mode;

[0110] The protein design module is used to read the design mode currently selected by the user in the design mode selection module. If the user chooses to execute the protein sequence generation mode, the protein sequence generation network obtained by training is called to execute the protein sequence generation task according to the protein design method based on the selection state space model as described in the above embodiment; if the user chooses to execute the protein sequence mutation prediction mode, the protein sequence generation network obtained by training is called to execute the protein sequence mutation prediction task according to the protein design method based on the selection state space model as described in the above embodiment.

[0111] The above protein design system can implement the design mode selection module through a GUI interface to facilitate user interaction, and the results of the protein design module can also be visualized through the GUI interface. The specific module design belongs to the existing technology and can be designed according to actual needs.

[0112] In addition, it should be noted that the protein design method based on the selection state space model in the above embodiments can essentially be implemented in the form of a computer program.

[0113] Therefore, based on the same inventive concept, Figure 5 As shown, the present invention also provides a computer electronic device corresponding to the protein design method based on the selection state space model provided in the above embodiment, which includes a memory and a processor;

[0114] The memory is used to store computer programs;

[0115] The processor is used to implement the protein design method based on the selection state space model as described above when executing the computer program;

[0116] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.

[0117] Therefore, based on the same inventive concept, the present invention provides a computer-readable storage medium corresponding to a protein design method based on a selection state space model, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it can implement the protein design method based on the selection state space model as described above.

[0118] Therefore, based on the same inventive concept, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement the protein design method based on the selection state space model as described above.

[0119] Specifically, in the computer-readable storage medium of the above three embodiments, the stored computer program is executed by the processor to perform the above steps S1 to S3.

[0120] It is understandable that the above storage medium may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. The storage medium may also be a U disk, a mobile hard disk, a magnetic disk or an optical disk, etc., which can store program codes.

[0121] It is understandable that the above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0122] It should also be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the various embodiments provided in this application, the division of steps or modules in the system and method is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules or steps can be combined or integrated together, and a module or step can also be split.

[0123] The above-described embodiments are only some preferred implementations of the present invention, but are not intended to limit the present invention. A person skilled in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.< / pad> < / eos> < / cls> < / eos> < / cls> < / pad> < / eos> < / cls> < / pad> < / eos> < / cls> < / pad> < / eos> < / cls> < / pad> < / eos> < / cls>

Claims

1. A protein design method based on a selection state space model for generating protein sequences, characterized in that: include: S1. Obtain protein sequence samples by random sampling in a protein sequence database. According to a predefined vocabulary, map each amino acid symbol and sequence tag symbol in the protein sequence sample to their corresponding unique identifiers to form an embedded representation of the protein sequence sample. Obtain a training data set by constructing training sample pairs. S2. Considering the protein sequence generation task as the next tag prediction task, the Mamba model is trained using the training data set, so that the Mamba model learns the biological laws of protein sequences and can capture the local and global dependencies between amino acids in the protein sequence; the trained Mamba model is used as a protein sequence generation network; S3. According to the length of the protein sequence to be generated, an input sequence that meets the sequence length requirement is constructed through a padding operation based on the known amino acid part in the protein sequence. The input sequence mapping is converted into an embedded representation according to a predefined vocabulary and input into the protein sequence generation network. The amino acid corresponding to each padding mark symbol is predicted to generate a complete protein sequence.

2. The protein design method based on the selection state space model according to claim 1, characterized in that: During the random sampling process, the length of a single protein sequence sample cannot exceed 1024 amino acids. If the length of the sampled protein sequence exceeds 1024 amino acids, a starting site is randomly selected to intercept 1024 amino acids as a protein sequence sample.

3. The protein design method based on the selection state space model according to claim 1, characterized in that: When the Mamba model is trained using the training data set, the Mamba model generates each site in the protein sequence respectively, and the loss function adopts the cross entropy loss between the generated protein sequence and the original input protein sequence.

4. A protein design method based on a selection state space model for predicting mutations in protein sequences, characterized in that: include: S1. Obtain protein sequence samples by random sampling in a protein sequence database. According to a predefined vocabulary, map each amino acid symbol and sequence tag symbol in the protein sequence sample to their corresponding unique identifiers to form an embedded representation of the protein sequence sample. Obtain a training data set by constructing training sample pairs. S2. Considering the protein sequence generation task as the next tag prediction task, the Mamba model is trained using the training data set, so that the Mamba model learns the biological laws of protein sequences and can capture the local and global dependencies between amino acids in the protein sequence; the trained Mamba model is used as a protein sequence generation network; S3. For the protein sequence to be predicted to be mutated, its mapping is converted into an embedded representation according to a predefined vocabulary and input into the protein sequence generation network. The protein generation task is performed for each mutation site in the protein sequence to obtain the probability distribution of the amino acid categories that may be generated at each amino acid site. The probability value of each amino acid category in the probability distribution is used to characterize the possibility that the protein sequence mutates to this amino acid at the mutation site.

5. The protein design method based on the selection state space model according to claim 4, characterized in that: During the random sampling process, the length of a single protein sequence sample cannot exceed 1024 amino acids. If the length of the sampled protein sequence exceeds 1024 amino acids, a starting site is randomly selected to intercept 1024 amino acids as a protein sequence sample.

6. The protein design method based on the selection state space model according to claim 4, characterized in that: When the Mamba model is trained using the training data set, the Mamba model generates each site in the protein sequence respectively, and the loss function adopts the cross entropy loss between the generated protein sequence and the original input protein sequence.

7. A protein design system based on a selection state space model, characterized in that: include: A design mode selection module, used for allowing a user to select a design mode to be executed, wherein the design mode includes a protein sequence generation mode and a protein sequence mutation prediction mode; A protein design module, used for reading the design mode currently selected by the user in the design mode selection module, and if the user selects to execute the protein sequence generation mode, calling the trained protein sequence generation network to execute the protein sequence generation task according to the protein design method based on the selection state space model as described in any one of claims 1 to 3; If the user chooses to execute the protein sequence mutation prediction mode, the protein sequence generation network obtained by training is called to execute the protein sequence mutation prediction task according to the protein design method based on the selection state space model as described in any one of claims 4 to 6.

8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, it can implement the protein design method based on the selection state space model as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the protein design method based on the selection state space model as described in any one of claims 1 to 6 is implemented.

10. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the protein design method based on the selection state space model as described in any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Method and device for training protein prediction model based on graph neural network

    CN116935952A

  • Protein reverse folding method and device based on multi-modal pre-training large model

    CN117727365A

  • Transitive and commutative multimodal models and uses

    WO2024223621A1