Multi-modal protein sequence generation method based on deep learning

Through a multimodal protein sequence generation method based on deep learning, combined with the multimodal interactive attention mechanism, the protein sequence and text description information are integrated, and the problem of difficulty in dealing with multimodal data in the existing technology is solved, and the high-precision and diversity of protein sequence generation is achieved, which improves the efficiency of protein design.

CN119943134APending Publication Date: 2025-05-06SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202411964342.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing protein language models have limitations when processing multimodal data, and it is difficult to fully reflect the complex characteristics of proteins and their diverse biological background, resulting in low accuracy and design efficiency of protein sequence generation.

Method used

A multimodal protein sequence generation method based on deep learning is used to integrate multimodal information such as protein sequences and text descriptions, and a protein generation model of multimodal interaction attention mechanism is used to generate high-quality protein sequences.

Benefits of technology

It significantly improves the accuracy and diversity of protein sequence generation, can capture the complex characteristics and biological background of proteins more comprehensively, and improves the efficiency and application value of protein design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943134A_ABST
    Figure CN119943134A_ABST
Patent Text Reader

Abstract

A multi-modal protein sequence generation method based on deep learning comprises the following steps: constructing an initial data set by fusing a protein sequence and multi-modal information described by a text, and training a protein sequence generation model based on a multi-modal interactive attention mechanism; in the online stage, according to text description input in real time and corresponding sequence data, a protein sequence conforming to expected functions and structures is generated through the trained model by means of an autoregression generation method. According to the method, the protein sequence and the text description and other multi-modal information are integrated as input, the protein generation model based on the multi-modal interaction attention mechanism is adopted, the protein sequence and the text data can be effectively fused, and therefore the high-quality protein sequence is generated; the method shows significant advantages in generation of distant homologous proteins and processing of complex tasks, and the improvement of the protein quality and structural similarity is verified through experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of bioengineering, specifically a multimodal protein sequence generation method based on deep learning. Background Art

[0002] Existing protein language models have limitations when processing multimodal data. They usually rely on only a single modality or lack the ability to effectively integrate multimodal data (such as protein sequence and text information), making it difficult to fully reflect the complex characteristics of proteins and their diverse biological backgrounds. Therefore, new evaluation methods are urgently needed to more accurately integrate multimodal data, thereby improving the accuracy of protein sequence generation and design efficiency. Summary of the invention

[0003] The present invention aims at the following deficiencies: the prior art cannot effectively integrate multimodal data such as protein sequences and text information, resulting in the inability to fully capture the multi-level characteristics of proteins; lacks fine-grained analysis capabilities and cannot deeply mine the complex associations and potential information in the data; when generating protein sequences based on multimodal data, the accuracy is low, affecting the efficiency and application effect of protein design; at the same time, the prior art is difficult to fully reflect the diverse biological background of proteins, resulting in the inability of the generated protein sequences to adapt to complex biological needs and many other deficiencies. A multimodal protein sequence generation method based on deep learning is proposed. By integrating multimodal information such as protein sequences and text descriptions as input, a protein generation model based on a multimodal interactive attention mechanism is adopted, which can effectively fuse protein sequences and text data to generate high-quality protein sequences. The present invention shows significant advantages in generating distantly homologous proteins and processing complex tasks, and its improvement in protein quality and structural similarity is verified by experiments.

[0004] The present invention is achieved through the following technical solutions:

[0005] The present invention relates to a multimodal protein sequence generation method based on deep learning. An initial data set is constructed by fusing multimodal information of protein sequence and text description, which is used to train a protein sequence generation model based on a multimodal interactive attention mechanism. In the online stage, a protein sequence that meets the expected function and structure is generated by the trained model using an autoregressive generation method according to the real-time input text description and the corresponding sequence data.

[0006] The initial data set is obtained in the following way: protein sequences with functional annotations, subcellular localization and protein family information are selected from the Swiss-Prot database, and redundancy removal is performed to finally screen out 469,395 protein text-sequence data pairs. The data set includes three parts of text description: protein functional annotation (FUNCTION), subcellular localization (SUBCELLULAR LOCATION) and protein family information (SIMILARITY), and the data pairs required for the model are constructed through these text information. The data set is then randomly divided into a training set, a validation set and a test set, where the training set contains 402,395 data pairs, the validation set contains 47,000 data pairs, and the test set contains 20,000 data pairs.

[0007] The protein sequence generation model based on the multimodal interactive attention mechanism includes: an encoder, a decoder, a linear mapping layer of each modal information and a protein generation module based on Top-p, temperature coefficient and repeated penalty comprehensive strategy, wherein: the encoder generates corresponding tensors and cross-modal tensors according to the protein sequence and protein description text; the decoder obtains the corresponding output tensor after layer normalization, rotation encoding and feedforward processing of multiple tensors by the multimodal interactive attention mechanism; the linear mapping layer of each modal data converts the output tensor of the corresponding modality into the generation conditional probability of the subsequent generation module, and the above-mentioned protein generation module uses the left-to-right autoregressive generation method to gradually construct the target protein sequence according to the generation conditional probability.

[0008] The cross-modal tensor is obtained by encoding the protein sequence and the protein description text respectively to obtain the sequence vector E s_input and the text vector E t_input After that, the two are connected through the cross-modal tensor E c_input To merge.

[0009] The encoder adopts but is not limited to the ESM1b model and the PubMedBERT model, which are used to encode protein sequences and texts respectively.

[0010] The decoder comprises 12 decoding layers connected in series, each decoding layer comprises: a layer normalization unit, a rotation position encoding unit, a residual connection unit, a multimodal interactive attention unit and a feedforward network, wherein: the layer normalization unit performs standardization processing according to the input multimodal data features to obtain a normalized data result; the rotation position encoding unit performs rotation position encoding processing according to the position information of the input sequence to obtain an embedding result with position information; the residual connection unit performs an addition operation according to the output of the previous layer and the current input information to obtain an output result containing residual information; the multimodal interactive attention unit performs attention mechanism processing according to the protein sequence, text description and cross-modal data to obtain an output result after multimodal interaction; the feedforward network performs nonlinear transformation processing according to the output information of the previous layer to obtain the final network output result.

[0011] The generated conditional probability Where: x t is the protein amino acid sequence generated by the model at time t, Text is the text information in the generation process, and p represents the probability information.

[0012] The autoregressive generation method refers to: using temperature coefficient T, repetition penalty R and Top-p decoding strategy, generating a protein sequence that meets the requirements through a text description of the protein or generating a sequence that meets the requirements through input text and protein sequence fragments. Specifically, the temperature coefficient controls the sampling diversity during the generation process, a higher temperature value increases the randomness of the generated sequence, and a lower temperature value makes the output more deterministic; the repetition penalty is used to reduce the recurrence of the same amino acid sequence during the generation process, thereby improving the diversity and quality of the generated sequence; the Top-p decoding strategy is based on probability distribution truncation, and selects words with a cumulative probability greater than p for sampling, thereby avoiding low-probability redundant outputs and ensuring that the generated sequence is more accurate and in line with expectations. Through the comprehensive application of these decoding strategies, it is possible to balance accuracy and diversity during the generation process to meet specific protein sequence design requirements.

[0013] The present invention relates to a protein sequence generation system for implementing the above method, comprising: a data processing unit for constructing and dividing a protein text-sequence data set, an encoding unit for encoding a protein sequence and a description text, a decoder unit for generating a protein sequence through cross-modal fusion, and a generation unit for generating a protein sequence using an autoregressive generation method and multiple decoding strategies. Technical Effects

[0014] The present invention efficiently fuses protein sequences, description texts and other biological information through a deep learning model, overcoming the limitations of the prior art in multimodal data processing. The method can capture the complex associations between different modalities in a fine-grained manner, and optimize the protein sequence generation process by comprehensively applying Top-p, temperature coefficient and repeated penalty decoding strategies, improve the accuracy and diversity of generation, and thus better meet the needs of protein design. Compared with the prior art, the present invention can more comprehensively capture the complex characteristics and biological background of proteins, and significantly improve the accuracy and diversity of protein sequence generation. Specifically, after fusing protein sequences, description texts and other biological information, the generated protein sequences are more biologically reasonable and can more accurately meet specific design requirements. In addition, the comprehensive application of Top-p, temperature coefficient and repeated penalty decoding strategies effectively reduces repeatability and improves the innovation of sequences, thereby greatly improving the efficiency and application value of protein design and promoting technological progress in the field of biomedicine. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flow chart of the present invention;

[0016] Figure 2 Preprocessing diagram for multimodal data;

[0017] Figure 3 Flowchart for model training;

[0018] Figure 4 This is a flowchart of the multimodal interactive attention mechanism;

[0019] Figure 5 It is the model loss curve;

[0020] Figure 6 Generate a module flow chart for the model;

[0021] Figure 7 Generate protein sequence metrics evaluation plots for the model. DETAILED DESCRIPTION

[0022] like Figure 1 As shown, this embodiment relates to a multimodal protein sequence generation method based on deep learning, which specifically includes:

[0023] In the first step, protein sequences with functional annotations, subcellular localization, and protein family information were selected from the Swiss-Prot database. These sequences were de-redundant to ensure the independence and accuracy of the data. After screening, 469,395 protein text-sequence data pairs were finally obtained. Each data pair contains three parts of text description information: protein functional annotation (FUNCTION), subcellular localization (SUBCELLULAR LOCATION), and protein family information (SIMILARITY), as shown in Table 1.

[0024] Table 1 Example of data set table

[0025] The second step is, Figure 2 As shown in the figure, the protein data in the model has two modes: text and sequence. In order to better integrate them, this paper introduces the cross-modal tensor E, which is a "bridge" for the integration between text and sequence. s_input , E t_input and E c_input Represent protein sequence vector, protein description text vector and cross-modal tensor respectively, which together constitute the input of the model. The model encodes protein sequence and protein description text through ESM1b and PubMedBERT models respectively and fuses the information of the two through cross-modal tensor. The introduction of cross-modal tensor promotes the interaction between protein sequence and text description.

[0026] The third step is Figure 3 and Figure 4As shown in the figure, the model decoder framework contains 12 decoding layers, each of which involves multiple important technical steps, such as layer normalization, rotation position embedding, residual connection, multimodal interactive attention mechanism and feedforward network. First, the embedding vectors of protein sequence and text are processed by the decoding layer to output vectors, and these output vectors are converted into probability distribution of protein sequence through linear mapping and softmax normalization function. The model adopts a new cross-stitching attention mechanism, which is significantly innovative compared with the traditional cross-stitching attention mechanism. The traditional cross-stitching attention mechanism usually only uses the information of one modality as the query vector, and the information of the other modality is processed as the key vector and value vector. The cross-stitching attention mechanism processes the information of both modalities as the key vector and the value vector at the same time, which can better capture the deep relationship between protein sequence and text information. On this basis, the present invention combines the self-attention mechanism of protein sequence and the cross-attention mechanism between protein sequence and text, so that the model can find the correspondence between amino acid groups and text rich in natural language information while learning the arrangement pattern of amino acid sequence. This improvement provides innovative ideas for the design of new proteins and a new methodology for text-guided protein generation. Through this unique cross-stitching attention mechanism, the present invention has achieved innovative breakthroughs in multimodal data processing, protein design and generation, and significantly improved the performance and application value of the model.

[0027] Under the specific environment settings of Python 3.9 in Linux environment, the training and validation loss function curves of the model in the training set and validation set of the data set are as follows Figure 5 As shown. The training and validation curves of the model show good convergence, and the convergence speed is from fast to slow, which is in line with the downward trend of the training generation model. At the same time, due to the large length of the protein, the loss function begins to converge at around 1.9, but it has learned the pattern of multimodal data guiding sequence generation very well. The loss curves of the validation set and the training set are highly consistent, indicating that the model performs well on both the training data and the validation data. There is no obvious overfitting or underfitting problem during the model training process, indicating that it has good generalization ability and stability.

[0028] The fourth step is Figure 6As shown in Figure 2, during the generation process, the model can effectively balance the randomness and accuracy of the generated sequence by adjusting the temperature coefficient, repetition penalty and Top-p decoding strategy, thereby improving the diversity and quality of the generated results. The model provides two generation modes: one is to generate a protein sequence that meets the requirements only through the protein text description; the other is that based on the text description, the user can provide an additional protein sequence fragment as a prompt to help generate a protein sequence that better meets specific needs. These generation modes enable the model to not only generate new protein sequences based on the description, but also to perform customized generation based on existing sequences, reflecting its flexibility and powerful capabilities in protein design and sequence generation. Examples of the two generation modes are shown in Table 2.

[0029] Table 2 Example model generation mode table

[0030] Based on the model, 40,000 protein sequences generated under the two modes were selected, and the evaluation of various indicators was as follows: Figure 7 As shown in Figure 2. The generated protein sequences show high global structural similarity and excellent local structural credibility. Although the global sequence identity is low, its structural quality is still high. The generated proteins have small deviations from the reference structure and high structural accuracy. The model is able to generate distantly homologous, highly innovative and high-quality protein sequences, showing strong adaptability and protein design capabilities.

[0031] Compared with the prior art, the present invention realizes fine-grained training of protein text-sequence datasets through deep learning algorithms, and provides an integrated model for the first time, which can be pre-trained from scratch on any specific protein text-sequence dataset. This model breaks through the problem of insufficient guidance of previous protein description texts, can generate new protein sequences, and effectively integrate protein sequences and text data, thereby improving the operability and accuracy of protein design. The present invention also proposes a new multimodal interactive attention mechanism, which realizes the deep fusion of protein sequence and text modalities, and promotes the innovative development of multimodal protein design. This model has significant improvements in generation quality, structural similarity and functionality compared with existing methods, especially in the generation of distant homologous proteins, showing its unique advantages.

[0032] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principle and purpose of the present invention. The protection scope of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. Each implementation scheme within its scope shall be subject to the constraints of the present invention.

Claims

1. A multimodal protein sequence generation method based on deep learning, characterized in that: By fusing the multimodal information of protein sequences and text descriptions, an initial dataset is constructed for training a protein sequence generation model based on a multimodal interactive attention mechanism. In the online stage, the trained model is used to generate protein sequences that meet the expected functions and structures using an autoregressive generation method based on the real-time input text descriptions and corresponding sequence data.

2. The method for generating multimodal protein sequences based on deep learning according to claim 1, characterized in that: The initial data set is obtained by: selecting protein sequences with functional annotations, subcellular localization and protein family information from the Swiss-Prot database, performing redundancy removal operations, and finally screening out protein text-sequence data pairs; The data set includes: protein function annotation (FUNCTION), subcellular localization (SUBCELLULARLOCATION) and protein family information (SIMILARITY), and the data pairs required for the model are constructed through these text information.

3. The method for generating multimodal protein sequences based on deep learning according to claim 1, characterized in that: The protein sequence generation model based on the multimodal interactive attention mechanism includes: an encoder, a decoder, a linear mapping layer of each modal information and a protein generation module based on Top-p, temperature coefficient and repeated penalty comprehensive strategy, wherein: the encoder generates corresponding tensors and cross-modal tensors according to the protein sequence and protein description text; the decoder obtains the corresponding output tensor after layer normalization, rotation encoding and feedforward processing of multiple tensors by the multimodal interactive attention mechanism; the linear mapping layer of each modal data converts the output tensor of the corresponding modality into the generation conditional probability of the subsequent generation module, and the above-mentioned protein generation module uses the left-to-right autoregressive generation method to gradually construct the target protein sequence according to the generation conditional probability.

4. The method for generating multimodal protein sequences based on deep learning according to claim 1 or 3, characterized in that: The cross-modal tensor is obtained by encoding the protein sequence and the protein description text respectively to obtain the sequence vector E s_input and the text vector E t_input After that, the two are connected through the cross-modal tensor E c_input To merge.

5. The method for generating multimodal protein sequences based on deep learning according to claim 3, characterized in that: The decoder comprises 12 decoding layers connected in series, each decoding layer comprises: a layer normalization unit, a rotation position encoding unit, a residual connection unit, a multimodal interactive attention unit and a feedforward network, wherein: the layer normalization unit performs standardization processing according to the input multimodal data features to obtain a normalized data result; the rotation position encoding unit performs rotation position encoding processing according to the position information of the input sequence to obtain an embedding result with position information; the residual connection unit performs an addition operation according to the output of the previous layer and the current input information to obtain an output result containing residual information; the multimodal interactive attention unit performs attention mechanism processing according to the protein sequence, text description and cross-modal data to obtain an output result after multimodal interaction; the feedforward network performs nonlinear transformation processing according to the output information of the previous layer to obtain the final network output result.

6. The method for generating multimodal protein sequences based on deep learning according to claim 3, characterized in that: The generated conditional probability Where: x t is the protein amino acid sequence generated by the model at time t, Text is the text information in the generation process, and p represents the probability information.

7. The method for generating multimodal protein sequences based on deep learning according to claim 1, characterized in that: The autoregressive generation method refers to: using a temperature coefficient T, a repetition penalty R and a Top-p decoding strategy, generating a protein sequence that meets the requirements through a text description of the protein or generating a sequence that meets the requirements through input text and protein sequence fragments. Specifically, the temperature coefficient controls the sampling diversity during the generation process, a higher temperature value increases the randomness of the generated sequence, and a lower temperature value makes the output more deterministic; the repetition penalty is used to reduce the occurrence of the same amino acid sequence during the generation process, thereby improving the diversity and quality of the generated sequence; the Top-p decoding strategy is based on probability distribution truncation, and selects words with a cumulative probability greater than p for sampling.

8. A protein sequence generation system for implementing the method according to any one of claims 1 to 7, characterized in that: include: A data processing unit for constructing and partitioning a protein text-sequence dataset, an encoding unit for encoding protein sequences and description texts, a decoder unit for generating protein sequences by cross-modal fusion, and a generation unit for generating protein sequences by using an autoregressive generation method and multiple decoding strategies.

Citation Information

Patent Citations

  • Multi-modal instruction guided protein design method and device

    CN118969059A

  • Cross-modal model training method and device, electronic equipment and storage medium

    CN119091971A

Cited By

  • Medical 3D image feature fusion method, device and system, and storage medium

    CN120689711A

  • Codon sequence design method and device based on large multi-modal model

    CN120727097A

  • Codon sequence design method and apparatus based on large multi-modal model

    CN120727097B