Protein sequence generation and representation learning

Through the diffusion protein language model (DPLM) based on the discrete diffusion probability model, the problem of insufficient generation and understanding ability in the protein sequence generation and prediction tasks is solved, and efficient protein sequence generation and feature representation is achieved, which improves the accuracy of downstream tasks.

WO2025175491A1PCT designated stage Publication Date: 2025-08-28BEIJING YOUZHUJU NETWORK TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/077839
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-20
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing language models are difficult to have both generation and understanding capabilities in protein sequence generation and prediction tasks, especially autoregressive language models cannot capture the complex global interactions of amino acids, resulting in poor generation and prediction results.

Method used

The diffusion protein language model (DPLM) based on the discrete diffusion probability model structure is used to achieve the generation and feature representation of protein sequences by pre-training on real protein sequence data, combining the discrete diffusion probability model and language model structure, and has two-way receptive field and generation ability.

Benefits of technology

Accurate generation and characteristic representation of protein sequences are achieved, the accuracy of downstream prediction tasks is improved, and protein sequences that meet specific conditions can be generated in the condition generation and controllable generation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024077839_28082025_PF_FP_ABST
    Figure CN2024077839_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure relate to protein sequence generation and representation learning. A method comprises: obtaining a target model, the target model being based on a discrete diffusion probabilistic model structure and a language model structure, and being pre-trained by using a set of real protein sequences; and using the target model to execute at least one of the following: at least on the basis of an input sequence in a protein sequence form, generating a first target protein sequence by using the target model; or, on the basis of a second target protein sequence, extracting a sequence feature representation corresponding to the second target protein sequence by using the target model.
Need to check novelty before this filing date? Find Prior Art

Description

Protein sequence generation and representation learning Technical Field

[0001] Example embodiments of the present disclosure relate generally to the field of computers, and more particularly to protein sequence generation and representation learning. Background Art

[0002] Proteins are three-dimensional (3D) folded linear sequences composed of amino acid residues. Proteins carry out many vital biological activities within organisms and play a crucial role in regulating various biological functions, including transcription, translation, signal transduction, and life cycle control. With the development of data-driven deep learning technologies, the field of protein analysis has evolved from traditional methods to the use of models to understand and design proteins. The goal is to develop powerful models to accomplish the diverse tasks of interest in protein research.

[0003] Summary of the Invention

[0004] In a first aspect of the present disclosure, a method for information processing is provided. The method comprises: obtaining a target model, the target model being based on a discrete diffusion probability model and a language model structure and pre-trained using a set of real protein sequences; and using the target model to perform at least one of the following: generating a first target protein sequence based on at least an input sequence in the form of a protein sequence using the target model; or extracting a sequence feature representation corresponding to a second target protein sequence based on the second target protein sequence using the target model.

[0005] In a second aspect of the present disclosure, an apparatus for information processing is provided. The apparatus includes: a model acquisition module configured to obtain a target model, the target model being based on a discrete diffusion probability model and a language model structure and pre-trained using a set of real protein sequences; and a model execution module configured to use the target model to perform at least one of the following: generating a first target protein sequence based on at least an input sequence in the form of a protein sequence using the target model; or extracting a sequence feature representation corresponding to a second target protein sequence based on the second target protein sequence using the target model.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method of the first aspect of the present disclosure.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium and can be executed by a processor to perform the method according to the first aspect of the present disclosure.

[0008] In a fifth aspect of the present disclosure, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method according to the first aspect of the present disclosure.

[0009] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0011] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIG2 shows a schematic diagram of a training architecture of a model for protein sequence generation and representation learning according to some embodiments of the present disclosure;

[0013] FIG3 shows an example of extracting sequence feature representations of protein sequences using a model according to some embodiments of the present disclosure;

[0014] FIG4A shows an example of a conditional generation task conditioned on a partial sequence according to some embodiments of the present disclosure;

[0015] FIG4B shows an example of a conditional generation task conditioned on a secondary structure according to some embodiments of the present disclosure;

[0016] FIG4C shows an example of a controllable sequence generation task with guidance information according to some embodiments of the present disclosure;

[0017] FIG5 shows a flowchart of a protein-related information processing process according to some embodiments of the present disclosure;

[0018] FIG6 shows a block diagram of a device for protein-related information processing according to some embodiments of the present disclosure; and

[0019] FIG7 illustrates a block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION

[0020] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0021] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below.

[0022] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.

[0023] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0024] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the information involved in this disclosure should be informed to relevant users and authorization should be obtained from relevant users in an appropriate manner in accordance with relevant laws and regulations. Relevant users can include any type of right holders, such as individuals, enterprises, and groups.

[0025] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly prompt the relevant user that the operation requested to be performed will require obtaining and using the information of the relevant user, so that the relevant user can independently choose whether to provide information to the software or hardware such as the electronic device, application, server or storage medium that executes the operation of the technical solution of the present disclosure based on the prompt message.

[0026] As an optional but non-limiting implementation, in response to receiving an active request from a relevant user, a prompt message may be sent to the relevant user in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide information to the electronic device.

[0027] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure. The activation of the digital assistant-related functions of the embodiment of the present disclosure, the data obtained, the processing and storage of the data, etc., shall all obtain the prior authorization of the user and other rights holders associated with the user, and shall comply with the provisions of relevant laws and regulations and the rules of agreement between rights holders.

[0028] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.

[0029] FIG1 shows a schematic diagram of an example environment 100 in which an embodiment of the present disclosure can be implemented. In the environment 100, a protein-related processing system 110 can utilize a target model 120 to perform protein-related analysis tasks based on a specified input 102. In protein-related analysis, protein sequence prediction, protein structure prediction, protein sequence representation learning, etc. are all tasks with practical application significance. In some implementations, the processing system 110 can generate a protein sequence based on the input 102 using the target model 120. The processing system 110 can generate a sequence representation corresponding to the protein sequence based on the input 102 (for example, the input protein sequence) using the target model 120. Based on the accurate sequence representation, various subsequent understanding tasks can be performed, including classification of protein sequences, classification of residues in proteins, protein sequence regression, protein contact point prediction, and the like.

[0030] In Figure 1, the processing system 110 can be any type of device with computing capabilities, including a terminal device or a server device. The terminal device can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. The server device can, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like.

[0031] It should be understood that the structure and functionality of environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0032] In the field of deep learning, language models (LMs) have made remarkable progress in natural language understanding. Given the similarities between protein sequences and human language, there has been discussion in the protein research community about using language models to achieve various tasks, including protein-related prediction tasks (e.g., detecting functional properties and predicting protein structure from a single sequence without the need for evolving homologous chromosomes) and protein-related generation tasks (e.g., attributing a given protein backbone structure, redesigning an optimized protein sequence, or synthesizing a completely new protein sequence).

[0033] The language model aims to learn the probability model p θ (x), to estimate the underlying distribution x~q(x) of the sequence data of interest (e.g., text or protein sequence). The language model can be parameterized by a neural network (e.g., a Transformer-based neural network) and represented as θ. When performing language modeling on protein sequences, the sequence data to be processed can be represented as x=(x1, x2, ..., x L )∈{0,1} L×|ν| , which consists of L elements. ν is a vocabulary of discrete data. For a protein sequence, ν includes multiple amino acids used to make up the protein, for example, it can be composed of 20 different amino acids, expressed as ν = {1, ..., 20}.

[0034] Current language models use different model structures for prediction and generation tasks. For prediction tasks (which require comprehension), the widely used pre-training target is mask prediction. For mask prediction, the masked language model (MLM) can detect the bidirectional receptive field of the sequence and gain comprehension. The bidirectional receptive field can consider both left-to-right and right-to-left directions, thereby taking into account the context on the left and right sides, and predicting the masked sequence (i.e., amino acids) through the self-encoding method of mask prediction. The pre-training target of mask prediction can be expressed as follows:

[0035] in It is a mask sequence that performs masking on the sequence x using a special mask symbol [X] (e.g., proportional masking, etc.). The masked sequence after masking is represented as This pre-training objective is based on the conditional independence assumption of each token. Mask prediction can achieve good model performance for sequence understanding tasks such as natural language and protein sequences.

[0036] For generative tasks, a widely used pre-training objective is autoregressive training. Autoregressive training uses the probabilistic chain rule to perform sequential factorization on the sequence. Model learning is then performed by maximizing the log-likelihood of the model on the training dataset, which can be expressed as follows:

[0037] Causal masking is used to ensure the sequence dependency structure. To sample the autoregressive model, it is necessary to follow a strict left-to-right unidirectional method from x1 to p θ (x1), x2~p θ (x2|x1) towards x L ~p(x L |x1,...,x L-1 ) performs L iterative steps.

[0038] By analyzing language models currently adopted in the field of natural language processing, the inventors have the following observations.

[0039] In mask prediction, the model lacks configuration for generative modeling, making it difficult to perform sequence generation tasks. Furthermore, the bidirectional nature of mask prediction makes it difficult to apply to sequence generation. Furthermore, proteins are structured macromolecules, not simple linear sequences. Therefore, autoregressive language models (ARLMs) in natural language processing only consider sequence context in a single direction. In protein-related tasks, the limitation of ARLMs is their inability to capture the complex global interactions of amino acids, resulting in unsatisfactory results in both protein generation and prediction.

[0040] Autoregressive language models can generate sequences, but they lack a good understanding of sequential data. In generative tasks, the model must continuously approximate the underlying data distribution to create new output samples. Therefore, it is desirable for generative models to simultaneously maintain a deep understanding of the data.

[0041] In the embodiments of the present disclosure, it is desirable to propose a unified universal model with both predictive and generative capabilities. In generative tasks, this universal model is a powerful, scalable generative model that can effectively capture protein sequences. In predictive tasks, this universal model has a bidirectional receptive field and can effectively model global interactions at the residue level.

[0042] In an embodiment of the present disclosure, a language model based on a discrete diffusion probability model structure is proposed, which is called a diffusion protein language model (DPLM). The model is pre-trained on a large amount of real protein sequence data, so that it can learn knowledge in protein sequences at the evolutionary level in the real world. The model has both the generation ability of a language model and the understanding ability of a diffusion model. The discrete diffusion probability model structure can have the generative generalization ability of language modeling. During pre-training, following the training principles of the diffusion model, the model is allowed to denoise the input protein sequence as training data at different noise levels, thereby obtaining a variety of capabilities for understanding and outputting proteins. After pre-training, the model can be used for protein sequence generation, as well as to provide accurate sequence feature representations of protein sequences for various downstream prediction tasks.

[0043] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0044] FIG2 shows a schematic diagram of a training architecture 200 of a model for protein sequence generation and representation learning according to some embodiments of the present disclosure.

[0045] In the architecture 200, a diffusion protein language model (DPLM) 210 (sometimes also referred to as a target model in this document) is constructed based on a discrete diffusion probability model structure and a language model structure. Specifically, the DPLM 210 may include multiple Transformer layers 212 used in the language model. Each Transformer layer 212 may include a multi-layer perceptron (MLP) 214 and a multi-head attention block 216. The Transformer layer also has other structural designs in addition to the structure shown in Figure 2. Based on the structure of the language model, the DPLM 210 also performs data processing through a subset of the diffusion probability model, namely, the discrete diffusion probability model.

[0046] As mentioned above, when language modeling is performed on protein sequences, the sequence data to be processed can be represented as x=(x1, x2, ..., xL )∈{0,1} L×|ν| , the sequence consists of L elements. ν is a vocabulary of discrete data. For a protein sequence, ν includes multiple amino acids used to compose the protein, for example, it can be composed of 20 different amino acids, represented as ν = {1, ..., 20}. Therefore, for DPLM 210, the element corresponding to each position in the input sequence and the output sequence (also called a token) can indicate any amino acid in the amino acid vocabulary, or can indicate a special symbol (for example, [X], which does not represent any amino acid).

[0047] For a better understanding, the following will first briefly introduce the diffusion probability model and then introduce the discrete diffusion probability model.

[0048] The diffusion probability model is a type of generative model, but its data generation process is based on a pair of Markov processes, namely the forward diffusion process and the backward denoising process. The forward diffusion process (expressed as: is the stepwise interference data x (0) ~q(x (0) ), through T stepwise noise addition steps x (1:T) =x1,...,x (t-1) , x (t) ,...,x (T) , and obtain the static noise distribution x (T) ~q noise Through model training, the learned backward denoising process (expressed as: Perform the reverse process and gradually denoise the samples towards the data distribution to obtain data x (0) ~q(x (0) ). It can be seen that the backward denoising process can correspond to the desired data modeling process and finally obtain the desired data.

[0049] In some implementations, to represent the model (denoted as: p θ (x (0) ))fits to the data distribution q(x (0) ), the learning of the backward denoising process is usually achieved by optimizing the variational constraint of the log-likelihood, which can be expressed as follows:

[0050] After learning is completed, the model performing the backward denoising process is able to first noise (x (T) ) starts sampling and by using p θ (x (t-1) |x (t) ) is iterated to denoise until the desired data is obtained.

[0051] In the diffusion probability model, a high-speed interference kernel is usually used for continuous diffusion, which has high performance when generating continuous data in Euclidean space and generalized Riemann manifolds. Considering that the bidirectional receptive field of the diffusion model is particularly suitable for global interactions at the residue level, the diffusion model is considered to be suitable for understanding protein sequences. On the other hand, the language model can achieve effective sequence learning that is diffusible and is suitable for protein sequence generation. Therefore, in the embodiments of the present disclosure, it is expected to obtain a model that has both the ability to understand (predict) protein sequences and the ability to generate sequences.

[0052] However, directly using a continuous diffusion model to model protein sequences does not yield good results. This is because protein sequences are discrete data, and Gaussian diffusion makes it difficult to model the discrete nature of sequence data in an embedded representation space, leading to discrete traps. Therefore, in the embodiments of the present disclosure, the DPLM 210 for protein sequence generation and representation learning proposes modeling protein sequences using a discrete diffusion probability model structure that operates directly in a discrete state space. Furthermore, the discrete diffusion probability model is implemented on the structure of a language model.

[0053] Specifically, assume that Cat(x;p) is the class distribution of the protein sequence x parameterized by the vector p on the (|ν|-1)-dimensional probability sample (this is a probability distribution used to describe the probability distribution of a random variable with a finite number of classes). The forward diffusion process of the discrete diffusion probability model defines a Markov process controlled by the transition kernel, which is expressed as follows: q(x (t) |x (t-1) )=Cat(x (t) βtx (t-1) +(1-β t )q noise ) (4)

[0054] where q noise is the static noise distribution q noise (x (t) ) probability vector (,4) that is, q(x (t) )=Cat(x (t) ; p = q noise , and 0<<β t <1 is the early noise control parameter, which controls the degree of noise interference at step t. Thus, given the original protein sequence x (0) , the disturbed sequence data x (t) has a closed form representation: q(x (t) x (0) )=Cat(x (t) ; α t x (0) +(1-αt )q noise ) (5)

[0055] in Make lim t→T α t →0, which means that at the Tth step, the information of the original protein sequence is not retained, but converges to the static noise distribution q noise This means that the entire diffusion process is a convex combination of the data (the original protein sequence) and the static noise prior distribution.

[0056] Different static noise distributions q noise In some embodiments, discrete diffusion is used with the following restrictions: if x (t) =[X],q(x (t) )=1; if x (t) ≠[X],q(x (t) )=0, where [X] represents the absorption state and does not indicate any specific amino acid in the protein sequence representation. The above formula (5) can make the sequence data x of step t (t) Able to pass the mask rate (1-α t ) is masked, or is identical to the original protein sequence x (0) same.

[0057] During the training of DPLM 210, discrete probability diffusion is related to the training of autoregressive language models and masked language models, but the learning objective of discrete probability diffusion can be further simplified. In some embodiments, the KL divergence between two classes can be transferred to the reweighted cross entropy through a reparameterized backward pass, which is expressed as:

[0058] where λ (t) is the weighting coefficient introduced by the specific noise interference. From the above formula (6), it can be seen that in the above formula (1) The mask language model in the case of , and in the above formula (2) and The autoregressive language model in the case of can be considered a special case of a generalized form of the discrete diffusion language model. Therefore, the learning process according to Equation (6) above includes the learning process of the masked language model and the autoregressive language model, which means that the learned model can simultaneously possess the understanding and prediction capabilities provided by the masked language model and the sequence generation capabilities provided by the autoregressive language model.

[0059] Based on the model structure discussed above, the DPLM 210 in the embodiment of the present disclosure is also pre-trained on a large-scale real protein sequence set. The real protein sequence set includes evolutionary level protein sequences 220, i.e., protein sequences that have been detected in the real world. During the pre-training process, the forward diffusion process is to extract the real protein sequences x from the pre-training dataset. (0) At the beginning, the noise is gradually added in T steps until a static noise distribution x is obtained. (T) , that is, all elements in the sequence are represented as [X]. The backward denoising process is to get the noise from the static noise distribution x (T) Start by gradually reducing the noise until you can generate the real protein sequence x in the pre-training dataset (0) .

[0060] By pre-training on a large number of real protein sequences, DPLM 210 is able to learn the characteristic representations of real protein sequences and learn how to generate protein sequences that conform to the characteristics of real protein sequences. Given a trained model, it can synthesize new protein sequences through the reverse iterative denoising process of discrete diffusion. Ultimately, discrete diffusion can sample from the following distribution:

[0061] Specifically, at step t, first start from p θ (·|x (t) )generate Then given x (t) and In the case of Can get x with less noise (t-1) . This process is repeated from T to 1.

[0062] The generative denoising process of DPLM 210 can be considered as an iterative mask prediction process. Specifically, the initial sequence can be a sequence with all noise states (i.e., all [X]). In each iteration, according to Model-based predictions of one or more elements in the initial sequence The elements are updated, that is, one or more elements are updated from the noise state to indicate a certain amino acid, while the remaining elements still maintain the noise state.

[0063] In some embodiments, DPLM 210 can support generating new protein sequences starting from a random input sequence in the form of a protein sequence. An input sequence in the form of a protein sequence can be input to DPLM 210, which then generates a first target protein sequence based on the target model. That is, throughout the entire backward denoising process, DPLM 210 can support denoising from arbitrary sequence data, ultimately generating a protein sequence. Therefore, for protein sequence generation tasks, DPLM 210 can be used to generate a target protein sequence based at least on an input sequence in the form of a protein sequence. In some embodiments, the random input sequence can be user-input or specified by other means. All elements in the random input sequence can indicate a noise state, or some elements can indicate a noise state while others indicate one or more amino acids.

[0064] In some embodiments, the generated protein sequence can be used to predict the corresponding protein structure 230. Specifically, various protein structure prediction methods or tools, including but not limited to the ESMfold tool, can be used to predict 3D protein structure from protein sequence.

[0065] In terms of supporting representation learning for prediction tasks, DPLM 210 is able to denoise the input protein sequence at various denoising levels, including the original noise-free protein sequence. Therefore, DPLM 210 can be used as a protein sequence representation learner for a large amount of protein sequence data at the same time. Given a target protein sequence, DPLM 210 can extract a sequence feature representation (also called a sequence embedding representation) of the target protein sequence. The sequence feature representation can be used for various downstream prediction tasks, including classification tasks or regression tasks at the sequence level or residue level. The extraction of sequence feature representation can be to let DPLM 210 take a given target protein sequence x as input for feature extraction, which is expressed as:

[0066] Where d is the dimension of the sequence feature representation, and h(x) is the extracted feature representation sequence.

[0067] FIG3 illustrates an example of extracting a sequence feature representation of a protein sequence using a model according to some embodiments of the present disclosure. Given a target protein sequence 310, it can be considered the initial state of DPLM 210, i.e., t = 0. Target protein sequence 310 is then input into DPLM 210, and a sequence feature representation 320 corresponding to target protein sequence 310 is obtained through processing according to Equation (8).

[0068] In some embodiments, the extracted sequence feature representation corresponding to the target protein sequence 310 can be used to perform sequence-level or residue-level classification or regression tasks related to the target protein sequence 310, including but not limited to residue classification tasks, sequence classification tasks, sequence regression tasks, contact point prediction tasks, etc. Because DPLM 210 can accurately understand and extract feature representations of protein sequences, this will significantly improve the accuracy of the results of subsequent downstream prediction tasks.

[0069] In some embodiments, DPLM 210 can also support conditional generation tasks. In some application scenarios, it may be desirable to generate protein sequences that meet certain conditions. Such conditional generation tasks include sequence-based conditional generation, extended-modal conditional generation, and plug-and-play preference-guided conditional generation.

[0070] FIG4A shows an example of a conditional generation task based on a partial sequence according to some embodiments of the present disclosure. In the conditional generation task based on a partial sequence, the input sequence of DPLM 210 includes a specified input sequence in the form of a protein sequence, and the element of at least one position in the specified input sequence indicates at least one specified amino acid. In actual application scenarios, the conditional generation task based on a partial sequence is suitable for generating a protein sequence with a given functional motif, filling an antibody complementary determining region (CDR) loop, or a protein sequence with expert knowledge as a priori condition. In these application scenarios, DPLM 210 needs to obtain the protein sequence from the conditional distribution. Sampling is performed, where Indicates that some amino acids have been specified for the input sequence. Each element in b i =0, then If b i =1|i∈[1,L], That is to say, b i ∈{0,1} indicates whether the predicted cleavage must retain the assignment to residue i in order for the element in the output protein sequence x to be

[0071] As shown in FIG4A , in some application scenarios, given a protein sequence 404 corresponding to a known protein structure 406, it is desirable to obtain a new protein sequence that has amino acids identical to those at certain positions in protein sequence 404. In this case, a designated input sequence 402 can be constructed, wherein the elements at corresponding positions indicate the amino acids at the same positions in protein sequence 404. Designated input sequence 402 is provided as input to DPLM 210. After processing by DPLM 210, the amino acids at certain positions in the generated target protein sequence 410 correspond to the designated amino acids at the same positions in protein sequence 404. Based on target protein sequence 410, a corresponding new protein structure 412 can be generated.

[0072] In some embodiments, the cross-modal conditional generation refers to the introduction of non-sequence modal protein data as constraints in protein sequence generation. The process of generating protein sequences is subject to the cross-modal constraint c, that is, x~p θ (x|c). For example, in reverse protein folding, it is desirable to generate a protein sequence for a given backbone structure (also known as protein secondary structure). Considering that DPLM 210 operates on sequences with amino acids as tokens, in this case, an additional adapter needs to be configured for DPLM 210 to support the generation of expanded model conditions.

[0073] Figure 4B illustrates an example of a conditional generation task conditioned on secondary structure, according to some embodiments of the present disclosure. As shown in Figure 4B , an adapter 420 is introduced into DPLM 210 and connected to the existing Transformer layer 212 in DPLM 210. In this case, the pre-trained DPLM 210 is considered to be a modality encoder ε(c). Adapter 420 is configured to predict the corresponding protein sequence based on the sequence feature representation extracted by modality encoder ε(c) and the input protein data in a non-sequential modality.

[0074] During training, the pre-trained DPLM 210 can be fixed, keeping its parameters unchanged. The adapter 420 is then trained based on supervised training. The training data for supervised training includes sample protein data in a non-sequential modality and the protein sequences corresponding to the sample protein data. In some embodiments, the non-sequential protein data includes protein secondary structures. When training the adapter 420, to support supervised training, a certain number of pairs of sample protein secondary structures and their corresponding protein sequences need to be collected, represented as (x, c), where the extended modal constraint c represents the protein secondary structure and x represents the corresponding protein sequence.

[0075] Since the DPLM 210 has been pre-trained, only a small number of paired training samples (x, c) are needed in the supervised training process of the adapter 420 to complete the training of the adapter 420. The DPLM 210 with the adapter 420 with the expanded mode can be expressed as the conditional DPLM p θ (x|ε(c)).

[0076] After training is complete, the input sequence input to the DPLM 210 with the modality-expanding adapter 420 can include a random input sequence in the form of a protein sequence or a specified input sequence (i.e., the amino acids of a portion of the sequence are specified). Figure 4B shows a random input sequence 424 being input to the DPLM 210. Furthermore, the DPLM 210 will function as a modality encoder to extract sequence feature representations for the random input sequence or the specified input sequence.

[0077] In addition, input protein data in a non-sequential modality, such as protein secondary structure 422, is input to adapter 420. Adapter 420 generates a target protein sequence 426 corresponding to protein secondary structure 422 based on the input protein data in a non-sequential modality (i.e., protein secondary structure 422) and the sequence feature representation extracted by Transformer layer 212. Target protein sequence 426 can be folded to have protein secondary structure 422.

[0078] It should be understood that, in addition to protein secondary structure, based on the pre-trained DPLM 210, adapters can also be trained to adapt to protein data of other non-sequential modalities.

[0079] In some embodiments, in protein sequence generation, it is also desirable to support plug-and-play controllable sequence generation tasks. Generally, it is difficult to directly construct a conditional generation model. Therefore, introducing classifier guidance in the continuous diffusion model is a way to implement the conditional generation model, by using pre-trained classifiers to guide the generation process towards the desired preference information. However, the premise of continuous classifier guidance is that it is logarithmically differentiable, that is, This will not hold for a discrete diffusion process.

[0080] In some embodiments of the present disclosure, in order to support controllable sequence generation in a discrete diffusion model structure, it is proposed to introduce discrete classifier guidance into the pre-trained DPLM 210 to achieve such controllable generation tasks.

[0081] Figure 4C illustrates an example of a controllable sequence generation task with guidance information, according to some embodiments of the present disclosure. The input sequence processed by DPLM 210 can include a random input sequence or a specified input sequence in the form of a protein sequence. Figure 4C shows a random input sequence 436 being input to DPLM 210. During protein sequence generation, a trained target classifier 430 is used to determine the reference attribute category corresponding to the reference protein data.

[0082] The reference protein data can correspond to protein data under any attribute. For example, the reference protein data includes protein secondary structure. In the example of FIG4C , annotation information 434 of the corresponding protein secondary structure can be determined from reference protein structure 432. In this example, target classifier 430, also referred to as a secondary structure classifier, classifies the protein secondary structure annotation information 434 into one of a plurality of protein secondary structure categories.

[0083] During the sequence generation process of DPLM 210, under the guidance of the reference attribute categories (e.g., protein secondary structure categories), DPLM 210 is used to generate a target protein sequence 438 based on an input sequence in the form of a protein sequence. Target protein sequence 438 will have the reference attribute categories as guiding information, although the target protein sequence may be different from the protein sequence corresponding to protein structure 432.

[0084] In addition to protein secondary structure, other protein properties can also be used as guidance information to guide DPLM 210 in generating protein sequences with the same attribute category. This allows for convenient generation of protein sequences based on preferences. In this approach, different classifiers can be easily introduced to classify protein data with corresponding attributes.

[0085] In the process of conditional generation based on guided information, we want to try to get the conditional distribution q(x (t-1) |x (t) ,y)∝q(x (t-1) |x (t) )q(y|x (t-1) ), where y represents the conditional guidance information (also called preference information), and such conditional distribution is approximately p θ (x (t-1) |x (t) )p φ (y|x (t-1) ), where p φ (y|x (t-1) ) is a discriminative guidance model (e.g., a classifier or regressor with guidance information about the user input). However, pφ (y|x (t-1) ) cannot be decomposed in all positions, so x (t-1) To this end, in the embodiment of the present disclosure, first, in the DPLM 210, the sequence data x associated with the discrete protein sequence is considered as a continuous random variable, and x is evaluated. (t) Perform a first-order Taylor expansion to obtain:

[0086] Where C(x (t) ) is not dependent on x (t-1) Use p φ (y|x (t) ) to estimate q(y|x (t) ) and insert it into Equation (8) above. This allows sampling from the conditional distribution instead of sampling at every step t, which can be expressed as follows:

[0087] Where η is an adjustable parameter used to control the degree of control of the guidance information on sequence generation.

[0088] FIG5 illustrates a flowchart of a protein-related information processing process according to some embodiments of the present disclosure. For ease of discussion, process 500 will be described with reference to environment 100 of FIG1 . Process 500 may be implemented in protein-related processing system 110. As previously mentioned, protein-related processing system 110 may be a server-side device or a terminal device, and the scope of the embodiments of the present disclosure is not limited in this regard.

[0089] At block 510 , the processing system 110 obtains a target model, which is based on a discrete diffusion probability model and a language model structure and is pre-trained using a set of real protein sequences.

[0090] In block 520 , the processing system 110 utilizes the target model to perform at least one of the following: utilizing the target model to generate a first target protein sequence based at least on an input sequence in the form of a protein sequence, or utilizing the target model to extract a sequence feature representation corresponding to the second target protein sequence based on the second target protein sequence.

[0091] In some embodiments, the input sequence comprises a random input sequence in the form of a protein sequence. Based on at least the input sequence in the form of a protein sequence, generating the first target protein sequence using the target model comprises: providing the random input sequence to the target model to generate the first target protein sequence generated by the target model.

[0092] In some embodiments, the input sequence comprises a designated input sequence in the form of a protein sequence, and an element at at least one position in the designated input sequence indicates at least one designated amino acid. Based on at least the input sequence in the form of a protein sequence, generating a first target protein sequence using the target model comprises: providing the designated input sequence to the target model to generate a first target protein sequence as output by the target model, wherein the element at at least one position in the first target protein sequence indicates the at least one designated amino acid.

[0093] In some embodiments, the input sequence includes a random input sequence or a designated input sequence in the form of a protein sequence. Generating a first target protein sequence using a target model based on at least the input sequence in the form of a protein sequence includes: determining a reference attribute category corresponding to reference protein data using a trained target classifier; and generating the first target protein sequence using the target model based on the input sequence in the form of a protein sequence, guided by the reference attribute category, wherein the first target protein sequence has the reference attribute category.

[0094] In some embodiments, the reference protein data includes protein secondary structure, and the reference property class includes one of a plurality of protein secondary structure classes.

[0095] In some embodiments, process 500 further includes: with the pre-trained target model fixed, training an adapter based on supervised training, the adapter being configured to predict a corresponding protein sequence based on sequence feature representation and protein data in a non-sequential modality, the training data for supervised training including sample protein data in a non-sequential modality and protein sequences corresponding to the sample protein data.

[0096] In some embodiments, the input sequence includes a random input sequence or a designated input sequence in the form of a protein sequence. Based on at least the input sequence in the form of a protein sequence, generating a first target protein sequence using a target model includes: providing the random input sequence or the designated input sequence to the target model to extract a sequence feature representation of the random input sequence or the designated input sequence; and generating the first target protein sequence corresponding to the input protein data based on the input protein data in a non-sequence modality and the sequence feature representation using an adapter.

[0097] In some embodiments, the non-sequence modality protein data includes protein secondary structure.

[0098] In some embodiments, process 500 further includes: based on the sequence feature representation corresponding to the second target protein sequence, performing at least one of the following tasks: a residue classification task for the second target protein sequence, a sequence classification task for the second target protein sequence, a sequence regression task for the second target protein sequence, and a contact point prediction task for the second target protein sequence.

[0099] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG6 shows a block diagram of an apparatus 600 for document querying according to certain embodiments of the present disclosure. Apparatus 600 can be implemented as or included in protein-related processing system 110. Each module / component in apparatus 600 can be implemented by hardware, software, firmware, or any combination thereof.

[0100] As shown in FIG6 , apparatus 600 includes a model acquisition module 610 configured to obtain a target model, where the target model is based on a discrete diffusion probability model and a language model structure and is pre-trained using a set of real protein sequences. Apparatus 600 also includes a model execution module 620 configured to use the target model to perform at least one of the following: generating a first target protein sequence based on at least an input sequence in the form of a protein sequence using the target model; or extracting a sequence feature representation corresponding to a second target protein sequence based on the second target protein sequence using the target model.

[0101] In some embodiments, the input sequence comprises a random input sequence in the form of a protein sequence. Based on at least the input sequence in the form of a protein sequence, generating the first target protein sequence using the target model comprises: providing the random input sequence to the target model to generate the first target protein sequence generated by the target model.

[0102] In some embodiments, the input sequence includes a specified input sequence in the form of a protein sequence, and an element at at least one position in the specified input sequence indicates at least one specified amino acid. The model execution module 620 is configured to provide the specified input sequence to the target model to generate a first target protein sequence as output by the target model, wherein the element at at least one position in the first target protein sequence indicates at least one specified amino acid.

[0103] In some embodiments, the input sequence includes a random input sequence or a specified input sequence in the form of a protein sequence. The model execution module 620 is configured to: determine a reference attribute category corresponding to the reference protein data using a trained target classifier; and generate a first target protein sequence having the reference attribute category based on the input sequence in the form of a protein sequence using the target model under the guidance of the reference attribute category.

[0104] In some embodiments, the reference protein data includes protein secondary structure, and the reference property class includes one of a plurality of protein secondary structure classes.

[0105] In some embodiments, the device 600 also includes an adapter training module, which is configured to: train the adapter based on supervised training when the pre-trained target model is fixed, and the adapter is configured to predict the corresponding protein sequence based on the sequence feature representation and the protein data of the non-sequential modality, and the training data of the supervised training includes the sample protein data of the non-sequential modality and the protein sequence corresponding to the sample protein data.

[0106] In some embodiments, the input sequence includes a random input sequence or a specified input sequence in the form of a protein sequence. The model execution module 620 is configured to: provide the random input sequence or the specified input sequence to the target model to extract a sequence feature representation of the random input sequence or the specified input sequence; and generate, using an adapter, a first target protein sequence corresponding to the input protein data based on the input protein data in a non-sequence modality and the sequence feature representation.

[0107] In some embodiments, the non-sequence modality protein data includes protein secondary structure.

[0108] In some embodiments, the device 600 also includes a downstream task execution module, which is configured to: based on the sequence feature representation corresponding to the second target protein sequence, perform at least one of the following tasks: a residue classification task for the second target protein sequence, a sequence classification task for the second target protein sequence, a sequence regression task for the second target protein sequence, and a contact point prediction task for the second target protein sequence.

[0109] The units and / or modules included in the device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 600 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0110] FIG7 shows a block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 700 shown in FIG7 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 700 shown in FIG7 can be used to implement the protein-related processing system 110 of FIG1 or the apparatus 600 of FIG6 .

[0111] As shown in FIG7 , electronic device 700 is a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 700.

[0112] The electronic device 700 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory) or some combination thereof. The storage device 730 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk or any other medium, which can be used to store information and / or data and can be accessed within the electronic device 700.

[0113] The electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0114] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 700 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 700 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0115] Input device 750 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 760 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 700 may also communicate with one or more external devices (not shown) via communication unit 740 as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with electronic device 700, or with any device that allows electronic device 700 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0116] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0117] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0118] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0119] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0120] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0121] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. An information processing method, comprising: Obtaining a target model, wherein the target model is based on a discrete diffusion probability model and a language model structure and is pre-trained using a set of real protein sequences; as well as Utilize the target model to perform at least one of the following: generating a first target protein sequence using the target model based on at least an input sequence in the form of a protein sequence, or Based on the second target protein sequence, the target model is used to extract a sequence feature representation corresponding to the second target protein sequence.

2. The method of claim 1 , wherein the input sequence comprises a random input sequence in the form of a protein sequence, and wherein generating a first target protein sequence using the target model based on at least the input sequence in the form of a protein sequence comprises: The random input sequence is provided to the target model to generate a first target protein sequence generated by the target model.

3. The method of claim 1 , wherein the input sequence comprises a specified input sequence in the form of a protein sequence, an element at at least one position in the specified input sequence indicates at least one specified amino acid, and wherein generating a first target protein sequence using the target model based on at least the input sequence in the form of a protein sequence comprises: The designated input sequence is provided to the target model to generate a first target protein sequence output by the target model, wherein the element of the at least one position in the first target protein sequence indicates the at least one designated amino acid.

4. The method of claim 1 , wherein the input sequence comprises a random input sequence or a designated input sequence in the form of a protein sequence, and wherein generating a first target protein sequence using the target model based on at least the input sequence in the form of a protein sequence comprises: Determine the reference attribute category corresponding to the reference protein data using the trained target classifier; as well as Under the guidance of the reference attribute category, the target model is used to generate a first target protein sequence based on an input sequence in the form of a protein sequence, where the first target protein sequence has the reference attribute category. 5 . The method of claim 4 , wherein the reference protein data comprises protein secondary structure, and the reference attribute category comprises one of a plurality of protein secondary structure categories.

6. The method according to claim 1, further comprising: With the pre-trained target model fixed, an adapter is trained based on supervised training, wherein the adapter is configured to predict a corresponding protein sequence based on sequence feature representation and protein data of a non-sequential modality, wherein the training data of the supervised training includes sample protein data of the non-sequential modality and a protein sequence corresponding to the sample protein data.

7. The method according to claim 6, wherein the input sequence comprises a random input sequence or a designated input sequence in the form of a protein sequence, and wherein generating a first target protein sequence using the target model based on at least the input sequence in the form of a protein sequence comprises: Providing the random input sequence or the specified input sequence to the target model to extract sequence feature representation of the random input sequence or the specified input sequence; as well as The adapter is used to generate a first target protein sequence corresponding to the input protein data based on the input protein data in the non-sequence modality and the sequence feature representation. The method according to claim 6 or 7, wherein the protein data in the non-sequence modality includes protein secondary structure.

9. The method according to any one of claims 1 to 8, further comprising: Based on the sequence feature representation corresponding to the second target protein sequence, perform at least one of the following tasks: For the residue classification task of the second target protein sequence, For the sequence classification task of the second target protein sequence, For the sequence regression task of the second target protein sequence, Contact point prediction task for the second target protein sequence.

10. An apparatus for information processing, comprising: A model acquisition module is configured to obtain a target model, wherein the target model is based on a discrete diffusion probability model and a language model structure and is pre-trained using a set of real protein sequences; as well as A model execution module is configured to use the target model to perform at least one of the following: generating a first target protein sequence using the target model based on at least an input sequence in the form of a protein sequence, or Based on the second target protein sequence, the target model is used to extract a sequence feature representation corresponding to the second target protein sequence.

11. An electronic device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processing unit.

12. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Molecular binding conformation determination method and device

    CN116580764A

  • Protein data processing method and device, electronic equipment and storage medium

    CN116978450A

  • Discrete graph probability denoising diffusion model for protein sequence generation

    CN116994642A

  • Protein sequence and structure generation with denoising diffusion probabilistic models

    US20230377690A1