Methods and apparatus for processing protein data
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YOUZHUJU NETWORK TECH CO LTD
- Filing Date
- 2024-09-30
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies struggle to effectively combine protein sequence and structural information for modeling, resulting in insufficient accuracy and efficiency in protein data processing, particularly in generation and prediction tasks.
By employing a multimodal machine learning model that combines a discrete diffusion probability model and a language model structure, the three-dimensional structure of a protein is converted into discrete labels through a structural encoder and decoder. The model is then trained using sequence features to form the multimodal protein language model DPLM-2, which solves the problem of joint modeling of protein sequences and structures.
It achieves high consistency in the simultaneous generation of protein structures and sequences, supports multiple tasks under both unconditional and conditional conditions, improves the performance of protein prediction tasks, reduces training data requirements, and enhances generation quality and accuracy for downstream tasks.
Smart Images

Figure CN122139224A_ABST
Abstract
Description
Method and apparatus for processing protein data TECHNICAL FIELD
[0001] Exemplary implementations of the present disclosure generally relate to machine learning, and in particular, to methods, apparatuses, devices, and computer-readable storage media for processing protein data using a multi-modal machine learning model. BACKGROUND
[0002] Proteins are folded linear (1D) sequences with three-dimensional (3D) structures composed of amino acid residues. Proteins perform many important life activities in living organisms and play an important role in regulating various biological functions, including transcription, translation, signal transduction, and life cycle control of cells. With the development of data-driven machine learning techniques, in the field of protein analysis, there has been a gradual evolution from traditional methods to using models to learn how to understand and design proteins. It is desirable to be able to develop powerful models to complete various tasks in the field of protein research.
[0003] SUMMARY
[0004] In a first aspect of the present disclosure, a method for processing protein data using a multi-modal machine learning model is provided. In the method, a target model is obtained, the target model is based on a discrete diffusion probability model and a language model structure, and is pre-trained using a protein data set. At least one of the following is performed using the target model: determining first target protein data using the target model based on a first feature representation corresponding to the protein data, the first feature representation including: first sequence features corresponding to a protein sequence of the first protein data, and first structure features corresponding to a protein structure of the first protein data; or extracting a second feature representation corresponding to second target protein data using the target model based on the second target protein data, the second feature representation including: second sequence features corresponding to a protein sequence of the second protein data, and second structure features corresponding to a protein structure of the second protein data.
[0005] In a second aspect of the disclosure, an apparatus for processing protein data using a multi-modal machine learning model is provided. The apparatus comprises: an obtaining module configured to obtain a target model, the target model being based on a discrete diffusion probability model and a language model structure, and being pre-trained using a protein dataset; and an executing module configured to perform at least one of: determining, using the target model, first target protein data based on a first feature representation corresponding to the protein data, the first feature representation comprising: first sequence features corresponding to a protein sequence of the first protein data, and first structure features corresponding to a protein structure of the first protein data; or extracting, using the target model, a second feature representation corresponding to second target protein data based on the second target protein data, the second feature representation comprising: second sequence features corresponding to a protein sequence of the second protein data, and second structure features corresponding to a protein structure of the second protein data.
[0006] In a third aspect of the disclosure, an electronic device is provided. The electronic device comprises: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the disclosure.
[0007] In a fourth aspect of the disclosure, a computer-readable storage medium is provided, having stored thereon a computer program, which, when executed by a processor, causes the processor to implement the method according to the first aspect of the disclosure.
[0008] In a fifth aspect of the disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the disclosure.
[0009] It is to be understood that the details set forth herein are not intended to limit the key or critical features of the implementations of the disclosure, nor are they intended to limit the scope of the disclosure. Other features of the disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other features, aspects, and advantages of various implementations of the present disclosure will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings, in which like reference numerals represent like elements, wherein:
[0011] FIG. 1 shows a block diagram of an application environment according to one example implementation of the disclosure;
[0012] FIG. 2 shows a block diagram of processing protein data using a multi-modal machine learning model according to some implementations of the disclosure;
[0013] FIG. 3 shows a block diagram for determining a structure encoder and a structure decoder according to some implementations of the present disclosure;
[0014] FIG. 4 shows a block diagram of a correspondence between sequence features and structure features according to some implementations of the present disclosure;
[0015] FIG. 5 shows a block diagram of determining a target model according to some implementations of the present disclosure;
[0016] FIG. 6 shows a block diagram of determining a target model based on an initial model according to some implementations of the present disclosure;
[0017] FIG. 7 shows a block diagram of extracting a feature representation of protein data with a target model according to some implementations of the present disclosure;
[0018] FIG. 8 shows a block diagram of a conditional generation task conditioned on partial sequences according to some implementations of the present disclosure;
[0019] FIG. 9 shows a flowchart of a method of processing protein data with a multi-modal machine learning model according to some implementations of the present disclosure;
[0020] FIG. 10 shows a block diagram of an apparatus of processing protein data with a multi-modal machine learning model according to some implementations of the present disclosure; and
[0021] FIG. 11 shows a block diagram of a device capable of implementing the multiple implementations of the present disclosure. DETAILED DESCRIPTION
[0022] Implementations of the present disclosure will be described below in detail with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the implementations set forth herein, but rather, these implementations are provided for more thorough and complete understanding of the present disclosure. It is understood that the drawings and implementations of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0023] In the description of the implementations of the disclosure, the term "comprising" and similar terms are to be construed as open-ended, that is, "including but not limited to". The term "based on" is to be construed as "based at least in part on". The term "one implementation" or "the implementation" is to be construed as "at least one implementation". The term "some implementations" is to be construed as "at least some implementations". Other explicit and implicit definitions can also be included below. As used herein, the term "model" can represent the association relationship between various data. For example, the above association relationship can be obtained based on various technical solutions that are currently known and / or will be developed in the future.
[0024] It can be understood that the data involved in the technical solutions of the present application (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0025] It can be understood that before using the technical solutions disclosed by the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to the relevant laws and regulations.
[0026] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as the electronic device, the application program, the server or the storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.
[0027] As an optional but non-limiting implementation, in response to receiving the active request of the user, the way of sending the prompt information to the user may, for example, be the way of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0028] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementations of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementations of the present disclosure.
[0029] The term "in response to" used herein indicates the state that the corresponding event occurs or the condition is met. It will be understood that the timing of the execution of the subsequent action performed in response to the event or condition is not necessarily strongly associated with the time when the event occurs or the condition is established. For example, in some cases, the subsequent action can be performed immediately when the event occurs or the condition is established; while in other cases, the subsequent action can be performed after a period of time after the event occurs or the condition is established.
[0030] Example Environment
[0031] Proteins are macromolecules that perform essential functions in living organisms. Proteins involve two modalities: amino acid sequence and three-dimensional structure. Therefore, the ultimate goal of generative protein modeling is to model, understand, and generate both modalities of sequence and structure of a protein simultaneously. Prior art solutions often focus on modeling each modality independently, e.g., language modeling for sequence or diffusion modeling for backbone structure. To model the whole protein, a cascaded workflow is needed, i.e., first generate the backbone structure, then process the sequence, or perform structure prediction or folding after generating the sequence.
[0032] It has been proposed to utilize models to learn how to understand and design proteins. See FIG. 1 depicting an application environment according to some implementations of the present disclosure, which illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the environment 100, a protein-related processing system 110 can utilize a target model 120 to perform a protein-related analysis task based on a specified input 102. In protein-related analysis, protein sequence prediction, protein structure prediction, representation learning of protein sequence, etc. are all practically meaningful tasks. In some implementations, the processing system 110 can utilize the target model 120 to generate a protein sequence based on the input 102. The processing system 110 can utilize the target model 120 to determine a sequence representation corresponding to the protein sequence based on the input 102 (e.g., an input protein sequence). Based on the accurate sequence representation, various understanding tasks can be performed, including classification of protein sequence, classification of residues in a protein, protein sequence regression, protein contact point prediction, etc.
[0033] In FIG. 1, the processing system 110 can be any type of computing-enabled device, including a terminal device or a server device. The terminal device can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a gaming device, or any combination of the foregoing, including accessories and peripherals of such devices or any combination thereof. The server device may, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, etc. It should be understood that the structure and functionality of the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0034] In the field of machine learning, language models (LMs) have made remarkable progress in the field of natural language understanding. Considering the similarity between protein sequences and human languages, language models have been utilized in the field of protein research to achieve various tasks, including protein-related prediction tasks (e.g., detecting functional attributes, predicting protein structures from a single sequence without evolutionary homolog chromosomes) and protein-related generation tasks (e.g., given a given protein backbone structure, redesigning an optimized protein sequence, or synthesizing a brand new protein sequence).
[0035] Proteins are macromolecules that play essential functions in all living organisms. Proteins are characterized by their amino acid sequences and three-dimensional structures, where the sequence determines the structure, which indicates the protein. Therefore, the ultimate goal of generative protein modeling is to model, understand, and generate both sequence and structure modalities of proteins simultaneously, rather than in isolation. Protein generation models have generally focused on a single modality, either sequence or structure. Diffusion models have shown great success in protein structure and sequence modeling and have become one of the protein base models for protein sequence representation learning and generation. It has been shown that discrete diffusion based on protein language models is well suited for modeling protein sequences at the evolutionary scale.
[0036] It has been proposed to utilize language models to process protein sequences in protein data. At this time, various performances of machine learning models are not satisfactory. Further, large-scale data poses a significant challenge to learning protein structures, and it is difficult to convert three-dimensional structure information to the feature space of language models. At this time, it is desirable that a pre-trained protein language model can be reused as a multi-modal protein language model, and various tasks focused on in the field of protein research can be accomplished using the model.
[0037] Summary of protein data processing
[0038] To at least partially address the deficiencies in the prior art, according to one example implementation of the present disclosure, a method of processing protein data using a multi-modal machine learning model is proposed. In the present disclosure, a multi-modal protein language model is proposed, which models both protein structure and sequence simultaneously. The proposed technical solution supports unconditional sampling design and can generate diverse protein structures and structurally plausible protein sequences. In other words, the model has high consistency between structure and sequence. In addition, the proposed technical solution supports various conditional downstream tasks, such as inverse folding, folding, and motif-scaffolding. The proposed technical solution can perform well as a supervised representation learner for protein prediction tasks. By introducing residue representations including spatial relationships and attention bias, the ability of the language model to capture complex spatial dependencies in protein structures can be enhanced. Further, the proposed technical solution adopts a hybrid training method of discrete diffusion models, which can alleviate the exposure bias problem. In addition, the technical solution only needs to utilize a relatively small amount of high-quality structure data to efficiently convert a pre-trained protein language model to a multi-modal protein language model.
[0039] A summary according to one example implementation of the present disclosure is described with reference to FIG. 2, which shows a block diagram 200 of processing protein data using a multi-modal machine learning model according to some implementations of the present disclosure. As shown in FIG. 2, a target model 250 can be obtained, which is based on a discrete diffusion probability model and a language model structure, and is pre-trained using a real protein dataset 262 (ground truth data). Here, the protein data in the protein dataset can include data of multiple modalities: protein sequence (1D data) and protein structure (3D data). The target model 250 can consider multiple modalities comprehensively, thereby describing aspects of the protein in a more accurate manner.
[0040] Further, the target model 250 can be used to perform at least one of: determining, using the target model, first target protein data based on a first feature representation corresponding to the protein data, the first feature representation including: first sequence features corresponding to a protein sequence of the first protein data, and first structure features corresponding to a protein structure of the first protein data; or extracting, using the target model, a second feature representation corresponding to second target protein data based on the second target protein data, the second feature representation including: second sequence features corresponding to a protein sequence of the second protein data, and second structure features corresponding to a protein structure of the second protein data.
[0041] As shown in FIG. 2, a first feature representation (e.g., feature representation 230) can be input to a target model 250 in order to determine first target protein data (e.g., protein data 260). It should be appreciated that the feature representation 230 herein can include two parts: sequence features (e.g., sequence features 212) corresponding to protein sequences (e.g., protein sequences 210) of the protein data, and structure features (e.g., structure features 222) corresponding to protein structures (e.g., protein structures 220) of the protein data. Generally speaking, the protein sequences and the protein structures have a correspondence relationship, and the corresponding protein structure can be determined from the protein sequence by a folding 240 operation. Alternatively and / or additionally, the corresponding protein sequence can be determined from the protein structure by an inverse folding 242 operation.
[0042] Alternatively and / or additionally, the target model 250 can be used in the opposite direction. For example, second target protein data (e.g., protein data 260) can be input to the target model 250 in order to extract a second feature representation (e.g., feature representation 230) corresponding to the second target protein data. The feature representation 230 herein can include two parts: sequence features (e.g., sequence features 212) corresponding to protein sequences (e.g., protein sequences 210) of the protein data, and structure features (e.g., structure features 222) corresponding to protein structures (e.g., protein structures 220) of the protein data. It should be appreciated that the protein sequences and the protein structures are important indicators for characterizing proteins. With some implementations of the present disclosure, the protein data can be described in a more accurate manner, thereby helping to improve the accuracy of downstream tasks.
[0043] On the basis of a first generation diffusion protein language model (DPLM), a second generation diffusion protein language model (DPLM-2) is proposed in the present disclosure. The model can be used for simultaneous diffusion generative training of protein structures and sequences, thereby realizing a multi-modal diffusion protein language model. Specifically, a structure encoder / decoder (i.e., a pair of tokenizer and de-tokenizer) can be used for discretization of protein structures. The structure encoder can convert 3D coordinates of the protein structure into discrete tokens, thereby facilitating language modeling. The structure decoder can perform the reverse operation.
[0044] For the training phase, DPLM-2 models the unimodal, conditional, and joint distributions of structure and sequence, which covers various unconditional and conditional generation scenarios. Notably, this joint modeling enables the simultaneous generation of structure and sequence with cross-conditions, which means that the cascading generation paradigm is no longer needed. However, discrete diffusion training will face the exposure bias problem, which means the mismatch between training and inference, hindering the generation quality. To address this issue, a hybrid training paradigm for discrete diffusion models is proposed, which enhances the consistency between training and inference. After training, DPLM-2 can be used for various structure or sequence generation needs and provides informative representations for downstream tasks. For example, DPLM-2 can generate discrete structure labels and amino acid sequences, where the discrete structure labels can then be converted to 3D space coordinates by a structure decoder.
[0045] Detailed procedure of protein data processing
[0046] Having described an overview of some implementations according to the present disclosure, in the following, more details of processing protein data will be described. According to some implementations of the present disclosure, the goal of generative protein modeling is to estimate the underlying distribution of protein data of interest y ~ q(y) by learning a probabilistic model P θ (y). L ) represents a protein with L residues (corresponding to amino acids), where each residue is represented as y i = [a i ,x i ] T , including two modalities: represents the amino acid type variable in , and represents the real-valued Cartesian coordinates (i.e., 3D coordinates) of the backbone atoms. At this time, there is the following relationship: θ (y) = p θ (a1,a2,…,a L ,x1,x2,…,x L ) = p θ (a,x)
[0047] Details about the diffusion protein language model are described below. Language models (LMs) parameterized by Transformers have become the conventional choice dominating different domains with scalable and expressive representations. Protein language models have been used as one of the base models for protein sequence learning and generation.
[0048] In particular, DPLM exhibits excellent performance in protein sequence generation and representation learning. DPLM's modeling is based on a discrete diffusion framework, characterized by forward and reverse Markov processes. Let Cat(y;p) represent the classification distribution on protein sequence y, which is derived from... The vector p on the probabilistic simplex is used for parameterization. The forward process of discrete diffusion is defined by the kernel q(y). (t) |y (t-1) ) = Cat(y (t) ;β t y (t-1) +(1-β t )q noise The Markov process governed by ) gradually transfers the data y (0) ~q(y (0) The disturbance is a fixed distribution y (T) ~q noise The backward learning process p θ (y (t-1) |y (t) ) will y (T) Towards the data distribution y (0) To denoise in reverse, this is typically optimized using the variational bound of the log-likelihood:
[0049] In the above formula, The learning objective is represented by this. The learning objective of discrete diffusion can be further simplified to reweighted cross-entropy, similar to masked language modeling under arbitrary noise levels:
[0050] In the above formula, λ (t) This represents the weighting coefficient caused by a specific noise dispatch. Considering absorption and discrete diffusion, q noise This represents the point mass representing all probabilities at the absorption state. Discrete diffusion from the following distribution, through a reverse iterative denoising process, allows DPLM to synthesize new amino acid sequences:
[0051] Specifically, at time t, first from p θ (·|y (t) )generate Then through For less noise y (t-1) Sampling is performed. In the form of discrete diffusion absorption, the generation process can be viewed as an iterative mask prediction method. DPLM can accept input sequences because its task is to remove noise from the input at all noise levels (including the original input data).
[0052] According to some implementations of the present disclosure, protein sequences in protein data can be processed in a manner similar to DPLM. Further, protein structures in protein data can be processed with a structure encoder and a structure decoder to convert 3D structures into tokens similar to those in a language model.
[0053] According to some implementations of the present disclosure, a first feature representation corresponding to protein data is determined based on: determining first sequence features with a sequence encoder based on protein sequences of the protein data; and determining first structure features with a structure encoder based on protein structures of the protein data, the structure encoder mapping continuous three-dimensional data of the protein structures to a discrete one-dimensional representation. Specifically, a structure decoder can be determined based on a variety of encoding-decoding algorithms, as long as a structure encoder corresponding to the structure decoder is able to recover from the encoded one-dimensional representation to a three-dimensional representation. With some implementations of the present disclosure, multiple modalities in protein data can be processed separately in different ways, and converted to token formats processable by a language model.
[0054] According to some implementations of the present disclosure, a structure encoder and a structure decoder can be trained. Specifically, an initial structure encoder and an initial structure decoder corresponding to the initial structure encoder can be obtained, and the initial structure encoder is updated to determine the structure encoder and the initial structure decoder is updated to determine the structure decoder based on protein structures in a protein structure dataset, the structure decoder corresponding to the structure encoder. With some implementations of the present disclosure, the structure encoder and the structure decoder can be trained simultaneously to map continuous three-dimensional data to discrete one-dimensional token representations with higher accuracy.
[0055] A specific procedure for processing protein structures is described with reference to FIG. 3, which illustrates a block diagram 300 for determining a structure encoder and a structure decoder according to some implementations of the present disclosure. In FIG. 3, parameters of a structure encoder (e.g., encoder 310) and a structure decoder (e.g., decoder 320) can be constantly updated to improve the accuracy of the encoding-decoding process. A structure dataset 330 can include a large number of protein structures, for one protein structure 220, the protein structure 220 can be encoded into structure features 222 with the encoder 310. Further, the structure features 222 can be decoded with the decoder 320 to recover a protein structure 222’. In the training process, differences between the protein structure 220 and the protein structure 222’ can be compared, and parameters of the encoder 310 and the decoder 320 can be updated towards a direction to reduce the differences.
[0056] In the initial stage of training, the encoder 310 and the decoder 320 can also be referred to as an initial structure encoder and an initial structure decoder. A large number of protein structures in the structure dataset 330 can be utilized to iteratively update the parameters of the encoder 310 and the decoder 320, so as to obtain the encoder 310 and the decoder 320 with higher precision.
[0057] It should be understood that the protein data can include a plurality of amino acids, and each amino acid can have a respective sequence and structure. At this time, the sequence feature and the structure feature can respectively include a plurality of positions respectively corresponding to the plurality of amino acids in the protein data. Specifically, the first sequence feature includes a first plurality of positions, the first plurality of positions respectively correspond to the plurality of amino acids in the protein data, the first structure feature includes a second plurality of positions, the first plurality of positions are equal to the second plurality of positions, and the second plurality of positions respectively correspond to the plurality of amino acids. With some implementations of the present disclosure, the protein data can be described in a more accurate manner.
[0058] More details are described with reference to FIG. 4, which shows a block diagram 400 of the correspondence between the sequence feature and the structure feature according to some implementations of the present disclosure. As shown in FIG. 4, it is assumed that the protein data 440 includes L amino acids, for example, a plurality of amino acids 410-1, 410-2, …, and 410-L represented in a circular shape. At this time, the sequence feature 212 can include L positions, and each position respectively corresponds to each amino acid. That is, the mark of the related amino acid sequence of the i-th amino acid is stored at the i-th position. Specifically, for example, the mark of the related amino acid sequence of the amino acid 410-1 is stored at the position 420-1, the mark of the related amino acid sequence of the amino acid 410-2 is stored at the position 420-2, …, and the mark of the related amino acid sequence of the amino acid 410-2 is stored at the position 420-2.
[0059] Further, the structure feature 222 can include L positions, and each position respectively corresponds to each amino acid. That is, the mark of the related amino acid structure of the i-th amino acid is stored at the i-th position. Specifically, for example, the mark of the related amino acid structure of the amino acid 410-1 is stored at the position 430-1, the mark of the related amino acid structure of the amino acid 410-2 is stored at the position 430-2, …, and the mark of the related amino acid structure of the amino acid 410-2 is stored at the position 430-2. With some implementations of the present disclosure, the sequence and structure of each amino acid in the protein data can be aligned, so that the feature representation of the protein data represented in this way is more accurate in describing aspects of the protein data.
[0060] The structure of the target model is described with reference to FIG. 5, which shows a block diagram 500 of determining a target model according to some implementations of the present disclosure. As shown in FIG. 5, the target model 510 (e.g., DPLM-2) can be determined based on a discrete diffusion probability model and a language model structure. According to some implementations of the present disclosure, DPLM-2 can employ a discrete diffusion probability framework similar to DPLM and simultaneously model protein sequences and their corresponding structures. In FIG. 5, the protein sequence 210 and the protein structure 220 of the protein data in the protein dataset 260 can be utilized to generate sequence features 212 and structure features 222. Specifically, the protein structure can be encoded into a structure sequence by the encoder 310.
[0061] As shown in FIG. 5, DPLM-2 can introduce a token-based representation for protein structures. This is achieved by a structure encoder 310, which can convert the 3D coordinates of a protein backbone geometry into a discrete structure token sequence, denoted as Here, each token s i represents a local structure element associated with the i-th residue.
[0062] Further, the feature representation 230 of the protein data can be determined by utilizing the sequence features 212 and the structure features 222. Specifically, the model processes the input by concatenating the structure token sequence s with the corresponding amino acid sequence a of the same protein. Notably, there is a position-position correspondence between s and a, where s i and a i denote the 3D and 1D representation of the same amino acid, respectively. To reinforce this correspondence, the same position encoding can be assigned to s i and a i , thus ensuring that structure and sequence information are aligned at the residue level. The training objective of the multi-modal model can be determined in a similar manner as the DPLM above, e.g., the multi-modal training objective of DPLM-2 can be derived from Equation (1) as follows:
[0063] During training, the task of DPLM-2 is to denoise the input sequence over a spectrum of noise levels, from complete noise to complete clean. By learning The model supports the simultaneous generation of highly correlated protein structures and sequences. This eliminates the need for a cascading generation paradigm, allowing both protein and sequence to be derived in a single step. DPLM-2 can be trained based on the training process of the diffusion model, from x (0) to x (T) Noise can be added to the feature representation step-by-step in the forward process from x (T) to x (0)The noise-free feature representation can be recovered step-by-step in the backward process. Further, the decoder 320 can be utilized to convert the feature representation to the protein structure 520. In this way, the DPLM-2 can perform the inference process using multiple modalities of the protein data.
[0064] According to some implementations of the present disclosure, in the process of determining the target model, an initial model can be obtained. The initial model is based on a discrete diffusion probability model and a language model structure, and the initial model is pre-trained using protein sequences of protein data in a protein data set; and the initial model is updated using at least any one of protein sequences of protein data in the protein data set and protein structures to determine the target model.
[0065] Regarding learning structure tokenization, tokenizing data into discrete representations has been proposed in various domains (e.g., image synthesis). The main motivation behind adopting discrete representations is that they can capture meaningful and compact information of the underlying data, and are used for efficient compression and efficient generation, especially using scalable sequence-based models like Transformers. In the context of the present disclosure, a discrete representation of protein structures, i.e., structure tokenization, can be obtained. This discrete representation allows language models to effectively learn the composition of local structural elements of protein structures.
[0066] In the protein structure tokenization method of the present disclosure, a lookup-free quantizer (LFQ) can be employed. LFQ operates by learning quantization boundaries directly from input data, presenting a novel paradigm for transforming continuous data (such as 3D coordinates of protein backbones) into discrete tokens. This significantly enhances the efficiency and effectiveness of tokenization in protein structure modeling. Specifically, given a set of 3D coordinates LFQ utilizes a learned encoder ε φ These coordinates are mapped to discrete indices s e {0, 1, …, |S|} within a codebook S. The quantization process can be represented as:
[0067] In the above formula, denotes a decoder that reconstructs the continuous representation from the discrete tokens.
[0068] According to some implementations of the present disclosure, a high-efficiency pre-training technical solution is proposed by utilizing a pre-trained DPLM. It should be understood that a large number of protein sequences found in nature encode key evolutionary information, reflecting a co-evolutionary process in which residue pairs mutate over time. These core evolutionary residue pairs usually interact in 3D space, thus providing valuable necessary structural information for predicting protein folding. Influenced by the link between evolutionary knowledge and structural interaction, transfer learning can be utilized to extract evolutionary information from a pre-trained protein language model, transforming a sequence-based model into a multi-modal model. A DPLM can be employed as a sequence language model for the model of the present disclosure.
[0069] According to some implementations of the present disclosure, a training process can be performed by utilizing a high-quality dataset including experimental structures and synthetic samples from protein data. The dataset is significantly smaller than the evolutionary scale dataset, thus enabling efficient fine-tuning of the pre-trained model. In order to preserve sequence knowledge and reduce the risk of catastrophic forgetting, a LoRA model can be used to constrain the modification to the original parameters. Models of different sizes (e.g., 150M, 650M, and 3B parameters) can be generated, and more than 100,000 training steps (batch size of 32,000) can be performed. Compared with training a model from scratch, this method not only reduces the training cost but also effectively transfers valuable evolutionary information.
[0070] According to some implementations of the present disclosure, the initial model herein can be a DPLM model, since the model already includes relevant knowledge of a single modality of protein data (only involving sequence features), supervised training can be performed on the basis of the single modality DPLM model, which can simplify the training complexity and generate a more accurate model. By utilizing some implementations of the present disclosure, a multi-modal target model can be supported to master the knowledge of performing different tasks, thus generating protein data that meets different needs.
[0071] More details about training are described with reference to FIG. 6, which shows a block diagram 600 of determining a target model based on an initial model according to some implementations of the present disclosure. As shown in FIG. 6, the target model 510 can be trained based on the initial model 640 by utilizing protein data 610. The protein data 610 includes multiple modalities: protein sequences 612 and protein structures 614. The target model 510 can be trained based on one or more of the multiple modalities.
[0072] According to some implementations of the present disclosure, in the process of updating the initial model to determine the target model, an update mode can be determined, the update mode comprising: updating the initial model with protein sequences, updating the initial model with protein structures, or updating the initial model with protein sequences and protein structures; and updating the initial model according to the update mode to determine the target model.
[0073] Specifically, to further enhance the ability of DPLM-2 to distinguish between two modalities s (structure) and a (sequence), each modality can have a different noise level, denoted as t aa and t ss This facilitates a more comprehensive understanding of the relationship between protein sequences and their corresponding structures. This design also allows for the exploration of arbitrary combinations of (t aa , t ss ), thereby providing flexible sampling options, including edge sampling from each modality and conditioning between them for various applications.
[0074] For example, update mode 1 can specify updating the initial model with protein sequences. At this time, the protein sequences of the protein data as training samples can be processed in a normal manner, thereby generating sequence features. The structure feature part can be set to empty, noise can be added to the structure feature, or smaller noise can be added to the sequence feature and larger noise can be added to the structure feature, so as to weaken the influence of the structure feature on the target model.
[0075] For another example, update mode 2 can specify updating the initial model with protein structures. At this time, the protein structures of the protein data as training samples can be processed in a normal manner, thereby generating structure features. The sequence feature part can be set to empty, noise can be added to the sequence feature, or smaller noise can be added to the structure feature and larger noise can be added to the sequence feature, so as to weaken the influence of the sequence feature on the target model.
[0076] For another example, update mode 3 can specify updating the initial model with protein sequences and protein structures. At this time, the protein sequences and protein structures of the protein data as training samples can be processed in a normal manner, thereby generating sequence features and structure features. Alternatively and / or additionally, noise can be added to the sequence features and structure features, and the like.
[0077] FIG. 7 shows a block diagram 700 of extracting feature representations of protein data with a target model according to some implementations of the present disclosure. For a given protein data 710, it can be considered as an initial state in the target model 510, i.e., t = 0. Then, the protein data 710 is input to the target model 510 in order to obtain the corresponding multi-modal feature representations 720.
[0078] According to some implementations of the present disclosure, based on the feature representation corresponding to the second target protein data, at least one of the following can be performed: a residue classification task for the second target protein data, a sequence classification task for the second target protein data, a sequence regression task for the second target protein data, a contact point prediction task for the second target protein data. Specifically, the multi-modal feature representation 720 corresponding to the extracted protein data 710 can be used to perform a classification task or a regression task at a sequence level or a residue level related to the protein data 710, including but not limited to a residue classification task, a sequence classification task, a sequence regression task, a contact point prediction task, etc. Since DPLM-2 can accurately understand and extract the feature representation of the protein sequence, this will significantly improve the result accuracy of the subsequent downstream prediction task.
[0079] According to some implementations of the present disclosure, the target model can also support unconditional generation tasks or conditional generation tasks. In some application scenarios, it is expected to generate protein sequences that meet certain conditions. Such conditional generation tasks include sequence-conditioned, expanded-modal, plug-and-play, preference-guided conditional generation.
[0080] According to some implementations of the present disclosure, the first feature representation includes a random feature representation, and wherein determining the first target protein data comprises: providing the random feature representation to the target model to determine the first target protein data using the target model. At this time, the random feature representation can be input to DPLM-2, and DPLM-2 will generate corresponding protein data, which can reflect the distribution of the training samples used in the training process.
[0081] Specifically, any one of the two modalities in the first feature representation can include a random feature. Assuming that the user specifies a first sequence feature (e.g., corresponding to a specified protein sequence) in the first feature representation, and the first structure feature includes a random feature, the protein structure in the output target protein data can correspond to the specified protein sequence. In this way, the folding process can be implemented using DPLM-2. Assuming that the user specifies a first structure feature (e.g., corresponding to a specified protein structure) in the first feature representation, and the first sequence feature includes a random feature, the protein sequence in the output target protein data can correspond to the specified protein structure. In this way, the folding process can be implemented using DPLM-2.
[0082] According to some implementations of the present disclosure, wherein the first sequence feature in the first feature representation comprises a plurality of positions, an element at at least one position of the plurality of positions represents at least one specified amino acid, and wherein determining the first target protein data comprises: providing the first feature representation to the target model to determine the first target protein data using the target model, an element in the first target protein data corresponding to the at least one position indicates the at least one specified amino acid. In this way, protein data comprising at least one specified amino acid can be generated in a more accurate manner.
[0083] FIG. 8 illustrates a block diagram 800 of a conditional generation task conditioned on a partial sequence, according to some embodiments of the present disclosure. In a conditional generation task conditioned on a partial sequence, an element of at least one position in the feature representation input to the DPLM-2 indicates at least one specified amino acid. In practical application scenarios, the conditional generation task conditioned on a partial sequence is suitable for generating a protein sequence with a given functional motif, filling an antibody complementarity determining region (CDR) loop, or a protein sequence with expert knowledge as a prior condition. In these application scenarios, the DPLM-2 can sample from the conditional distribution.
[0084] As shown in FIG. 8, in some application scenarios, known protein data 810 corresponds to a protein sequence 812, and it is desired to obtain a new protein sequence with the same amino acids at certain positions of the protein sequence 812. At this time, a protein sequence 816 can be constructed, the elements of which at the corresponding positions indicate the amino acids at the same positions of the protein sequence 812, and other positions (e.g., shown in block 814) are masked (as shown in FIG. 8, replaced with “X”). The protein sequence 816 is provided to the DPLM-2 as input. After processing by the DPLM-2, as shown by arrow 830, the amino acids at certain positions of the generated target protein sequence 822 correspond to the specified amino acids at the same positions of the protein sequence 812. Based on the protein sequence 822, corresponding new protein data 820 can be generated. At this time, as shown by arrow 840, the new protein data 820 has the same motif as the known protein data 810.
[0085] According to some implementations of the present disclosure, the first structure feature in the first feature representation comprises a plurality of positions, an element at at least one position of the plurality of positions represents at least one specified protein structure, and wherein determining the first target protein data comprises: providing the first feature representation to the target model to determine the first target protein data using the target model, an element in the first target protein data corresponding to the at least one position indicates the at least one specified protein structure. In this way, protein data comprising at least one specified protein structure can be generated in a more accurate manner.
[0086] Specifically, protein data with a certain protein structure (e.g., motif 818 in FIG. 8) can be specified to be generated. At this time, the protein structure at other positions other than the protein structure can be masked, and the masked protein structure can be input to the DPLM-2. At this time, the generated new protein data can have motif 828. With some implementations of the present disclosure, the DPLM has knowledge of multiple modalities, thus supporting the specification of generation conditions in multiple modalities respectively, thereby performing conditional generation tasks.
[0087] It should be understood that although only examples of specifying sequence conditions and specifying structure conditions are described in detail above, alternatively and / or additionally, sequence conditions and structure conditions can be specified respectively in a manner similar to that described above. In this way, more complex generation conditions can be specified, thereby enabling the generated protein data to have more rich characteristics.
[0088] The performance of the DPLM-2 can be evaluated in various generation and understanding scenarios, including unconditional protein generation (structure, sequence, and structure-sequence co-generation) and various conditional tasks such as folding, inverse folding and motif scaffolding, and a series of protein prediction tasks.
[0089] The purpose of unconditional protein generation is to generate the 3D structure and 1D sequence of a protein. In existing technical solutions, this can be achieved by a cascading method: first generate the structure, then use the inverse folding model to generate the sequence; or first sample the sequence, then the folding model to obtain the structure. However, existing technical solutions cannot generate structure and sequence simultaneously, and have low performance. In the context of the present disclosure, the DPLM-2 outperforms existing technical solutions in the tasks of unconditional structure generation, unconditional sequence generation, and structure-sequence co-generation. Experiments show that the DPLM-2 has better experimental results in terms of quality, novelty, and diversity.
[0090] Experimental data shows that the DPLM-2 can unconditionally generate highly designable and diverse structures. The DPLM-2 can perform structure-sequence co-generation in a flexible manner, where the structure and sequence have great consistency. The DPLM-2 can sample protein sequences with high structural rationality at all lengths.
[0091] It should be understood that protein sequences and protein structures are important indicators for characterizing proteins. With some implementations of the present disclosure, protein data can be described in a more accurate manner, thereby helping to improve the accuracy of downstream tasks.
[0092] Example process
[0093] FIG. 9 illustrates a flowchart of a method 900 of processing protein data with a multi-modal machine learning model, according to some implementations of the present disclosure. At block 910, a target model is obtained, the target model is based on a discrete diffusion probability model and a language model structure, and is pre-trained with a protein dataset. At block 920, the target model is utilized to perform at least one of: determining first target protein data with the target model based on a first feature representation corresponding to the protein data, the first feature representation including: first sequence features corresponding to a protein sequence of the first protein data, and first structure features corresponding to a protein structure of the first protein data; or extracting a second feature representation corresponding to second target protein data with the target model based on the second target protein data, the second feature representation including: second sequence features corresponding to a protein sequence of the second protein data, and second structure features corresponding to a protein structure of the second protein data.
[0094] According to some implementations of the present disclosure, the first feature representation includes a random feature representation, and wherein determining the first target protein data includes: providing the random feature representation to the target model to determine the first target protein data with the target model.
[0095] According to some implementations of the present disclosure, the first sequence features in the first feature representation include a plurality of positions, an element at at least one position of the plurality of positions represents at least one specified amino acid, and wherein determining the first target protein data includes: providing the first feature representation to the target model to determine the first target protein data with the target model, an element in the first target protein data corresponding to the at least one position indicates the at least one specified amino acid.
[0096] According to some implementations of the present disclosure, the first structure features in the first feature representation include a plurality of positions, an element at at least one position of the plurality of positions represents at least one specified protein structure, and wherein determining the first target protein data includes: providing the first feature representation to the target model to determine the first target protein data with the target model, an element in the first target protein data corresponding to the at least one position indicates the at least one specified protein structure.
[0097] According to some implementations of the present disclosure, the first feature representation corresponding to the protein data is determined based on: determining the first sequence features with a sequence encoder based on a protein sequence of the protein data; and determining the first structure features with a structure encoder based on a protein structure of the protein data, the structure encoder mapping continuous three-dimensional data of the protein structure to a discrete one-dimensional representation.
[0098] According to some implementations of the present disclosure, the first sequence feature includes a first plurality of positions, the first plurality of positions respectively correspond to a plurality of amino acids in the protein data, the first structure feature includes a second plurality of positions, the first plurality of positions is equal to the second plurality of positions, and the second plurality of positions respectively correspond to the plurality of amino acids.
[0099] According to some implementations of the present disclosure, the structure encoder is determined by: obtaining an initial structure encoder and an initial structure decoder, the initial structure decoder corresponds to the initial structure encoder; and updating the initial structure encoder to determine the structure encoder and updating the initial structure decoder to determine the structure decoder based on the protein structures in the protein structure dataset, the structure decoder corresponds to the structure encoder.
[0100] According to some implementations of the present disclosure, obtaining the target model includes: obtaining an initial model, the initial model is based on a discrete diffusion probability model and a language model structure, and the initial model is pre-trained by using protein sequences of protein data in a protein dataset; and updating the initial model to determine the target model by using at least any one of the protein sequences of the protein data in the protein dataset and the protein structures.
[0101] According to some implementations of the present disclosure, updating the initial model to determine the target model includes: determining an update mode, the update mode includes: updating the initial model by using the protein sequences, updating the initial model by using the protein structures, or updating the initial model by using the protein sequences and the protein structures; and updating the initial model to determine the target model according to the update mode.
[0102] According to some implementations of the present disclosure, the method further includes performing at least any one of the following based on the feature representation corresponding to the second target protein data: a residue classification task for the second target protein data, a sequence classification task for the second target protein data, a sequence regression task for the second target protein data, or a contact point prediction task for the second target protein data.
[0103] Example apparatus and devices
[0104] FIG. 10 shows a block diagram of an apparatus 1000 for processing protein data with a multi-modal machine learning model, according to some implementations of the present disclosure. The apparatus 1000 includes an obtaining module 1110 configured to obtain a target model, the target model being based on a discrete diffusion probability model and a language model structure, and pre-trained with a protein dataset; and an performing module 1120 configured to perform at least one of: determining, with the target model, first target protein data based on a first feature representation corresponding to the protein data, the first feature representation including first sequence features corresponding to protein sequences of the first protein data and first structure features corresponding to protein structures of the first protein data, or extracting, with the target model, a second feature representation corresponding to second target protein data based on the second target protein data, the second feature representation including second sequence features corresponding to protein sequences of the second protein data and second structure features corresponding to protein structures of the second protein data.
[0105] According to some implementations of the present disclosure, the first feature representation includes a random feature representation, and the performing module is further configured to provide the random feature representation to the target model to determine the first target protein data with the target model.
[0106] According to some implementations of the present disclosure, the first sequence features in the first feature representation include a plurality of positions, an element at at least one position of the plurality of positions representing at least one specified amino acid, and the performing module is further configured to provide the first feature representation to the target model to determine the first target protein data with the target model, an element in the first target protein data corresponding to the at least one position indicating the at least one specified amino acid.
[0107] According to some implementations of the present disclosure, the first structure features in the first feature representation include a plurality of positions, an element at at least one position of the plurality of positions representing at least one specified protein structure, and the performing module is further configured to provide the first feature representation to the target model to determine the first target protein data with the target model, an element in the first target protein data corresponding to the at least one position indicating the at least one specified protein structure.
[0108] According to some implementations of the present disclosure, the obtaining module is further configured to determine the first sequence features with a sequence encoder based on protein sequences of the protein data; and determine the first structure features with a structure encoder based on protein structures of the protein data, the structure encoder mapping continuous three-dimensional data of the protein structures to a discrete one-dimensional representation.
[0109] According to some implementations of the present disclosure, the first sequence feature includes a first plurality of positions, the first plurality of positions respectively correspond to a plurality of amino acids in the protein data, the first structure feature includes a second plurality of positions, the first plurality of positions is equal to the second plurality of positions, and the second plurality of positions respectively correspond to the plurality of amino acids.
[0110] According to some implementations of the present disclosure, the obtaining module is further configured to: obtain an initial structure encoder and an initial structure decoder, the initial structure decoder corresponds to the initial structure encoder; and update the initial structure encoder to determine the structure encoder and update the initial structure decoder to determine the structure decoder based on the protein structures in the protein structure dataset, the structure decoder corresponds to the structure encoder.
[0111] According to some implementations of the present disclosure, the obtaining module is further configured to: obtain an initial model, the initial model is based on a discrete diffusion probability model and a language model structure, and the initial model is pre-trained by using protein sequences of protein data in a protein dataset; and update the initial model to determine the target model by using at least any one of the protein sequences of the protein data and the protein structures in the protein dataset.
[0112] According to some implementations of the present disclosure, the obtaining module is further configured to: determine an update mode, the update mode includes: updating the initial model by using the protein sequences, updating the initial model by using the protein structures, or updating the initial model by using the protein sequences and the protein structures; and update the initial model to determine the target model according to the update mode.
[0113] According to some implementations of the present disclosure, the performing module is further configured to perform at least any one of the following based on the feature representation corresponding to the second target protein data: a residue classification task for the second target protein data, a sequence classification task for the second target protein data, a sequence regression task for the second target protein data, or a contact point prediction task for the second target protein data.
[0114] FIG. 11 shows a block diagram of a device 1100 that can implement a number of implementations of the present disclosure. It should be understood that the computing device 1100 illustrated in FIG. 11 is merely an example and should not be construed to limit the functionality and scope of the implementations described herein. The computing device 1100 illustrated in FIG. 11 can be used to implement the methods described above.
[0115] As shown in FIG. 11, computing device 1100 is in the form of a general-purpose computing device. Components of computing device 1100 can include, but are not limited to, one or more processors or processing units 1110, memory 1120, storage 1130, one or more communication units 1140, one or more input devices 1150, and one or more output devices 1160. Processing unit(s) 1110 can be actual or virtual processors and capable of executing various processing in accordance with programs stored in memory 1120. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of computing device 1100.
[0116] Computing device 1100 typically includes a plurality of computer storage media. Such media can be removable and / or non-removable, and can include volatile and / or non-volatile media. Memory 1120 can be volatile (such as, for example, registers, cache, random access memory (RAM)), non-volatile (such as, for example, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage 1130 can be removable or non-removable and can include machine-readable media, such as a flash drive, a magnetic disk drive, or any other media capable of storing information and / or data (e.g., training data for training) and accessible by computing device 1100.
[0117] Computing device 1100 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 11, a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non-removable, volatile, or non-volatile media. In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. Memory 1120 can include computer program product 1125 having one or more program modules configured to carry out the various methods or actions of the various implementations of the present disclosure.
[0118] Communication unit(s) 1140 enable communication over communication media to other computing devices. Additionally, the functionality of components of computing device 1100 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating over a communication connection. Thus, computing device 1100 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in a distributed environment.
[0119] Input device 1150 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. Output device 1160 can be one or more output devices, such as a display, a speaker, a printer, etc. Computing device 1100 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through communication unit 1140, and with one or more devices that enable a user to interact with computing device 1100, and / or any devices (e.g., a network card, a modem, etc.) that enable computing device 1100 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface (not shown).
[0120] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is provided having a computer program stored thereon, which when executed by a processor implements the method described above.
[0121] Various aspects of the disclosure can be described in the context of flow diagrams and / or block diagrams that illustrate the functions and / or acts performed by the methods, apparatus, devices, and computer program products according to the present disclosure. It is to be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer readable program instructions.
[0122] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flow diagrams and / or block diagrams. These computer readable program instructions can also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flow diagrams and / or block diagrams.
[0123] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0124] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0125] The above-described implementations of this disclosure are illustrative and not exhaustive, and are not limited to the disclosed implementations. Numerous modifications and adaptations will be apparent to those skilled in the art without departing from the scope and spirit of the disclosed implementations. The choice of words in this document is intended to best explain the principles of the implementations, practical application, or improvement over the technology in the market, or to enable other ordinary skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for processing protein data, comprising: obtaining a target model, the target model being based on a discrete diffusion probability model and a language model structure, and being pre-trained with a protein dataset; and utilizing the target model to perform at least one of: determining first target protein data based on a first feature representation corresponding to protein data, the first feature representation comprising: first sequence features corresponding to a protein sequence of the first protein data, and first structure features corresponding to a protein structure of the first protein data; or extracting, based on second target protein data, a second feature representation corresponding to the second target protein data using the target model, the second feature representation comprising: second sequence features corresponding to a protein sequence of the second protein data, and second structure features corresponding to a protein structure of the second protein data.
2. The method of claim 1, wherein the first feature representation comprises a random feature representation, and wherein determining the first target protein data comprises: providing the random feature representation to the target model to determine the first target protein data using the target model.
3. The method of claim 1, wherein the first sequence features in the first feature representation comprise a plurality of positions, an element at at least one of the plurality of positions representing at least one specified amino acid, and wherein determining the first target protein data comprises: providing the first feature representation to the target model to determine the first target protein data using the target model, an element in the first target protein data corresponding to the at least one position indicating the at least one specified amino acid.
4. The method of claim 1, wherein the first structure features in the first feature representation comprise a plurality of positions, an element at at least one of the plurality of positions representing at least one specified protein structure, and wherein determining the first target protein data comprises: providing the first feature representation to the target model to determine the first target protein data using the target model, an element in the first target protein data corresponding to the at least one position indicating the at least one specified protein structure.
5. The method of claim 1, wherein the first feature representation corresponding to the protein data is determined based on: determining the first sequence features based on a protein sequence of the protein data using a sequence encoder; and determining the first structure features based on a protein structure of the protein data using a structure encoder, the structure encoder mapping continuous three-dimensional data of the protein structure to a discrete one-dimensional representation. 6. The method of claim 5, wherein the first sequence feature comprises a first plurality of positions, the first plurality of positions respectively corresponding to a plurality of amino acids in the protein data, the first structure feature comprises a second plurality of positions, the first plurality of positions is equal to the second plurality of positions, and the second plurality of positions respectively correspond to the plurality of amino acids.
7. The method of claim 5, wherein the structure encoder is determined using: obtaining an initial structure encoder and an initial structure decoder, the initial structure decoder corresponding to the initial structure encoder; and updating the initial structure encoder based on protein structures in a protein structure dataset to determine the structure encoder, and updating the initial structure decoder based on the protein structures in the protein structure dataset to determine the structure decoder, the structure decoder corresponding to the structure encoder.
8. The method of claim 1, obtaining the target model comprises: obtaining an initial model, the initial model based on a discrete diffusion probability model and a language model structure, and the initial model pre-trained using protein sequences of protein data in the protein dataset; and updating the initial model using at least either one of protein sequences and protein structures of the protein data in the protein dataset to determine the target model.
9. The method of claim 8, wherein updating the initial model to determine the target model comprises: determining an update mode, the update mode comprising: updating the initial model using the protein sequences, updating the initial model using the protein structures, or updating the initial model using the protein sequences and the protein structures; and updating the initial model according to the update mode to determine the target model.
10. The method of claim 1, further comprising performing at least one of the following based on a corresponding feature representation of the second target protein data: a residue classification task for the second target protein data, a sequence classification task for the second target protein data, a sequence regression task for the second target protein data, or a contact prediction task for the second target protein data.
11. An apparatus for processing protein data using a multi-modal machine learning model, comprising: an obtaining module configured to obtain a target model, the target model based on a discrete diffusion probability model and a language model structure, and pre-trained using a protein dataset; and an executing module configured to perform at least one of the following using the target model: determining a first target protein data using the target model based on a first feature representation corresponding to the protein data, the first feature representation comprising: a first sequence feature corresponding to protein sequences of the first protein data, and a first structure feature corresponding to protein structures of the first protein data; or Based on the second target protein data, a second feature representation corresponding to the second target protein data is extracted using the target model, the second feature representation including: a second sequence feature corresponding to a protein sequence of the second protein data, and a second structure feature corresponding to a protein structure of the second protein data.
12. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions which, when executed by the at least one processing unit, cause the electronic device to carry out the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 10.
14. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 10.