A method for generating property-decoupled proteins based on a language model

By constructing an amino acid property knowledge map and training multiple language models, decoupling the amino acid properties, the problem that it is difficult for amino acid language models to model amino acid properties is solved, and the generation of functionally specific proteins and model flexibility is improved.

CN116013407BActive Publication Date: 2025-06-27HANGZHOU ZHIKE HUICHUANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211686617.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-06-27
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

The amino acid language model based on 20 independent amino acid symbols is difficult to effectively model the properties of amino acids, such as steric hindrance and hydrophilicity, resulting in increased model learning difficulty and limited functional specific protein production.

Method used

By constructing an amino acid property knowledge map, decoupling the amino acid properties and training multiple language models on the protein dataset, using the property representations output by these models to train samplers for specific amino acid families to generate proteins in different fields.

Benefits of technology

A finer-grained protein representation is achieved, which improves the model's ability to model amino acid properties, and can generate proteins with specific functions, which improves the flexibility and interpretability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116013407B_ABST
    Figure CN116013407B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating property-decoupled proteins based on a language model, including: constructing an amino acid property knowledge graph according to amino acid properties; obtaining protein data, decoupling each protein data into an amino acid property sequence according to the amino acid property knowledge graph, and mapping the amino acid property sequence from the property space to the vector space to obtain the vector representation of the amino acid property sequence; using the vector representation of the amino acid property sequence to model and train the language model for a causal relationship prediction task to optimize the parameters of the language model; generating proteins based on the language model with optimized parameters, and this method can generate specific proteins based on amino acid properties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of proteins, and particularly to a method for generating property-decoupled proteins based on a language model. Background Art

[0002] A protein is a sequence data composed of several amino acids. From this perspective, it has a certain similarity to natural language. Therefore, a large number of studies have migrated the methods used for natural language to protein sequences. The language model is the most concerned paradigm for language modeling at present. Its core idea is to use known sequences to obtain the probability distribution of unknown sequences. There are two most common language models: the masked language model and the causal language model. The masked language model predicts the probability distribution of the masked word at the masked position according to the context of the masked word, and it is very effective in understanding text. In the field of sequence generation, the causal language model dominates. It models the probability of the subsequent text using the previous text and generates text through continuous iteration. Nowadays, the modeling of natural language has changed from the statistics based on the co-occurrence frequency of words to the neural network fitting based on word vectors. Experiments have proved that the distributed features and non-linear mapping of neural networks have stronger generalization ability. GPT-3 expands the parameters to 175B and finds that at such a large scale, the model can generate sentences no less than those spoken by humans. Inspired by this, multiple research teams have tried to apply this paradigm to proteins and trained models such as ProGen2, ProtGPT2, and RITA. It is found that the larger the scale of the model parameters, the better the modeling effect on protein sequences, and the more natural-law-compliant proteins can be generated. It is expected that these models can obtain sequences different from those in nature but with expected functions through sampling.

[0003] However, the amino acid language model based on 20 mutually independent amino acid symbols cannot well model the properties of amino acids themselves, such as the steric hindrance of amino acids and the hydrophilicity of amino acids, etc., thus increasing the difficulty of model learning. Secondly, since the amino acid symbol embedding cannot decouple the properties, the probability of each amino acid appearing at the current position predicted is the superposition of the probabilities of each property appearing at that position. However, for proteins with different functions, the weight of probability superposition should be different. Therefore, the inability to decouple the properties will lead to the inability to generate proteins with specific functions directionally, restricting the flexibility in model use.

[0004] Current generation methods based on causal language models mainly focus on how to design samplers. Traditional maximization-based samplers aim to generate sequences that best match the model's expectations, including greedy generation and its improved method, beam search. However, maximization-based methods can lead to text degradation such as dullness, incoherence, or getting stuck in repetitive loops. To make the generated text more flexible, researchers have designed Top-k and Nucleus sampling methods. The main idea is to first select multiple candidate words and then select based on probability among the candidate words. However, it is noted that these sampling methods are all based on probability and cannot determine the meaning of the generated sequence, nor can they generate the desired sequence in a biased manner, resulting in a decrease in the practicality of the model. Summary of the Invention

[0005] In view of the above, the object of the present invention is to provide a method for generating proteins by decoupling based on the properties of a language model. By decoupling the amino acid property knowledge graph, multiple language models are trained on a protein dataset, and the property representations output by the multiple language models are used as inputs to train a sampler for a specific family on a specific amino acid family dataset to produce proteins in different fields.

[0006] To achieve the above object of the invention, an embodiment provides a method for generating proteins by decoupling based on the properties of a language model, including the following steps:

[0007] Construct an amino acid property knowledge graph according to amino acid properties;

[0008] Obtain protein data, decouple each protein data into an amino acid property sequence according to the amino acid property knowledge graph, and map the amino acid property sequence from the property space to the vector space to obtain the vector representation of the amino acid property sequence;

[0009] Use the vector representation of the amino acid property sequence to model and train the language model for a causal relationship prediction task to optimize the parameters of the language model;

[0010] Use the language model with optimized parameters to predict the probability distribution of the next amino acid property based on the vector representation of the known amino acid property sequence, use the sampler to predict the amino acid based on the probability distribution of the next amino acid property, and supplement the property of the predicted amino acid to the known amino acid property sequence. Repeat this step until the end, and convert the final amino acid property sequence into an amino acid sequence as the generated protein.

[0011] Preferably, in the amino acid property knowledge graph, each amino acid and its property are represented as a triple (amino acid, property strength, property category). An amino acid property knowledge graph is constructed according to the triple. In the property space, the property strength is represented by the modulus of the vector, and the property category is represented by the direction of the vector to obtain the embedding of the property in the amino acid property knowledge graph.

[0012] Preferably, each protein data is decoupled into an amino acid property sequence according to the amino acid property knowledge graph, including:

[0013] Find the corresponding amino acid property of each amino acid in each protein data in the amino acid property knowledge graph, and use the amino acid property to replace the amino acid to obtain the amino acid property sequence.

[0014] Preferably, the amino acid property sequence is mapped from the property space to the vector space, including:

[0015] According to the different amino acid properties, different mapping methods are adopted. For discrete properties, a dictionary-based embedding method is used; for continuous properties, the vector direction represents the property and the vector magnitude represents the magnitude of the property value; for properties represented in graph type, a graph neural network embedding method is used.

[0016] Preferably, the language model is a pluggable model capable of encoding sequences, including LSTM, Transformer, GPT3;

[0017] The language model is used to predict the probability distribution of the next amino acid property representation based on the vector representation of the known amino acid properties. In the language model, the top layer and the bottom layer independently encode different amino acid properties without sharing parameters. In the middle part except the top layer and the bottom layer, through sparse self-attention, the embeddings of different amino acid properties are shared for information, and through the information interaction between multiple amino acid properties, the prediction of a single amino acid property representation is enhanced by the known amino acid property information.

[0018] Preferably, when training the language model, a loss function is constructed by minimizing the error between the property label and the property prediction result, and the parameters of the language model are updated according to the loss function;

[0019] For discrete amino acid properties, the constructed loss function l b (p0: m ) is:

[0020]

[0021] Among them, b represents the batch number, y represents the property label, c represents a predicted property, C is the total number of predicted property categories, represents the probability that the model predicts the i-th amino acid type as c, represents the probability that the model predicts the i-th amino acid property type as y, p0: m represents the sum of the probabilities predicted by the model for the amino acid property sequence of length m.

[0022] For continuous amino acid properties, the constructed loss function is:

[0023]

[0024] Among them, m represents the total amount of amino acids, represents that the average value of amino acids is and the variance is the probability density function of the amino acid property of a normal distribution, x represents the input amino acid property sequence, μ and σ represent the average value and variance of amino acid properties, represents that the average value of amino acids is the probability density function of the amino acid property of a normal distribution with a variance of 1.

[0025] Preferably, the output of the language model is the predicted amino acid property representation, and then a single-layer linear network is used to map the predicted amino acid property representation to the property space, and the loss function is calculated based on the predicted amino acid properties.

[0026] Preferably, the sampler uses a neural network and a multi-layer perceptron.

[0027] Preferably, after obtaining the predicted amino acids according to the probability distribution of amino acid properties by the sampler, an amino acid property knowledge graph is used to determine the properties of the predicted amino acids, and the properties of the predicted amino acids are supplemented into the known amino acid property sequence.

[0028] Compared with the prior art, the beneficial effects of the present invention at least include:

[0029] (1) For the first time, an amino acid property knowledge graph is constructed based on the existing amino acid data, providing more fine-grained prior knowledge for protein representation.

[0030] (2) Currently, the probability of amino acid properties is predicted through multiple language models for the first time, and domain-specific proteins are generated through a domain sampler. Different from the existing generation models based on amino acid symbols, the generated amino acid sequences are specific and cannot reflect the properties of the amino acids required at the current position, and the generation space is limited to 20 natural amino acids. The language model in the present invention can not only describe the attributes of the amino acids required at the current position, but also has interpretability, and can also enable biologists to design new artificial amino acids through the properties described by the model, improving the diversity of biological materials. The sampler in the present invention takes multiple amino acid property signals as inputs and can better weigh and select what kind of amino acids are more in line with the requirements of the current position.

[0031] (3) The property embedding method of the present invention uses the means of knowledge graph enhancement to map continuous and discrete properties to a vector space for use by the language model.

[0032] (4) Different from existing single-property language modeling generation models, the present invention proposes to use a mixture of experts system under different amino acid properties, allowing limited communication among various properties to learn unique sequence patterns of various properties to guide generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0034] Figure 1 is a flowchart of the property decoupled protein generation method based on a language model provided by the embodiment;

[0035] Figure 2 is a schematic diagram of amino acid properties provided by the embodiment;

[0036] Figure 3 is a schematic diagram of the pre-training and fine-tuning processes of the language model provided by the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0038] Figure 1 is a flowchart of the property decoupled protein generation method based on a language model provided by the embodiment. As Figure 1 shown, the property decoupled protein generation method based on a language model provided by the embodiment includes the following steps:

[0039] Step 1, construct an amino acid property knowledge graph according to amino acid properties.

[0040] In the embodiment, based on experiments on amino acids in the chemical field, the physicochemical properties and importance degrees of amino acids that play important roles in protein function are obtained, and an amino acid property knowledge graph is constructed based on the physicochemical properties and importance degrees. This amino acid property knowledge graph is used as the basis for decoupling amino acids.

[0041] Among them, as Figure 2As shown in the figure, the properties of amino acids include: category, solubility, radius, charge, polarity, and molecular weight, etc. In the amino acid property knowledge graph, each amino acid and its properties are represented as a triple (amino acid, property strength, property category). Based on the triple, an amino acid property knowledge graph is constructed. In the property space, the property strength is represented by the modulus of the vector, and the property category is represented by the direction of the vector, obtaining the embedding of the property in the amino acid property knowledge graph.

[0042] Step 2: Obtain protein data and construct a pre-training dataset.

[0043] In the embodiment, the protein data comes from a protein sequence corpus, including all protein data measured from sequencing experiments on proteins in the biological field. Each protein data is an amino acid sequence composed of amino acids. In the embodiment, except for the protein data that cannot be used as a sample, that is, the protein data with an amino acid sequence length greater than 2048 is removed, and the remaining amino acid sequences with similar lengths and a total length not exceeding 2048 are used as samples.

[0044] Step 3: After decoupling the protein data based on the amino acid property knowledge graph, perform vector representation.

[0045] In the embodiment, each protein data is decoupled into an amino acid property sequence according to the amino acid property knowledge graph, and the amino acid property sequence is mapped from the property space to the vector space to obtain the vector representation of the amino acid property sequence.

[0046] Specifically, decoupling each protein data into an amino acid property sequence according to the amino acid property knowledge graph includes: finding the corresponding amino acid property of each amino acid in the protein data in the amino acid property knowledge graph, and using the amino acid property to replace the amino acid to obtain the amino acid property sequence.

[0047] For a given protein data, each amino acid in the protein data can be mapped by the amino acid property knowledge graph into three vectors related to solubility, radius, and polarity. In this way, a protein data can be mapped into three groups of amino acid property sequence vectors.

[0048] In the embodiment, in order to input to the language model, it is also necessary to map the amino acid property from the property space to a dense vector space. Specifically, according to the different amino acid properties, different mapping methods are adopted. For discrete properties, a dictionary-based embedding method is used; for continuous properties, the vector direction represents the property and the vector magnitude represents the magnitude of the property value; for properties represented in graph type, a graph neural network embedding method is used. In this way, a conversion processing method from the property space to the vector space is established through the property and the embedding method to obtain the vector representation of the amino acid property sequence.

[0049] Step 4: Use the vector representation of the amino acid property sequence to model and train the language model for the causal relationship prediction task to optimize the parameters of the language model.

[0050] In the embodiment, the language model is used to predict the probability distribution of the next amino acid property representation based on the vector representation of the known amino acid properties. It is necessary to exchange information between properties to obtain the property representation at the current position based on the known property information. The language model is a pluggable model capable of encoding sequences, including LSTM, Transformer, GPT3, etc. In the embodiment, GPT3 is adopted, and this GPT3 is a Transformer model with only a decoder. As Figure 3 shown, in the language model, the top layer and the bottom layer independently encode different amino acid properties without sharing parameters. In the middle part except for the top layer and the bottom layer, through sparse self-attention, the embeddings of different amino acid properties are shared in information, and through the information interaction between multiple amino acid properties, the single amino acid property representation is predicted through the enhancement of the known amino acid property information.

[0051] In the embodiment, based on different property expression methods, the amino acid property representation generated by the language model is mapped to the corresponding property space to calculate the loss function. Specifically, a single-layer linear network is used as the mapping head for both discrete properties and continuous properties to map the amino acid property representation to the corresponding property space to obtain the amino acid properties.

[0052] In the embodiment, when training the language model, the loss function is constructed by minimizing the error between the property label and the property prediction result, and the parameters of the language model are updated according to the loss function.

[0053] In the embodiment, for amino acid property causal language modeling, a training batch consists of N protein data. For each amino acid chain among them, only the property information of the previous positions can be obtained at the current position. The model predicts the amino acid properties at each position one by one to reduce the perplexity of the model for the entire amino acid property.

[0054] The causal language modeling loss function for the discrete amino acid properties of a protein is:

[0055]

[0056] where b represents the batch number, y represents the property label, c represents a predicted property, C is the total number of predicted property categories, represents the probability that the model predicts the i-th amino acid type as c, represents the probability that the model predicts the i-th amino acid property type as y, p 0:mrepresents the sum of the probabilities predicted by the representation model for the amino acid property sequence of length m, and the total loss is the sum of all prediction losses in a training batch.

[0057] The causal language modeling loss function for the consecutive amino acid properties of a protein is:

[0058]

[0059] where m represents the total amount of amino acids, represents the probability density function of the amino acid property of a normal distribution with the mean of amino acids being and the variance being x represents the input amino acid property sequence, μ and σ represent the mean and variance of the amino acid properties, represents the probability density function of the amino acid property of a normal distribution with the mean of amino acids being and the variance being 1. The total loss is the sum of all prediction losses in a training batch.

[0060] Step 5, generate proteins based on the parameter-optimized language model.

[0061] In downstream tasks, as Figure 3 shown, use domain-specific protein families as the fine-tuning dataset. Based on the fine-tuning dataset, use the parameter-optimized language model to generate amino acid property representations, and use the amino acid property representations and amino acids as samples to optimize the parameters of the domain-specific sampler. The sampler with optimized parameters can generate proteins with certain functions.

[0062] In the embodiment, the protein generation process includes: using the parameter-optimized language model to predict the probability distribution of the next amino acid property based on the vector representation of the known amino acid property sequence, using the sampler to predict amino acids based on the probability distribution of the next amino acid property, and supplementing the properties of the predicted amino acids to the known amino acid property sequence. Repeat this step until the end, and convert the final amino acid property sequence into an amino acid sequence as the generated protein. In this way, domain-specific proteins are generated by simultaneously using the language model and the amino acid sequence.

[0063] In the embodiment, the sampler uses a multi-layer perceptron and a neural network.

[0064] The above-described specific embodiments have elaborated on the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating property - decoupled proteins based on a language model, characterized in that, It includes the following steps: Construct an amino acid property knowledge graph based on amino acid properties; Obtain protein data, decouple each protein data into an amino acid property sequence according to the amino acid property knowledge graph, and map the amino acid property sequence from the property space to the vector space to obtain the vector representation of the amino acid property sequence; Use the vector representation of the amino acid property sequence to model and train the language model for the causal relationship prediction task to optimize the parameters of the language model; Use the language model with optimized parameters to predict the probability distribution of the next amino acid property based on the vector representation of the known amino acid property sequence, use the sampler to predict the amino acid based on the probability distribution of the next amino acid property, and supplement the properties of the predicted amino acid to the known amino acid property sequence. Repeat this step until the end, and convert the final amino acid property sequence into an amino acid sequence as the generated protein.

2. The method for generating proteins by decoupling properties based on a language model according to claim 1, wherein In the amino acid property knowledge graph, each amino acid and its properties are represented as a triple (amino acid, property strength, property category). Construct the amino acid property knowledge graph according to the triple. In the property space, the property strength is represented by the modulus of the vector, and the property category is represented by the direction of the vector to obtain the embedding of the property in the amino acid property knowledge graph.

3. The method for generating proteins by decoupling properties based on a language model according to claim 1, characterized in that, Decouple each protein data into an amino acid property sequence according to the amino acid property knowledge graph, including: Find the corresponding amino acid property of each amino acid in each protein data in the amino acid property knowledge graph, and use the amino acid property to replace the amino acid to obtain the amino acid property sequence.

4. The method for generating proteins by decoupling properties based on a language model according to claim 1, wherein The mapping of the amino acid property sequence from the property space to the vector space includes: According to the different amino acid properties, different mapping methods are adopted. For discrete properties, the dictionary embedding method is adopted; for continuous properties, the vector direction represents the property and the vector size represents the size of the property value; for properties represented in graph type, the graph neural network embedding method is adopted.

5. The method for generating proteins by decoupling properties based on a language model according to claim 1, characterized in that, The language model is a pluggable model capable of encoding sequences, including LSTM, Transformer, GPT3; The language model is used to predict the probability distribution of the next amino acid property representation based on the vector representation of the known amino acid property. In the language model, the top layer and the bottom layer independently encode different amino acid properties without sharing parameters. In the middle part except the top layer and the bottom layer, through sparse self-attention, the embeddings of different amino acid properties are shared for information, and through the information interaction between multiple amino acid properties, the prediction of a single amino acid property representation is enhanced by the known amino acid property information.

6. The method for generating proteins by decoupling properties based on a language model according to claim 1, wherein, When training the language model, a loss function is constructed by minimizing the error between the property label and the property prediction result, and the parameters of the language model are updated according to the loss function; The loss function l constructed for discrete amino acid properties b (p 0:m ) is as follows: Among them, b represents the batch serial number, y represents the property label, c represents a predicted property, and C is the total amount of predicted property categories. represents the probability that the model predicts that the i-th amino acid type is c. represents the probability that the model predicts that the i-th amino acid property type is y, p 0:m represents the sum of the probabilities predicted by the model for the amino acid property sequence with length m. For continuous amino acid properties, the constructed loss function is: where m represents the total amount of amino acids, represents the probability density function of the amino acid property of the normal distribution with the mean value of amino acids being and the variance being For the amino acid property sequence x input, μ and σ represent the mean value and variance of amino acid properties. represents the probability density function of the amino acid property of the normal distribution with the mean value of amino acids being and the variance being 1.

7. The method for generating proteins by decoupling properties based on a language model according to claim 6, characterized in that, The output of the language model is the predicted amino acid property representation, and then a single-layer linear network is used to map the predicted amino acid property representation to the property space, and the loss function is calculated based on the predicted amino acid property.

8. The method for generating a protein by decoupling properties based on a language model according to claim 1, characterized in that, The sampler uses a neural network and a multi-layer perceptron.

9. The method for generating a protein by decoupling properties based on a language model according to claim 1, wherein After using a sampler to obtain a predicted amino acid based on the probability distribution of amino acid properties, an amino acid property knowledge graph is used to determine the properties of the predicted amino acid, and the properties of the predicted amino acid are supplemented into the known amino acid property sequence.