A method and device for training a capsid protein sequence generation model

By training a capsid protein sequence generation model through data augmentation and noise addition, the problem of low capsid sequence transduction efficiency in existing technologies has been solved, enabling the efficient generation of a large number of capsid sequences with specific functions and improving the efficiency and accuracy of sequence design.

CN116525011BActive Publication Date: 2026-04-28HANGZHOU CARBON SILICON SMART TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU CARBON SILICON SMART TECH DEV CO LTD
Filing Date
2023-03-27
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing technologies for designing adeno-associated virus capsid sequences suffer from low transduction efficiency and difficulty in efficiently screening capsid sequences with specific functions. Random mutation methods are inefficient and rely on the performance of binary classifiers, which may lead to the omission of potential functional sequences.

Method used

By acquiring training data, performing data augmentation and noise addition, and using specified generation paths and loss functions to train a capsid protein sequence generation model, a large number of capsid protein sequences with specified functions are generated, including discrete and continuous noise addition methods. The model training is then optimized by combining different loss functions.

Benefits of technology

It increases the proportion of sequences with specific functions in the generated capsid sequences, increases the number of sequences available in the same amount of time, and improves the efficiency and accuracy of capsid sequence design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116525011B_ABST
    Figure CN116525011B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a training method and device of a capsid protein sequence generation model. The method comprises: obtaining training data, the training data being a sequence set having a specified capsid function in a selected region of a protein molecule; performing data enhancement processing on the training data to obtain model training data; performing noise addition processing on the model training data at different times based on a set noise addition processing mode to obtain noise-added training data at different times; training a to-be-trained capsid protein sequence generation model according to the noise-added training data and a loss function under a specified generation path to obtain a capsid protein sequence generation model of the specified capsid function under the specified generation path. Embodiments of the present application can generate a large number of capsid protein sequences required by the function of the capsid, and the generated capsid protein sequences have a high sequence proportion, and the number of available sequences found is much larger than the number of sequences found by random mutation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a training method and apparatus for a capsid protein sequence generation model. Background Technology

[0002] AAV (adeno-associated virus) capsids have become a powerful tool for therapeutic in vivo gene delivery. However, the transduction efficiency of natural capsids still limits therapeutic applications, and the engineering of enhanced capsids has proven challenging due to the complexity of genotype-phenotype relationships and the many functional properties that must be optimized simultaneously.

[0003] Currently, the physical interactions that determine protein function are not well understood, so most existing methods rely on directed evolution. When the understanding of the mechanisms is limited, repeatedly applying random mutations and artificial selection is often the default engineering strategy. However, capsids designed in this way are unusable in a high proportion, resulting in low production efficiency. A more recent approach involves first training a binary classifier using a small amount of data with a specific function (in this method, the function of the capsid is determined by the activity of the virus corresponding to the designed capsid sequence) to determine whether the sequence is active or inactive. Then, a randomly divided mutation subspace is created (the subspace construction is illustrated in the diagram). Figure 1 Random sampling is performed within the range shown. If a binary classifier determines that a sample is active, it is retained; otherwise, it is deleted. Finally, through continuous iterative filtering, a set of capsid sequences that are expected to be active is selected. The capsid sequence set constructed in this way has a higher proportion of active sequences than the capsid sequence set constructed through random mutation. However, this proportion of active sequences is highly dependent on the performance of the previously trained binary classifier. If the performance of the binary classifier is poor, the proportion of active sequences in the final selected set will be lower. Furthermore, because the number of mutation combinations is very high—even without considering insertions—the number of combinations reaches 20. seq_len Where seq_len is the sequence length. This order of magnitude of data cannot be filtered within a predictable timeframe. Therefore, in practice, a subspace must be randomly partitioned from this size of space before further filtering. Since the proportion of sequences with a specific function in the entire sequence space is extremely low, potential sequences with a specific function may be missed when partitioning the subspace. Summary of the Invention

[0004] The technical problem to be solved by the embodiments of this application is to provide a training method and apparatus for a capsid protein sequence generation model, so as to achieve the goal of generating a large number of capsid protein sequences with the required capsid function, with a high proportion of usable capsid protein sequences, and at the same time, the number of usable sequences found is much larger than the number of sequences found by random mutation.

[0005] In a first aspect, embodiments of this application provide a method for training a capsid protein sequence generation model, the method comprising:

[0006] Acquire training data, which is a set of sequences that have a specified capsid function in a selected region of a protein molecule;

[0007] The training data is augmented to obtain model training data;

[0008] The training data of the model is subjected to noise processing at different times based on the set noise processing method, so as to obtain noisy training data at different times.

[0009] The capsid protein sequence generation model to be trained is trained based on the noisy training data and the loss function under the specified generation path to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path.

[0010] Optionally, the step of performing data augmentation on the training data to obtain model training data includes:

[0011] Add a set character at any position in the sequence corresponding to each training data to obtain model training data with the same sequence length;

[0012] The set character is a character without meaning.

[0013] Optionally, after training the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path, the method further includes:

[0014] Obtain the protein sequence with complete noise, along with the specified model function;

[0015] Obtain the target capsid protein sequence generation model corresponding to the specified model function;

[0016] The completely noisy protein sequence is input into the target capsid protein sequence generation model, so that the target capsid protein sequence generation model processes the completely noisy protein sequence according to the specified generation path to obtain the denoised predicted capsid protein sequence corresponding to the completely noisy protein sequence.

[0017] The predicted capsid protein sequence is post-processed to obtain the final capsid protein sequence.

[0018] Optionally, the post-processing of the predicted capsid protein sequence to obtain the final capsid protein sequence includes:

[0019] The predetermined character contained in the predicted capsid protein sequence was detected;

[0020] Replace the specified character with an empty character to generate the final capsid protein sequence.

[0021] Optionally, the noise addition processing method includes either a discrete noise addition method or a continuous noise addition method.

[0022] Optionally, when the noise processing method is set to discrete noise processing,

[0023] The step of adding noise to the model training data at different times based on a set noise-adding method to obtain noisy training data at different times includes:

[0024] The model training data is noise-added based on the predefined transition probability matrix at different times to obtain the noise-added training data at different times corresponding to the model training data.

[0025] Optionally, when the noise processing method is set to continuous noise processing,

[0026] The step of adding noise to the model training data at different times based on a set noise-adding method to obtain noisy training data at different times includes:

[0027] The model training data is noise-added based on the predefined noise weight at different times and the latent vector of each amino acid in the sequence set under the specified capsid function, to obtain the noise-added training data corresponding to the model training data at different times.

[0028] Optionally, when the specified generation path is a predicted noise path,

[0029] The step of training the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path, to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path, includes:

[0030] Obtain the first loss function corresponding to the predicted noise path;

[0031] The capsid protein sequence generation model is trained based on the noisy training data and the first loss function to obtain the capsid protein sequence generation model with the specified capsid function under the predicted noise path.

[0032] Optionally, when the specified generation path is the path for predicting a noisy sequence,

[0033] The step of training the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path, to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path, includes:

[0034] Based on the set noise processing method and the predicted path of the un-noiseed sequence, a second loss function is determined;

[0035] The capsid protein sequence generation model is trained based on the noisy training data and the second loss function to obtain the capsid protein sequence generation model with the specified capsid function under the path of the predicted unnoisy sequence.

[0036] Optionally, when the specified generation path is a path that predicts the probability of the noisy sequence at the previous time step,

[0037] The step of training the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path, to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path, includes:

[0038] Based on the set noise processing method and the path of predicting the probability of the noise-added sequence at the previous time step, the third loss function is determined;

[0039] The capsid protein sequence generation model is trained based on the noisy training data and the third loss function to obtain the capsid protein sequence generation model under the path of the probability of the noisy sequence at the previous prediction time for the specified capsid function.

[0040] Secondly, embodiments of this application provide a training apparatus for a capsid protein sequence generation model, the apparatus comprising:

[0041] The training data acquisition module is used to acquire training data, which is a set of sequences that have a specified capsid function in a selected region of a protein molecule.

[0042] The model data acquisition module is used to perform data augmentation processing on the training data to obtain model training data;

[0043] The noise-added data acquisition module is used to add noise to the model training data at different times based on a set noise-adding processing method, so as to obtain noise-added training data at different times.

[0044] The model acquisition module is used to train the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path, so as to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path.

[0045] Optionally, the model data acquisition module includes:

[0046] The model data acquisition unit is used to add a set character at any position in the sequence corresponding to each training data to obtain the model training data with the same sequence length.

[0047] The set character is a character without meaning.

[0048] Optionally, the device further includes:

[0049] The protein sequence acquisition module is used to acquire completely noisy protein sequences, along with specified model functions;

[0050] The target generation model acquisition module is used to acquire the target capsid protein sequence generation model corresponding to the specified model function.

[0051] The predicted protein sequence acquisition module is used to input the completely noisy protein sequence into the target capsid protein sequence generation model, so that the target capsid protein sequence generation model processes the completely noisy protein sequence according to the specified generation path to obtain the denoised predicted capsid protein sequence corresponding to the completely noisy protein sequence.

[0052] The capsid protein sequence acquisition module is used to post-process the predicted capsid protein sequence to obtain the final capsid protein sequence.

[0053] Optionally, the capsid protein sequence acquisition module includes:

[0054] A character detection unit is set up to detect the set character contained in the predicted capsid protein sequence;

[0055] The capsid protein sequence generation unit is used to replace the set character with an empty character to generate the final capsid protein sequence.

[0056] Optionally, the noise addition processing method includes either a discrete noise addition method or a continuous noise addition method.

[0057] Optionally, when the noise processing method is set to discrete noise processing,

[0058] The noisy data acquisition module includes:

[0059] The first noisy data acquisition unit is used to add noise to the model training data according to the predefined transition probability matrix at different times, so as to obtain the noisy training data corresponding to the model training data at different times.

[0060] Optionally, when the noise processing method is set to continuous noise processing,

[0061] The noisy data acquisition module includes:

[0062] The second noise-added data acquisition unit is used to add noise to the model training data according to the noise ratio at different times and the latent vector of each amino acid in the sequence set under the specified capsid function, so as to obtain the noise-added training data corresponding to the model training data at different times.

[0063] Optionally, when the specified generation path is a predicted noise path,

[0064] The generative model acquisition module includes:

[0065] The first loss function acquisition unit is used to acquire the first loss function corresponding to the predicted noise path;

[0066] The first generative model acquisition unit is used to train the capsid protein sequence generation model to be trained based on the noisy training data and the first loss function, so as to obtain the capsid protein sequence generation model with the specified capsid function under the predicted noise path.

[0067] Optionally, when the specified generation path is the path for predicting a noisy sequence,

[0068] The generative model acquisition module includes:

[0069] The second loss function acquisition unit is used to determine the second loss function based on the set noise processing method and the predicted path of the unnoised sequence.

[0070] The second generative model acquisition unit is used to train the capsid protein sequence generation model to be trained based on the noisy training data and the second loss function, so as to obtain the capsid protein sequence generation model with the specified capsid function under the path of the predicted unnoisy sequence.

[0071] Optionally, when the specified generation path is a path that predicts the probability of the noisy sequence at the previous time step,

[0072] The generative model acquisition module includes:

[0073] The third loss function acquisition unit is used to determine the third loss function based on the set noise processing method and the path of predicting the probability of the noise sequence at the previous time step.

[0074] The third generative model acquisition unit is used to train the capsid protein sequence generation model to be trained based on the noisy training data and the third loss function, so as to obtain the capsid protein sequence generation model under the path of the probability of the noisy sequence at the previous prediction time for the specified capsid function.

[0075] Thirdly, embodiments of this application provide an electronic device, including:

[0076] A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the training method for the capsid protein sequence generation model described in any of the preceding claims.

[0077] Fourthly, embodiments of this application provide a computer-readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the training method for the capsid protein sequence generation model described in any of the preceding claims.

[0078] Compared with the prior art, the embodiments of this application have the following advantages:

[0079] In this embodiment, training data is acquired, which is a set of sequences with a specified capsid function within a selected region of a protein molecule. Data augmentation is performed on the training data to obtain model training data. Noise is added to the model training data at different times based on a set noise-adding method, resulting in noisy training data at different times. The capsid protein sequence generation model is trained based on the noisy training data and a loss function under a specified generation path, resulting in a capsid protein sequence generation model with the specified capsid function under the specified generation path. This embodiment constructs a model capable of generating a large number of desired capsid functions based on a small amount of existing data (functions may include, but are not limited to, active capsids, deimmunogenic capsids, or capsids targeting specific cells), thereby achieving the goal of designing capsid sequences. The proportion of usable sequences generated is higher than in existing methods. Furthermore, since the model provided in this embodiment is a capsid sequence generator that can directly generate capsid sequences containing the required functions of the capsid, it is expected that in the same amount of time, the number of usable sequences found by this application (searching along the direction of the capsid with the function) will be much larger than the number of sequences found by random mutation (searching in the entire domain).

[0080] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0081] Figure 1 A schematic diagram illustrating a subspace construction method provided in an embodiment of this application;

[0082] Figure 2 A flowchart illustrating the steps of a training method for a capsid protein sequence generation model provided in this application embodiment;

[0083] Figure 3 This is a schematic diagram showing the appearance of a folded capsid protein sequence provided in an embodiment of this application;

[0084] Figure 4 A schematic diagram of the genome structure of an AAV vector provided in this application embodiment;

[0085] Figure 5 A schematic diagram illustrating a process for generating a diffusion model, as provided in an embodiment of this application;

[0086] Figure 6 A schematic diagram of the structure of a training device for a capsid protein sequence generation model provided in an embodiment of this application;

[0087] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0088] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0089] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0090] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes said element.

[0091] First, it can be combined Figure 3 The capsid protein sequence is described as follows.

[0092] Reference Figure 3 This illustration shows a schematic diagram of the appearance of a folded capsid protein sequence provided in an embodiment of this application. Figure 3 As shown, the capsid of adeno-associated virus (AAV) is a T=1 icosahedron (an icosahedron is a convex polyhedron formed by 20 triangles, with every 5 triangles forming a five-fold vertex, and a five-fold rotation axis passing through each pair of opposing five-fold vertices. There is a three-fold rotation axis passing through the center of each pair of opposing triangles; and a two-fold rotation axis passing through the midpoint of each pair of opposing edges). It is assembled from 60 VP monomers through the interaction of these rotation axes. These VP monomers can be all VP3, or they can be composed of VP1, VP2, and VP3 together. Figure 4 The genome structure of an AAV vector is shown, such as Figure 4 As shown in the figure, the sequence relationships of the monomers VP1, VP2, and VP3 are as follows: VP3 is entirely contained within VP2 and VP1; therefore, VP3 is the common region of the three VPs, also known as the VP3 common region. VP2 is approximately 57 amino acids longer than VP3, and this sequence is contained within VP1; therefore, it is also called the VP1 / VP2 common region. Only the N-terminus of VP1 has a unique sequence of approximately 138 amino acids; therefore, it is also called the VP1 unique region (VP1u). Previous studies have shown that the active mutation regions of VP monomers are generally located in the overlapping VP3 regions of different monomers. Therefore, potential mutation regions are selected within the VP3 region, and then relevant methods are used to design sequences for these mutation regions.

[0093] The capsid protein sequence generation model provided in this application embodiment can be a diffusion model. The implementation process of the capsid protein sequence generation model can be as follows: Figure 5As shown, the generative diffusion model consists of a diffusion process and a denoising process. The denoising process can be understood as a prediction process. Given a noise sequence of length [length of the sequence], the denoising model trained during the diffusion process continuously denoises the noise sequence. After time steps T, the noise sequence will be restored to a meaningful sequence under a given function. The diffusion process is a noise addition process. For a single sequence x0 under a given function, we add noise step by step to obtain x1, x2, ..., finally obtaining a completely noisy x. T The significance of the diffusion process is to help neural networks learn the denoising process. Thinking more closely, the noisy sequence obtained during the diffusion process is actually the generated label, because in this process, the real sequence already exists, and the noisy result of the real sequence is generated; therefore, it is possible for the network to learn such a mapping, that is, to recover the original real sequence from the noisy sequence.

[0094] The training and inference processes of the capsid protein sequence generation model will be described in detail below with reference to specific embodiments.

[0095] Reference Figure 2 The diagram illustrates a flowchart of the training method for a capsid protein sequence generation model provided in this application embodiment. Figure 2 As shown, the training method for this capsid protein sequence generation model may include the following steps:

[0096] Step 201: Obtain training data, which is a set of sequences that have a specified capsid function in a selected region of a protein molecule.

[0097] The embodiments of this application can be applied to scenarios where capsid protein sequence generation models with different capsid functions are trained.

[0098] Training data refers to the data used to train the capsid protein sequence generation model. In this embodiment, the training data can be a set of sequences that have a specified capsid function in a selected region of a protein molecule.

[0099] This application embodiment uses a set of sequences with a specified capsid function in a selected region of a protein molecule as model training data, thereby enabling the trained model to sample along the sample direction of a given function in the sequence space. Therefore, this generation method may be able to sample almost all potential samples with a given function within a foreseeable time.

[0100] When training a capsid protein sequence generation model, the corresponding training data can be obtained.

[0101] After obtaining the training data, proceed to step 202.

[0102] Step 202: Perform data augmentation on the training data to obtain model training data.

[0103] After obtaining the training data, data augmentation can be performed on the training data to obtain model training data, which can then be applied to the subsequent model training process.

[0104] The specific process of data augmentation for training data can be as follows: Add a specified character at any position in the sequence corresponding to each training data point to obtain model training data with consistent sequence lengths. The specified character can be a character without any meaning.

[0105] The embodiments of this application achieve the following technical effects by adding specified characters to the sequence of training data:

[0106] 1. It can ensure that the length of the final sequence samples remains consistent;

[0107] 2. The meaning of the character can be the deletion of an amino acid at that site. Therefore, this insertion method does not change the meaning of the original protein sequence, but increases the diversity of the sample and can achieve the effect of data augmentation.

[0108] After performing data augmentation on the training data to obtain the model training data, proceed to step 203.

[0109] Step 203: Based on the set noise processing method, perform noise processing on the model training data at different times to obtain noisy training data at different times.

[0110] After obtaining the model training data, noise can be added to the model training data at different times based on the set noise addition method to obtain noisy training data at different times.

[0111] In this example, the noise addition method can be either discrete noise addition or continuous noise addition.

[0112] Next, the noise addition process will be described in detail below, combining the two noise addition methods.

[0113] I. When the noise addition method is set to discrete noise addition, the process of adding noise to the model training data can be as follows: Noise is added to the model training data according to the predefined transition probability matrix at different times to obtain the noisy training data at different times. The specific implementation process is as follows:

[0114] When adding discrete noise to the training data, the transition probability matrix Q at different time points is determined in advance. t |Q t |mn =q(x t =m|x t-1 =n) means the probability that the amino acid type is n at the current time and changes to m at the next time, where Q t The schematic matrix is ​​shown below. The matrix content can be arbitrarily set, as long as the sum of the same column is 1. Then, according to Q... t x0 obtains the noisy data x at any given time. t .

[0115]

[0116]

[0117] Here, v(x) represents a one-hot vector that is 1 in state x and 0 in all other states.

[0118] II. When the noise addition method is set to continuous noise addition, the process of adding noise to the model training data can be as follows: Noise is added to the training data based on the predefined noise weight at different times and the latent vector of each amino acid in the sequence set under the specified capsid function, resulting in noisy training data at different times corresponding to the model training data. The specific implementation process is as follows:

[0119] When adding continuous noise to the model training data, the proportion α of normally distributed noise introduced at different times can be determined in advance. t ,β t =1-α t And by obtaining the latent vector x0 for each amino acid in the sequence, the noisy data x at any given time can be obtained from x0. t .

[0120]

[0121] Where N(x; μ,σ) 2 ) represents a mean of μ and a variance of σ. 2 The normal distribution

[0122] Understandably, the above two noise-adding processing methods are merely examples listed to better understand the technical solutions of the embodiments of this application. In specific implementations, the noise-adding processing method may include, but is not limited to, the above two methods. Specifically, the specific method for adding noise to the model training data can be determined according to business needs, and this embodiment does not impose any restrictions on it.

[0123] After adding noise to the model training data at different times based on the set noise addition method, and obtaining noisy training data at different times, step 204 is executed.

[0124] Step 204: Train the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path.

[0125] After adding noise to the model training data at different times based on the set noise addition method, and obtaining noisy training data at different times, the capsid protein sequence generation model to be trained can be trained according to the noisy training data and the loss function under the specified generation path, so as to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path.

[0126] In this implementation, the specified generation path may include any one of the following: a path for predicting noise, a path for predicting the unnoised sequence, and a path for predicting the probability of the noisy sequence at the previous time step.

[0127] Next, the model training process will be described in detail below, combining the three paths mentioned above.

[0128] First, when the specified generation path is a predicted noise path, the specific implementation process of model training can be as follows: obtain the first loss function corresponding to the predicted noise path, and then train the capsid protein sequence generation model to be trained based on the noisy training data and the first loss function to obtain a capsid protein sequence generation model with a specified capsid function under the predicted noise path. The specific implementation process is as follows:

[0129] When the model's specified generation path is predicted noise, the loss function is as follows, where ε0 is the initial noise added during the diffusion process. For the model based on x t The model can be trained by minimizing the initial noise predicted by t.

[0130]

[0131] Second, when the specified generation path is the path of the predicted un-noised sequence, the specific implementation process of model training can be as follows: Determine the second loss function based on the set noise processing method and the path of the predicted un-noised sequence. Train the capsid protein sequence generation model to be trained based on the noisy training data and the second loss function to obtain the capsid protein sequence generation model with the specified capsid function under the path of the predicted un-noised sequence. The specific implementation process is as follows:

[0132] The specified generation path is direct prediction p(x0|x t (indicates that according to x) t Predict x0, x tWhen x0 is the predicted sequence (i.e., the path of the un-noised sequence), the loss function is as follows:

[0133] 1. When the noise addition method is set to discrete noise addition, the loss function is as follows. This loss function is also the baseline loss function for the generation diffusion model. Other different loss function forms are gradually derived from the following formula. Where, as t approaches infinity, q(x) T |x0) approaches a stationary distribution, so for L T The item can be ignored, then minimized. This allows for model training and enables the model to predict p. θ (x0|x t Substitute p θ (x t-1 |x t ), then p θ (x t-1 |x t ), q(x) t-1 |x t Substitute x0) This allows for the calculation of the loss function.

[0134]

[0135]

[0136] 2. When the noise addition method is set to continuous noise addition, the loss function is as follows, where To train the model, minimize the following formula to obtain the predicted values:

[0137]

[0138] Third, when the specified generation path is the path that predicts the probability of the noisy sequence at the previous time step, the specific implementation process of model training can be as follows: Determine the third loss function based on the set noise processing method and the path that predicts the probability of the noisy sequence at the previous time step. Train the capsid protein sequence generation model to be trained based on the noisy training data and the third loss function to obtain the capsid protein sequence generation model with the specified capsid function under the path that predicts the probability of the noisy sequence at the previous time step. The specific implementation process is as follows:

[0139] When the specified generation path is direct prediction p(x) t-1 |x t When ), the loss function is as follows:

[0140] 1. When the noise processing method is set to discrete noise addition, the loss function is as follows, where p θ (xt-1 |x t ), q(x) t-1 |x t Substitute x0) This allows for loss function calculation:

[0141]

[0142]

[0143] 2. When the noise processing method is set to continuous noise addition, the loss function is as follows, where μ θ Let μ be the mean value predicted by the model at time t-1. q This is the mean value at the actual time t-1.

[0144]

[0145] Understandably, the first, second, and third loss functions in the above process are merely qualifiers "first," "second," and "third" added to distinguish different generation paths, and have no practical significance.

[0146] Through the above-described method, this application embodiment can obtain a model that can generate a large number of required caption functions, thereby achieving the purpose of designing caption sequences.

[0147] The model reasoning process will now be described in detail with reference to specific embodiments.

[0148] In one specific implementation of this application, after step 204 above, the following may also be included:

[0149] Step S1: Obtain the protein sequence with complete noise, and the specified model function.

[0150] In this embodiment, after the capsid protein sequence generation model is obtained through the above steps 201 to 204, the capsid protein sequence generation model can be applied to the design of capsids with arbitrary functions. As long as there is data on the capsid with this function, the method can be used for training, and after training, a capsid with this function can be generated.

[0151] During model inference, a completely noisy protein sequence and a specified model function can be obtained. In this example, the model function may include, but is not limited to, the ability to bind to target cells.

[0152] After obtaining the completely noisy protein sequence and the specified model function, proceed to step S2.

[0153] Step S2: Obtain the target capsid protein sequence generation model corresponding to the specified model function.

[0154] After obtaining a completely noisy protein sequence and a specified model function, a target capsid protein sequence generation model corresponding to the specified model function can be obtained. In specific implementations, different capsid functions correspond to different capsid protein sequence generation models. After the user specifies the model function, the corresponding capsid protein sequence can be directly generated using the model corresponding to the user-specified model function.

[0155] After obtaining the target capsid protein sequence corresponding to the specified model function to generate the model, proceed to step S3.

[0156] Step S3: Input the completely noisy protein sequence into the target capsid protein sequence generation model, so that the target capsid protein sequence generation model processes the completely noisy protein sequence according to the specified generation path to obtain the denoised predicted capsid protein sequence corresponding to the completely noisy protein sequence.

[0157] After obtaining the target capsid protein sequence generation model corresponding to the specified model function, the completely noisy protein sequence can be input into the target capsid protein sequence generation model. The target capsid protein sequence generation model will process the completely noisy protein sequence according to the specified generation path to obtain the denoised predicted capsid protein sequence corresponding to the completely noisy protein sequence.

[0158] The model reasoning process can be described in detail in conjunction with the following implementation steps.

[0159] The process of obtaining the denoised sequence based on the trained models under different functions is as follows: according to x t ,t, first obtain x t-1 The sample, and then based on x t-1 ,t-1 yields x t-2 The samples are traversed in this way until t=0, then the x0 sequence is obtained, which is the sequence generated after removing all noise.

[0160] The specific process is as follows:

[0161] a) When the model function is to predict noise When x is obtained t-1 The sample format is as follows:

[0162] Given x t-1 Obey N(x) t-1 μ θ ,∑ q (t) can be obtained by sampling from this normal distribution to obtain sample x. t-1, where μ θ ,∑ q (t) is shown below:

[0163]

[0164] b) When the model function is to directly predict p(x0|x) t When x is obtained, t-1 The sample format is as follows:

[0165] 1. When the noise addition method is discrete, p is obtained. θ (x t-1 |x t )=∑ x0 p(x t-1 |x t ,x0)p θ (x0|x t ), and then from p θ (x t-1 |x t Sample x was obtained by sampling from ) t-1 Where p θ (x t-1 |x t The following is an example of how to calculate x0 by simply replacing the symbol q with p.

[0166]

[0167] 2. When the noise addition method is continuous, the model prediction output is: That is, the model prediction. Given x t-1 obey Sample x can be obtained by sampling from this normal distribution. t-1 Then you will receive

[0168]

[0169] c) When the model function is to directly predict p(x) t-1 |x t When x is obtained, t-1 The sample format is as follows:

[0170] 1. When the noise addition method is discrete, the model predicts p. θ (x t-1 |x t If x is a probability distribution, then x is obtained by sampling directly from that probability distribution. t-1 .

[0171] 2. When the noise addition method is continuous, the model predicts μ. θ Given that M is a constant, then according to x t-1obey x can be obtained by sampling directly from this normal distribution. t-1 arrive.

[0172] in,

[0173] After obtaining the predicted capsid protein sequence, proceed to step S4.

[0174] Step S4: Post-process the predicted capsid protein sequence to obtain the final capsid protein sequence.

[0175] After obtaining the predicted capsid protein sequence, post-processing can be performed on the predicted capsid protein sequence to obtain the final capsid protein sequence. Specifically, after obtaining the predicted capsid protein sequence, the specified characters contained in the predicted capsid protein sequence can be detected and replaced with empty characters to generate the final capsid protein sequence.

[0176] The above process can be combined Figure 5 Examples are described below.

[0177] like Figure 5 As shown, assuming the original sequence of the mutation region of a selected sample is "DEEEIR", the diffusion process involves adding noise to the sequence at each time step. At t=0, the original sequence is used. Assuming discrete noise is selected, after noise addition, at t=1, sampling is performed based on the distribution calculated from the transition probability matrix. Assuming the first amino acid sampled is [M] (set as a special type) and is different from the original D, and the remaining amino acids remain unchanged, then at t=1, the sequence becomes "[M]EEEIR". After adding noise for T time steps, at t=T, the original sequence stably transforms into "[M][M][M][M][M][M]". The denoising process, i.e., the generation process, is as follows: given a sequence consisting entirely of "[M][M][M][M][M][M]", a pre-trained model is used to denoise step by step to obtain "[M]EEEIR", and finally, another round of denoising yields "DEEEIR".

[0178] The training method for capsid protein sequence generation model provided in this application involves acquiring training data, which is a set of sequences with a specified capsid function within a selected region of a protein molecule. Data augmentation is performed on the training data to obtain model training data. Noise is added to the model training data at different times based on a set noise addition method, resulting in noisy training data at different times. The capsid protein sequence generation model to be trained is then trained based on the noisy training data and a loss function under a specified generation path, resulting in a capsid protein sequence generation model with the specified capsid function under the specified generation path. This application embodiment constructs a model capable of generating a large number of desired capsid functions based on a small amount of existing data on the required capsid function (functions may include, but are not limited to, active capsid, deimmunogenic capsid, or capsid targeting a specific cell), thereby achieving the goal of designing capsid sequences. The proportion of usable sequences generated is higher than in existing methods. Furthermore, since the model provided in this embodiment is a capsid sequence generator that can directly generate capsid sequences containing the required functions of the capsid, it is expected that in the same amount of time, the number of usable sequences found by this application (searching along the direction of the capsid with the function) will be much larger than the number of sequences found by random mutation (searching in the entire domain).

[0179] Reference Figure 6 The diagram shows a schematic representation of the structure of a training device for a capsid protein sequence generation model provided in an embodiment of this application. Figure 6 As shown, the training device 600 for the capsid protein sequence generation model may include the following modules:

[0180] The training data acquisition module 610 is used to acquire training data, which is a set of sequences that have a specified capsid function in a selected region of a protein molecule.

[0181] The model data acquisition module 620 is used to perform data augmentation processing on the training data to obtain model training data;

[0182] The noise-added data acquisition module 630 is used to perform noise-added processing on the model training data at different times based on a set noise-added processing method, so as to obtain noise-added training data at different times.

[0183] The model acquisition module 640 is used to train the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path, so as to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path.

[0184] Optionally, the model data acquisition module includes:

[0185] The model data acquisition unit is used to add a set character at any position in the sequence corresponding to each training data to obtain the model training data with the same sequence length.

[0186] The set character is a character without meaning.

[0187] Optionally, the device further includes:

[0188] The protein sequence acquisition module is used to acquire completely noisy protein sequences, along with specified model functions;

[0189] The target generation model acquisition module is used to acquire the target capsid protein sequence generation model corresponding to the specified model function.

[0190] The predicted protein sequence acquisition module is used to input the completely noisy protein sequence into the target capsid protein sequence generation model, so that the target capsid protein sequence generation model processes the completely noisy protein sequence according to the specified generation path to obtain the denoised predicted capsid protein sequence corresponding to the completely noisy protein sequence.

[0191] The capsid protein sequence acquisition module is used to post-process the predicted capsid protein sequence to obtain the final capsid protein sequence.

[0192] Optionally, the capsid protein sequence acquisition module includes:

[0193] A character detection unit is set up to detect the set character contained in the predicted capsid protein sequence;

[0194] The capsid protein sequence generation unit is used to replace the set character with an empty character to generate the final capsid protein sequence.

[0195] Optionally, the noise addition processing method includes either a discrete noise addition method or a continuous noise addition method.

[0196] Optionally, when the noise processing method is set to discrete noise processing,

[0197] The noisy data acquisition module includes:

[0198] The first noisy data acquisition unit is used to add noise to the model training data according to the predefined transition probability matrix at different times, so as to obtain the noisy training data corresponding to the model training data at different times.

[0199] Optionally, when the noise processing method is set to continuous noise processing,

[0200] The noisy data acquisition module includes:

[0201] The second noise-added data acquisition unit is used to add noise to the model training data according to the noise ratio at different times and the latent vector of each amino acid in the sequence set under the specified capsid function, so as to obtain the noise-added training data corresponding to the model training data at different times.

[0202] Optionally, when the specified generation path is a predicted noise path,

[0203] The generative model acquisition module includes:

[0204] The first loss function acquisition unit is used to acquire the first loss function corresponding to the predicted noise path;

[0205] The first generative model acquisition unit is used to train the capsid protein sequence generation model to be trained based on the noisy training data and the first loss function, so as to obtain the capsid protein sequence generation model with the specified capsid function under the predicted noise path.

[0206] Optionally, when the specified generation path is the path for predicting a noisy sequence,

[0207] The generative model acquisition module includes:

[0208] The second loss function acquisition unit is used to determine the second loss function based on the set noise processing method and the predicted path of the unnoised sequence.

[0209] The second generative model acquisition unit is used to train the capsid protein sequence generation model to be trained based on the noisy training data and the second loss function, so as to obtain the capsid protein sequence generation model with the specified capsid function under the path of the predicted unnoisy sequence.

[0210] Optionally, when the specified generation path is a path that predicts the probability of the noisy sequence at the previous time step,

[0211] The generative model acquisition module includes:

[0212] The third loss function acquisition unit is used to determine the third loss function based on the set noise processing method and the path of predicting the probability of the noise sequence at the previous time step.

[0213] The third generative model acquisition unit is used to train the capsid protein sequence generation model to be trained based on the noisy training data and the third loss function, so as to obtain the capsid protein sequence generation model under the path of the probability of the noisy sequence at the previous prediction time for the specified capsid function.

[0214] The training apparatus for the capsid protein sequence generation model provided in this application acquires training data, which is a set of sequences with a specified capsid function within a selected region of a protein molecule. Data augmentation processing is performed on the training data to obtain model training data. Noise is added to the model training data at different times based on a set noise addition method, resulting in noisy training data at different times. The capsid protein sequence generation model to be trained is then trained based on the noisy training data and a loss function under a specified generation path, resulting in a capsid protein sequence generation model with the specified capsid function under the specified generation path. This application embodiment constructs a model capable of generating a large number of desired capsid functions based on a small amount of existing data on the required capsid function (functions may include, but are not limited to, active capsid, deimmunogenic capsid, or capsid targeting a specific cell), thereby achieving the purpose of designing capsid sequences. The proportion of usable sequences generated is higher than in existing methods. Furthermore, since the model provided in this embodiment is a capsid sequence generator that can directly generate capsid sequences containing the required functions of the capsid, it is expected that in the same amount of time, the number of usable sequences found by this application (searching along the direction of the capsid with the function) will be much larger than the number of sequences found by random mutation (searching in the entire domain).

[0215] This application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the training method of the capsid protein sequence generation model described above.

[0216] Figure 7 A schematic diagram of the structure of an electronic device 700 according to an embodiment of the present invention is shown. Figure 7 As shown, the electronic device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 702 or loaded from storage unit 708 into random access memory (RAM) 703. The RAM 703 can also store various programs and data required for the operation of the electronic device 700. The CPU 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0217] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, microphone, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0218] The various processes and handling described above can be executed by processing unit 701. For example, the methods of any of the above embodiments can be implemented as computer software programs, which are tangibly contained in a computer-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more actions of the methods described above can be performed.

[0219] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the training method for the capsid protein sequence generation model described above.

[0220] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0221] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0222] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminals (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal, generate instructions for implementing the flowchart illustrations. Figure 1One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0223] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0224] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal, causing a series of operational steps to be executed on the computer or other programmable terminal to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0225] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0226] The foregoing has provided a detailed description of a training method for a capsid protein sequence generation model, a training device for a capsid protein sequence generation model, an electronic device, and a computer-readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A training method for a capsid protein sequence generation model, characterized in that, The method includes: Acquire training data, which is a set of sequences that have a specified capsid function in a selected region of a protein molecule; The training data is augmented to obtain model training data; The training data of the model is subjected to noise processing at different times based on the set noise processing method, so as to obtain noisy training data at different times. The capsid protein sequence generation model to be trained is trained based on the noisy training data and the loss function under the specified generation path to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path. The noise addition processing method includes either a discrete noise addition method or a continuous noise addition method. When the noise processing method is set to discrete noise processing method. The step of adding noise to the model training data at different times based on a set noise-adding method to obtain noisy training data at different times includes: The model training data is noise-added according to the predefined transition probability matrix at different times to obtain the noise-added training data at different times corresponding to the model training data. Wherein, when the specified generation path is the path for predicting a noisy sequence, The step of training the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path, to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path, includes: Based on the set noise processing method and the predicted path of the un-noiseed sequence, a second loss function is determined; The capsid protein sequence generation model is trained based on the noisy training data and the second loss function to obtain the capsid protein sequence generation model with the specified capsid function under the path of the predicted unnoisy sequence.

2. The method according to claim 1, characterized in that, The step of performing data augmentation on the training data to obtain model training data includes: Add a set character at any position in the sequence corresponding to each training data to obtain model training data with the same sequence length; The set character is a character without meaning.

3. The method according to claim 2, characterized in that, After training the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path, the method further includes: Obtain the protein sequence with complete noise, along with the specified model function; Obtain the target capsid protein sequence generation model corresponding to the specified model function; The completely noisy protein sequence is input into the target capsid protein sequence generation model, so that the target capsid protein sequence generation model processes the completely noisy protein sequence according to the specified generation path to obtain the denoised predicted capsid protein sequence corresponding to the completely noisy protein sequence. The predicted capsid protein sequence is post-processed to obtain the final capsid protein sequence.

4. The method according to claim 3, characterized in that, The post-processing of the predicted capsid protein sequence to obtain the final capsid protein sequence includes: The predetermined character contained in the predicted capsid protein sequence was detected; Replace the specified character with an empty character to generate the final capsid protein sequence.

5. The method according to claim 1, characterized in that, When the noise addition processing method can also be a continuous noise addition method. The step of adding noise to the model training data at different times based on a set noise-adding method to obtain noisy training data at different times includes: The model training data is noise-added based on the predefined noise weight at different times and the latent vector of each amino acid in the sequence set under the specified capsid function, to obtain the noise-added training data corresponding to the model training data at different times.

6. The method according to claim 1, characterized in that, When the specified generation path is a predicted noise path, The step of training the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path, to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path, includes: Obtain the first loss function corresponding to the predicted noise path; The capsid protein sequence generation model is trained based on the noisy training data and the first loss function to obtain the capsid protein sequence generation model with the specified capsid function under the predicted noise path.

7. The method according to claim 1, characterized in that, When the specified generation path can also be a path for predicting the probability of the noisy sequence at the previous time step, The step of training the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path, to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path, includes: Based on the set noise processing method and the path of predicting the probability of the noise-added sequence at the previous time step, the third loss function is determined; The capsid protein sequence generation model is trained based on the noisy training data and the third loss function to obtain the capsid protein sequence generation model under the path of the probability of the noisy sequence at the previous prediction time for the specified capsid function.

8. A training device for a capsid protein sequence generation model, characterized in that, The device includes: The training data acquisition module is used to acquire training data, which is a set of sequences that have a specified capsid function in a selected region of a protein molecule. The model data acquisition module is used to perform data augmentation processing on the training data to obtain model training data; The noise-added data acquisition module is used to add noise to the model training data at different times based on a set noise-adding processing method, so as to obtain noise-added training data at different times. The model acquisition module is used to train the capsid protein sequence generation model to be trained based on the noisy training data and the loss function under the specified generation path, so as to obtain the capsid protein sequence generation model with the specified capsid function under the specified generation path. The noise addition processing method includes either a discrete noise addition method or a continuous noise addition method. When the noise processing method is set to discrete noise processing method. The noisy data acquisition module includes: The first noisy data acquisition unit is used to add noise to the model training data according to the predefined transition probability matrix at different times, so as to obtain the noisy training data corresponding to the model training data at different times. Wherein, when the specified generation path is the path for predicting a noisy sequence, The generative model acquisition module includes: The second loss function acquisition unit is used to determine the second loss function based on the set noise processing method and the predicted path of the unnoised sequence. The second generative model acquisition unit is used to train the capsid protein sequence generation model to be trained based on the noisy training data and the second loss function, so as to obtain the capsid protein sequence generation model with the specified capsid function under the path of the predicted unnoisy sequence.

Citation Information

Patent Citations

  • Training method and device of self-supervised learning model, equipment and storage medium

    CN112420123A

  • DNA origami subunits and their use for encapsulation of filamentous virus particles

    WO2022261312A1