An antibody sequence generation method and device based on deep learning diffusion generation
Through the antibody sequence generation method based on deep learning diffusion, the problems of high computational costs, strong data dependence, insufficient sequence diversity and limited model generalization capabilities in the prior art are solved, efficient and accurate antibody sequence generation is achieved, and a comprehensive evaluation system is provided.
Patent Information
- Application Number
- CN202411426755.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-10-14
AI Technical Summary
The existing antibody design technology has problems such as high computational cost, strong data dependence, insufficient sequence diversity and limited model generalization capabilities, and lacks a comprehensive evaluation system.
Using the antibody sequence generation method based on deep learning diffusion generation, high-quality antibody sequences are generated through diffusion deep learning model, complete data preprocessing process and modular design. The method includes initializing the generation of the initial antibody CDR sequence, performing step-by-step denoising through the trained antibody sequence generator, generating the target antibody sequence, and screening through protein structure prediction and antigen antibody binding evaluation.
It improves the efficiency and accuracy of antibody sequence generation, reduces calculation costs, enhances data utilization, improves sequence diversity and model generalization capabilities, and provides a comprehensive evaluation system.
Smart Images

Figure CN118942533B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technologies, and particularly to an antibody sequence generation method and device based on deep learning diffusion generation. Background Art
[0002] Antibody design is one of the important research directions in the field of biomedical technologies. Antibodies are key defense tools in the immune system, and their main function is to recognize and neutralize antigens such as foreign pathogens. By precisely designing and optimizing antibody sequences, highly effective treatment means against specific antigens can be developed. Therefore, antibody design has great application potential in the treatment of cancer, infectious diseases, autoimmune diseases, etc.
[0003] The technologies in the field of antibody design mainly include traditional experimental methods and modern computational methods. Traditional antibody design mainly relies on laboratory screening techniques such as phage display technology and monoclonal antibody preparation. These methods find antibodies that bind efficiently to the target antigen through a large number of experimental cultivations and screenings. However, experimental methods are time-consuming, laborious, costly, and have low efficiency and success rates.
[0004] In recent years, with the improvement of computing power and the development of bioinformatics, computational methods have been widely used in antibody design. The main computational methods include molecular dynamics simulation, computer-aided antibody engineering, and deep learning methods. Molecular dynamics simulation predicts the binding force and stability of antibodies by simulating the interaction between antibodies and antigens, thereby screening suitable antibodies. This method requires high-performance computing resources, has a long simulation time, and cannot optimize and design unknown antigens, belonging to an accelerated screening process. Computer-aided antibody engineering aims to overcome the problems of low antibody performance and poor resistance of unknown antibodies, and uses algorithms and databases, combined with experimental data, to design and optimize antibody sequences. This method can improve the efficiency of antibody design, but depends on the quality of existing data and the optimization degree of algorithms. The development of deep learning methods has brought new progress to computer-aided antibody engineering. In recent years, deep learning has made remarkable progress in protein structure prediction and design. For example, the success of AlphaFold has demonstrated the great potential of deep learning in protein structure prediction. These methods can quickly predict the binding mode between antibodies and antigens and generate optimized antibody sequences by training neural network models. Compared with past annealing algorithms, genetic algorithms, etc., deep learning methods can often design more effective antibody sequences.
[0005] Although the existing technologies have made important progress in antibody design, there are still the following main problems:
[0006] High computational cost: Whether it is molecular dynamics simulation or deep learning model training, a large amount of computing resources are required, which limits large-scale applications and popularization and requires professional computing centers to provide services and support.
[0007] Strong data dependence: The effectiveness of computational methods depends to a large extent on the quality and quantity of sample data. The antibody-antigen complex data in existing databases may be biased or insufficient, affecting the design effectiveness of the model.
[0008] Insufficient sequence diversity: The antibody sequences generated by traditional and some computational methods have limited diversity, and may not be able to design antibodies that effectively target antigen binding sites, resulting in poor performance of the designed antibodies in practical applications.
[0009] Limited generalization ability of the model: Deep learning models often perform better on sample data than on new data. How to improve the generalization ability of the model on new antigens is an important research direction.
[0010] Lack of a comprehensive evaluation system: Currently, the evaluation of antibody design mainly relies on computational simulation and a small amount of experimental verification, lacking a systematic and comprehensive evaluation system, making it difficult to effectively evaluate the function and stability of designed antibodies. Summary of the Invention
[0011] The purpose of the present invention is to overcome the deficiencies in the prior art and provide an antibody sequence generation method and device based on deep learning diffusion generation. By introducing diffusion deep learning, a complete data preprocessing process, and modular design, it aims to improve the efficiency and accuracy of antibody sequence generation and solve some bottleneck problems in the prior art.
[0012] To achieve the above object, the present invention is implemented by the following technical solutions:
[0013] In the first aspect, the present invention provides an antibody sequence generation method based on deep learning diffusion generation, including:
[0014] Initialize to generate an initial antibody CDR sequence; the initial antibody CDR sequence is a completely random noise data;
[0015] Gradually denoise the initial antibody CDR sequence through a trained antibody sequence generator, and generate a target antibody CDR sequence at each step;
[0016] Splice each of the target antibody CDR sequences with a template antibody sequence to generate a target antibody sequence;
[0017] Perform protein structure prediction on each of the target antibody sequences through a protein structure prediction model to generate corresponding protein structures;
[0018] Physically stitch together the protein structures of the antigens in the protein structures of the antibodies and antigens complex data generated by each prediction and the antigens input by the user to generate a target antigen-antibody complex;
[0019] Evaluate each of the target antigen-antibody complexes using an antigen-antibody binding evaluation method, screen out the optimal target antigen-antibody complex, and output the corresponding target antibody sequence.
[0020] Optionally, the training of the antibody sequence generator includes:
[0021] Construct a deep learning diffusion generation model, which includes a diffusion model, a residue embedding module, and a residue sequence prediction module; the residue embedding module and the residue sequence prediction module constitute the antibody sequence generator;
[0022] Obtain antibody-antigen complex data for training, and perform preprocessing to generate antibody-antigen complex sample data;
[0023] The diffusion model includes a diffusion process and an inverse diffusion process. In the diffusion process, a step-by-step random noise addition operation is performed on the antibody CDR sequence in the antibody-antigen complex sample data. In the inverse diffusion process, a step-by-step denoising operation is performed on the antibody CDR sequence finally obtained in the diffusion process by the antibody sequence generator;
[0024] Calculate the model loss based on the antibody CDR sequence generated by the noise addition operation and the antibody CDR sequence generated by the antibody sequence generator, and backpropagate to optimize the model parameters of the antibody sequence generator to complete the training.
[0025] Optionally, the preprocessing of the trained antibody-antigen complex data includes:
[0026] Select the target antigen type from the antibody-antigen complex data file;
[0027] Parse the antibody sequences corresponding to the target antigen type, determine whether there is an antibody Fab region, whether there are missing values, whether there are multiple protein structure conformations, and whether there are unconventional amino acids, and perform repair processing;
[0028] Perform adjustment processing on the repaired antibody sequences to generate antibody-antigen complex sample data. The adjustment processing includes amino acid sequence number rearrangement, antibody sequence floating-point vectorization, continuity check, and CDR region marking;
[0029] Classify the antibody-antigen complex sample data through a sequence classifier, and construct a training set and a test set according to the classification results.
[0030] Optionally, the preprocessed antibody-antigen complex data file also needs to undergo feature construction, which includes generating a mask for the labeled CDR region, antigen region labeling, random antibody mutation, and binding site feature calculation.
[0031] Optionally, the residue embedding module includes a first feature extraction unit, a second feature extraction unit, and a third feature extraction unit;
[0032] The first feature extraction unit includes a first linear layer, a first Than layer, an identity mapping layer, a second linear layer, and a second Than layer connected in sequence; obtaining dihedral angle information and position encoding from the antibody-antigen complex sample data, and fusing them to obtain dihedral angle structure features; inputting the dihedral angle structure features into the first feature extraction unit to generate structure feature embedding vectors;
[0033] The second feature extraction unit includes a first embedding layer; obtaining atom types and position encoding from the antibody-antigen complex sample data, and fusing them to obtain atom type features; inputting the atom type features into the second feature extraction unit to generate atom type embedding vectors;
[0034] The third feature extraction unit includes a second embedding layer; obtaining residue types and position encoding from the antigen-antibody complex sample data, and fusing them to obtain residue type features; inputting the residue type features into the third feature extraction unit to generate residue type embedding vectors;
[0035] Align the dimensions of the structure feature embedding vectors, the atom type embedding vectors, and the residue type embedding vectors through a broadcasting mechanism, and add them to generate sequence feature embedding vectors.
[0036] Optionally, the residue sequence prediction module includes a plurality of Transformer decoders and a predictor connected in sequence;
[0037] The Transformer decoder includes a multi-head attention layer, a first normalization layer, a feed-forward neural network layer, and a second normalization layer connected in sequence, and the input of the multi-head attention layer and the output of the first normalization layer are added and connected, and the input of the feed-forward neural network and the input of the second normalization layer are added and connected; the feed-forward neural network layer includes an amplification linear layer, a first ReLU layer, and a contraction linear layer connected in sequence;
[0038] The predictor includes a fifth linear layer, a second ReLU layer, a sixth linear layer, and a third ReLU layer connected in sequence;
[0039] Input the sequence feature embedding vectors obtained by the residue embedding module into the residue sequence prediction module to generate the antibody CDR sequence.
[0040] Optionally, the random noise addition operation in the diffusion process is a predefined Markov chain process with a transition probability of , being the antibody CDR sequences at the -th step in the diffusion process respectively; the noise removal operation in the reverse diffusion process is a parameterized Markov chain process with a transition probability of , being the antibody CDR sequences at the -th step in the reverse diffusion process respectively, being the model parameters of the antibody sequence generator.
[0041] Optionally, the antibody sequence generator minimizes the model loss by maximizing the variational lower bound;
[0042] The variational lower bound is:
[0043] ;
[0044] In the formula, is the expectation of the approximate posterior distribution obeyed by the antibody CDR sequence at the 0-th step in the diffusion process. The probability distribution of the antibody CDR sequences in the antibody-antigen complex sample data is used as the approximate posterior distribution ; is the expected value of the conditional probability distribution of the antibody CDR sequence at the 1-st step in the diffusion process given the antibody CDR sequence at the 0-th step; is the log-likelihood of the conditional probability distribution of the antibody CDR sequence at the 0-th step in the reverse diffusion process given the antibody CDR sequence at the 1-st step; is the KL divergence function, is the conditional probability distribution of the antibody CDR sequence T at the -th step in the diffusion process given the antibody CDR sequence , T is the total number of steps in the diffusion process or the reverse diffusion process; is the preset prior distribution. The probability distribution of the antibody CDR sequence finally obtained in the diffusion process is used as the prior distribution ; is the expected value of the conditional probability distribution of the antibody CDR sequence at the -th step in the diffusion process given the antibody CDR sequence ; The antibody CDR sequence at the th step during the diffusion process The conditional probability distribution given the antibody CDR sequence ; The antibody CDR sequence at the th step during the inverse diffusion process The conditional probability distribution given the antibody CDR sequence at the th step ;
[0045] A loss function for constructing the model loss based on the variational lower bound : :
[0046] ;
[0047] In the formula, is a hyperparameter, is the log-likelihood of the conditional probability distribution of the antibody CDR sequence during the inverse diffusion process given the antibody CDR sequence .
[0048] In a second aspect, the present invention provides an antibody sequence generation device based on deep learning diffusion generation, including:
[0049] An initialization module configured to initialize and generate an initial antibody CDR sequence; the initial antibody CDR sequence is a completely random noise data;
[0050] A target generation module configured to gradually denoise the initial antibody CDR sequence through a trained antibody sequence generator, and generate a target antibody CDR sequence at each step; respectively splice each of the target antibody CDR sequences with a template antibody sequence to generate a target antibody sequence;
[0051] A structure prediction module configured to perform protein structure prediction on each of the target antibody sequences through a protein structure prediction model to generate corresponding protein structures;
[0052] A physical splicing module configured to physically splice each of the predicted protein structures and the protein structure of the antigen in the antibody-antigen complex data input by the user to generate a target antigen-antibody complex;
[0053] An antibody screening module configured to evaluate each of the target antigen-antibody complexes by using an antigen-antibody binding evaluation method, screen out the optimal target antigen-antibody complex, and output the corresponding target antibody sequence.
[0054] In a third aspect, the present invention provides an electronic device, including a processor and a storage medium;
[0055] The storage medium is used for storing instructions;
[0056] The processor is used to operate according to the instructions to execute the steps of the above method.
[0057] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0058] The present invention provides a method and device for generating antibody sequences based on deep learning diffusion generation. 1) The antibody CDR sequence is generated through an antibody sequence generator based on diffusion deep learning, and then combined with the template antibody sequence to generate the required target antibody sequence. The sequence generation process is efficient and accurate; then, the target antigen-antibody complex is generated based on the target antibody sequence, and screening is performed through the antigen-antibody binding evaluation method to obtain the final required target antibody sequence. 2) In the diffusion deep learning process, through the forward and reverse diffusion processes, noise can be gradually removed, so as to generate an antibody CDR sequence that matches the input data distribution. This method can generate higher-quality sequences, and the training process can more stably optimize the model parameters. 3) Through the residue embedding module and the residue sequence prediction module, the antibody sequence can be accurately predicted. The residue embedding module uses the protein molecular structure and species characteristics for residue embedding, combined with the multi-layer stacking processing of the Transformer Decoder, so that the predictor can work better. 4) Through detailed data preprocessing and feature extraction steps, the quality and consistency of the input data are ensured, thereby improving the effect of model training and the accuracy of the generated sequence. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is a schematic flowchart of the method for generating antibody sequences based on deep learning diffusion generation provided by an embodiment of the present invention;
[0060] Figure 2 is a schematic structural diagram of the antibody sequence generator provided by an embodiment of the present invention;
[0061] Figure 3 is a schematic structural diagram of the residue embedding module provided by an embodiment of the present invention;
[0062] Figure 4 is a schematic structural diagram of the residue sequence prediction module provided by an embodiment of the present invention;
[0063] Figure 5 is a schematic diagram of the diffusion process and the inverse diffusion process of the antibody sequence generator provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and cannot be used to limit the protection scope of the present invention.
[0065] Embodiment 1:
[0066] As Figure 1 shown, the present invention provides an antibody sequence generation method based on deep learning diffusion generation, including the following steps:
[0067] Step S1, initialize to generate an initial antibody CDR sequence; the initial antibody CDR sequence is a completely random noise data.
[0068] The initial antibody CDR sequence is the starting point for the generation process of the target antibody CDR sequence.
[0069] Step S2, gradually denoise the initial antibody CDR sequence through a trained antibody sequence generator, and generate a target antibody CDR sequence at each step.
[0070] For the training of the antibody sequence generator, it is necessary to first collect the original antibody-antigen complex data files from the database, dataset or user input. The file format of the antibody-antigen complex data files is generally the Protein Data Bank (PDB) file. The antibody-antigen complex in the PDB file is specifically the product of the physical binding of an antibody and an antigen.
[0071] The antibody in the antibody-antigen complex is specifically the Fab region (Fragment antigen-binding region) of a human antibody, including the variable regions (VH and VL) of the antibody and their binding sites. The antigen in the antibody-antigen complex is specifically the fragment of the ReceptorBinding Domain (RBD) region on the selected Antigen Protein.
[0072] Specifically, in this embodiment, the training of the antibody sequence generator includes:
[0073] Step S2.1, construct a deep learning diffusion generation model, the deep learning diffusion generation model includes a diffusion model, a residue embedding module and a residue sequence prediction module; as Figure 2 shown, the residue embedding module and the residue sequence prediction module constitute the antibody sequence generator.
[0074] As Figure 3 shown, the residue embedding module includes a first feature extraction unit, a second feature extraction unit and a third feature extraction unit.
[0075] The first feature extraction unit includes a first linear layer, a first Than layer, an identity mapping layer, a second linear layer, and a second Than layer connected in sequence; it obtains dihedral angle information and position encoding from the antibody-antigen complex sample data, and fuses them to obtain dihedral angle structure features; it inputs the dihedral angle structure features into the first feature extraction unit to generate structure feature embedding vectors;
[0076] In the molecular structure features, the angle between two planes formed by four consecutive atoms in a protein molecule is called the dihedral angle. This feature is very important for the dynamic properties and structural stability of proteins. Therefore, the dihedral angle is added to the position encoding as the protein molecular structure feature.
[0077] The second feature extraction unit includes a first embedding layer; it obtains atom types and position encoding from the antibody-antigen complex sample data, and fuses them to obtain atom type features; it inputs the atom type features into the second feature extraction unit to generate atom type embedding vectors;
[0078] The third feature extraction unit includes a second embedding layer; it obtains residue types and position encoding from the antigen-antibody complex sample data, and fuses them to obtain residue type features; it inputs the residue type features into the third feature extraction unit to generate residue type embedding vectors;
[0079] The dimensions of the structure feature embedding vectors, atom type embedding vectors, and residue type embedding vectors are aligned through a broadcasting mechanism and added together to generate sequence feature embedding vectors.
[0080] As Figure 4 shown, the residue sequence prediction module includes multiple Transformer decoders and a predictor connected in sequence; the residue sequence prediction module is used to predict and generate an antibody sequence based on the sequence feature embedding vectors.
[0081] The Transformer decoder includes a multi-head attention layer, a first normalization layer, a feed-forward neural network layer, and a second normalization layer connected in sequence. The input of the multi-head attention layer is added and connected to the output of the first normalization layer, and the input of the feed-forward neural network is added and connected to the input of the second normalization layer; the feed-forward neural network layer includes an amplification linear layer, a first ReLU layer, and a contraction linear layer connected in sequence;
[0082] The predictor includes a fifth linear layer, a second ReLU layer, a sixth linear layer, and a third ReLU layer connected in sequence;
[0083] The sequence feature embedding vectors obtained by the residue embedding module are input into the residue sequence prediction module to generate the antibody CDR sequence.
[0084] The output after the multi-head self-attention mechanism is processed by layer normalization to standardize the activation values of each layer to stabilize the training process. The normalized output is further processed by a feed-forward neural network. This network typically consists of two linear transformations and a ReLU activation function, which increases the model's information expression ability by performing non-high-dimensional linear transformations on the input. The output of the feed-forward neural network is again processed by layer normalization to further stabilize and normalize the output. Residual connections are introduced between the inputs and outputs of the multi-head self-attention mechanism and the feed-forward neural network, thereby alleviating the vanishing gradient problem and improving the training effect of the model. Each decoder layer repeats the above structure. By stacking multiple such layers, the Transformer Decoder can gradually construct more complex feature representations, achieve efficient sequence feature to sequence feature conversion, and enable the predictor to work better. The stacking layer number parameter of the Transformer Decoder module is 7, the multi-head parameter is 8, the dimension of the hidden layer in the multi-head self-attention layer is 336; the scaling factor of the residual connection is 1; the predictor dimension parameter is 128.
[0085] Step S2.2: Obtain the antibody-antigen complex data for training, and perform preprocessing to generate antibody-antigen complex sample data.
[0086] The preprocessing process specifically includes:
[0087] Step S2.2.1: Select the target antigen type from the antibody-antigen complex data file.
[0088] Step S2.2.2: Analyze the antibody sequence corresponding to the target antigen type, and determine whether there is an antibody Fab region, whether there are missing values, whether there are multiple protein structural conformations, and whether there are unconventional amino acids, and perform repair processing. If it cannot be repaired, the calculation is aborted and the link problem is prompted.
[0089] Step S2.2.3: Perform adjustment processing on the repaired antibody sequence to generate antibody-antigen complex sample data. The adjustment processing includes amino acid sequence number rearrangement, antibody sequence floating-point vectorization, continuity check, and CDR region marking.
[0090] Step S2.2.4: Classify the antibody-antigen complex sample data through a sequence classifier, and construct a training set and a test set according to the classification results. When training for the first time or when there is no classifier, use MMseqs2 to construct an unsupervised classifier and perform similarity clustering training on the sequences.
[0091] The preprocessed antibody-antigen complex data file also needs to perform feature construction, and the feature construction includes generating a mask for the labeled CDR region, antigen region marking, antibody random mutation, and binding site feature calculation.
[0092] Step S2.3: The diffusion model includes a diffusion process and an inverse diffusion process. As shown in Figure 5 , during the diffusion process, a step-by-step random noise addition operation is performed on the antibody CDR sequence in the antibody-antigen complex sample data. During the inverse diffusion process, a step-by-step noise removal operation is performed on the antibody CDR sequence finally obtained in the diffusion process through an antibody sequence generator.
[0093] The random noise addition operation during the diffusion process is a predefined Markov chain process, and its transition probability is , are the antibody CDR sequences at the th step during the diffusion process respectively; the noise removal operation during the inverse diffusion process is a parameterized Markov chain process, and its transition probability is , are the antibody CDR sequences at the th step during the inverse diffusion process respectively, is the model parameter of the antibody sequence generator.
[0094] Step S2.4: Calculate the model loss using the antibody CDR sequences generated by the noise addition operation and the antibody CDR sequences generated by the antibody sequence generator, and backpropagate to optimize the model parameters of the antibody sequence generator to complete the training.
[0095] The antibody sequence generator minimizes the model loss by maximizing the variational lower bound;
[0096] The variational lower bound is:
[0097] ;
[0098] In the formula, is the expectation of the approximate posterior distribution obeyed by the antibody CDR sequence at the 0th step during the diffusion process. The probability distribution of the antibody CDR sequence in the antibody-antigen complex sample data is used as the approximate posterior distribution ; is the conditional probability distribution expectation of the antibody CDR sequence at the 1st step during the diffusion process given the antibody CDR sequence at the 0th step; is the log-likelihood of the conditional probability distribution of the antibody CDR sequence at the 0th step during the inverse diffusion process given the antibody CDR sequence at the 1st step; is the KL divergence function, is the KL divergence function, is at the TThe CDR sequence of the antibody at step The conditional probability distribution given the CDR sequence of the antibody is the total number of steps in the diffusion process or the inverse diffusion process; T is the preset prior distribution, and the probability distribution of the CDR sequence of the antibody finally obtained in the diffusion process is used as the prior distribution ; ; is the CDR sequence of the antibody at step in the diffusion process The expected value of the conditional probability distribution given the CDR sequence of the antibody ; is the CDR sequence of the antibody at step in the diffusion process The conditional probability distribution given the CDR sequence of the antibody ; is the CDR sequence of the antibody at step in the inverse diffusion process The conditional probability distribution given the CDR sequence of the antibody at step ; ;
[0099] ;
[0100] In the formula, , , is matrix transpose, is element-wise multiplication, is the categorical distribution of specified by the parameter ; is the transition probability matrix at step , is the element in the th row and th column of the transition probability matrix ;
[0101] ;
[0102] is the number of categorical classes of the antibody sequence, is the sequence at step , and the sequence decreases with the number of steps.
[0103] The loss function for constructing the model loss based on the variational lower bound is :
[0104] ;
[0105] Wherein, is a hyperparameter, is the log-likelihood of the conditional probability distribution of the antibody CDR sequence in the reverse diffusion process given the antibody CDR sequence under the given antibody CDR sequence.
[0106] Step S3: Splices each target antibody CDR sequence with the template antibody sequence respectively to generate a target antibody sequence.
[0107] Step S4: Performs protein structure prediction on each target antibody sequence through a protein structure prediction model (Alphafold) to generate the corresponding protein structure.
[0108] Step S5: Physically splices each predicted protein structure with the protein structure of the antigen in the antibody-antigen complex data input by the user (ZDOCK) to generate a target antigen-antibody complex.
[0109] Step S6: Evaluates each target antigen-antibody complex using an antigen-antibody binding evaluation method (ROSETTA), screens out the optimal target antigen-antibody complex, and outputs the corresponding target antibody sequence.
[0110] Example 2:
[0111] Based on the antibody sequence generation method provided in Example 1, an antibody sequence generation device based on deep learning diffusion generation provided in an embodiment of the present invention includes:
[0112] An initialization module, configured to initialize and generate an initial antibody CDR sequence; the initial antibody CDR sequence is a completely random noise data;
[0113] A target generation module, configured to gradually denoise the initial antibody CDR sequence through a trained antibody sequence generator, and generate a target antibody CDR sequence at each step; splice each of the target antibody CDR sequences with the template antibody sequence respectively to generate a target antibody sequence;
[0114] A structure prediction module, configured to perform protein structure prediction on each of the target antibody sequences through a protein structure prediction model to generate the corresponding protein structure;
[0115] A physical splicing module, configured to physically splice each of the predicted protein structures with the protein structure of the antigen in the antibody-antigen complex data input by the user to generate a target antigen-antibody complex;
[0116] The antibody screening module is configured to evaluate each of the target antigen-antibody complexes by using an antigen-antibody binding evaluation method, screen out the optimal target antigen-antibody complex, and output the corresponding target antibody sequence.
[0117] Embodiment 3:
[0118] Based on the antibody sequence generation method provided in Embodiment 1, an embodiment of the present invention further provides an electronic device, which is characterized by including a processor and a storage medium;
[0119] The storage medium is used to store instructions;
[0120] The processor is configured to operate according to the instructions to execute the steps of the above method.
[0121] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0122] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0123] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the specified functions in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the process Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for the functions specified in one block or a plurality of blocks.
[0125] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for generating antibody sequences based on deep learning diffusion generation, characterized in that: include: Initializing and generating an initial antibody CDR sequence; the initial antibody CDR sequence is a completely random noise data; The initial antibody CDR sequence is gradually denoised by a trained antibody sequence generator, and a target antibody CDR sequence is generated at each step; Separately combining each of the target antibody CDR sequences with the template antibody sequence to generate a target antibody sequence; Performing protein structure prediction on each of the target antibody sequences using a protein structure prediction model to generate a corresponding protein structure; Physically splicing the predicted protein structures and the protein structure of the antigen in the antibody-antigen complex data input by the user to generate a target antigen-antibody complex; Using an antigen-antibody binding evaluation method to evaluate each of the target antigen-antibody complexes, screen out the optimal target antigen-antibody complex, and output the corresponding target antibody sequence; Wherein, the training of the antibody sequence generator includes: Constructing a deep learning diffusion generation model, wherein the deep learning diffusion generation model includes a diffusion model, a residue embedding module and a residue sequence prediction module; the residue embedding module and the residue sequence prediction module constitute the antibody sequence generator; Obtaining antibody-antigen complex data for training, and performing preprocessing to generate antibody-antigen complex sample data; The diffusion model includes a diffusion process and a reverse diffusion process. In the diffusion process, the antibody CDR sequence in the antibody-antigen complex sample data is subjected to a step-by-step random noise addition operation. In the reverse diffusion process, the antibody CDR sequence finally obtained by the diffusion process is subjected to a step-by-step denoising operation by the antibody sequence generator. The antibody CDR sequence generated by the noise addition operation and the antibody CDR sequence generated by the antibody sequence generator are used to calculate the model loss and back-propagate to optimize the model parameters of the antibody sequence generator to complete the training; The antibody sequence generator minimizes the model loss by maximizing the variational lower bound; The variational lower bound for: ; In the formula, is the antibody CDR sequence at step 0 in the diffusion process The approximate posterior distribution The probability distribution of the antibody CDR sequence in the antibody-antigen complex sample data is used as the approximate posterior distribution ; is the antibody CDR sequence in step 1 of the diffusion process Given the antibody CDR sequence at step 0 The conditional probability distribution expectation under ; is the antibody CDR sequence at step 0 in the reverse diffusion process Given the antibody CDR sequence in step 1 The log-likelihood of the conditional probability distribution under ; is the KL divergence function, In the diffusion process T Antibody CDR sequences In a given antibody CDR sequence The conditional probability distribution under T is the total number of steps in the diffusion process or reverse diffusion process; The probability distribution of the antibody CDR sequence finally obtained by the diffusion process is used as the prior distribution ; In the diffusion process Antibody CDR sequences In a given antibody CDR sequence The conditional probability distribution expectation under ; In the diffusion process Antibody CDR sequences In a given antibody CDR sequence The conditional probability distribution under ; In the reverse diffusion process Antibody CDR sequences In a given Antibody CDR sequences The conditional probability distribution under ; ; In the formula, , , is the matrix transpose, is element-wise multiplication, For the parameter Specified description The distribution of classification categories; In the diffusion process The transition probability matrix of the step, is the transition probability matrix Middle Line Elements of a column; ; is the number of classification categories of antibody sequences, For the The number of steps decreases with the number of steps; Based on the variational lower bound Constructing the loss function of the model loss : ; In the formula, is a hyperparameter, is the antibody CDR sequence in the reverse diffusion process In a given antibody CDR sequence The log-likelihood of the conditional probability distribution under ; Wherein, the residue embedding module includes a first feature extraction unit, a second feature extraction unit and a third feature extraction unit; The first feature extraction unit includes a first linear layer, a first Than layer, an identity mapping layer, a second linear layer and a second Than layer connected in sequence; dihedral angle information and position encoding are obtained from the antibody-antigen complex sample data, and are fused to obtain dihedral angle structural features; the dihedral angle structural features are input into the first feature extraction unit to generate a structural feature embedding vector; The second feature extraction unit includes a first embedding layer; obtaining atomic species and position codes from the antibody-antigen complex sample data, and fusing them to obtain atomic species features; inputting the atomic species features into the second feature extraction unit to generate an atomic species embedding vector; The third feature extraction unit includes a second embedding layer; obtaining residue types and position codes from the antigen-antibody complex sample data, and fusing them to obtain residue type features; inputting the residue type features into the third feature extraction unit to generate a residue type embedding vector; Aligning the dimensions of the structural feature embedding vector, the atom type embedding vector, and the residue type embedding vector through a broadcast mechanism, and adding them together to generate a sequence feature embedding vector; Wherein, the residue sequence prediction module includes a plurality of Transformer decoders and predictors connected in sequence; The Transformer decoder includes a multi-head attention layer, a first normalization layer, a feedforward neural network layer, and a second normalization layer connected in sequence, and the input of the multi-head attention layer is additively connected to the output of the first normalization layer, and the input of the feedforward neural network is additively connected to the input of the second normalization layer; the feedforward neural network layer includes an amplification linear layer, a first ReLU layer, and a contraction linear layer connected in sequence; The predictor includes a fifth linear layer, a second ReLU layer, a sixth linear layer, and a third ReLU layer connected in sequence; Inputting the sequence feature embedding vector obtained by the residue embedding module into the residue sequence prediction module to generate an antibody CDR sequence; The random noise addition operation in the diffusion process is a predefined Markov chain process, and its transition probability is , In the diffusion process, The operation of removing noise in the reverse diffusion process is a parameterized Markov chain process, and its transition probability is , are the first The antibody CDR sequence of the step are the model parameters of the antibody sequence generator.
2. The method for generating antibody sequences based on deep learning diffusion generation according to claim 1, characterized in that: The preprocessing of the training antibody-antigen complex data includes: Selecting a target antigen type from the antibody-antigen complex data file; Parse the antibody sequence corresponding to the target antigen type to determine whether there is an antibody Fab region, whether there are missing values, whether there are multi-protein structural conformations, and whether there are unconventional amino acids, and perform repair processing; The repaired antibody sequence is adjusted to generate antibody-antigen complex sample data, wherein the adjustment process includes amino acid sequence rearrangement, antibody sequence floating point vectorization, continuity check, and CDR region marking; The antibody-antigen complex sample data is classified by a sequence classifier, and a training set and a test set are constructed based on the classification results.
3. The method for generating antibody sequences based on deep learning diffusion generation according to claim 1, characterized in that: The pre-processed antibody-antigen complex data file also needs to undergo feature construction, which includes generating a mask for the marked CDR region, antigen region marking, random mutation of the antibody, and calculation of binding site features.
4. An antibody sequence generation device based on deep learning diffusion generation, characterized in that: The generating device is used to perform the steps of the method according to any one of claims 1 to 3, and the generating device comprises: An initialization module is configured to initialize and generate an initial antibody CDR sequence; the initial antibody CDR sequence is a completely random noise data; The target generation module is configured to gradually denoise the initial antibody CDR sequence through a trained antibody sequence generator, and generate a target antibody CDR sequence at each step; each of the target antibody CDR sequences is respectively combined with a template antibody sequence to generate a target antibody sequence; A structure prediction module is configured to perform protein structure prediction on each of the target antibody sequences using a protein structure prediction model to generate a corresponding protein structure; A physical splicing module is configured to physically splice the protein structures generated by the predictions and the protein structure of the antigen in the antibody-antigen complex data input by the user to generate a target antigen-antibody complex; The antibody screening module is configured to evaluate each of the target antigen-antibody complexes using an antigen-antibody binding evaluation method, screen out the optimal target antigen-antibody complex, and output its corresponding target antibody sequence.
5. An electronic device, characterized in that: including processor and storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-3.
Citation Information
Patent Citations
Deep learning model-based antibody structure optimization method and device
CN116741260A