Training method for antibody structure prediction model, antibody structure prediction method and device

By employing linear units for feature extraction and an equivariant attention module, the anti-body structure prediction model training process is optimized, reducing computational and storage needs, thus enhancing efficiency and resource utilization.

CN119229974BActive Publication Date: 2025-07-15INT DIGITAL ECONOMY ACAD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411130742.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2025-07-15
Estimated Expiration
2044-08-16

AI Technical Summary

Technical Problem

The training process of the existing antibody structure prediction model requires a large amount of GPU resources and time, the storage space needs a large amount, and hardware resources limit the efficiency of use.

Method used

Linear units are used for feature extraction. By converting the high-dimensional vector representation into two low-dimensional first representation vectors and second representation vectors, the model parameters and calculation amount are reduced, and high-dimensional vector representation is extracted using a pre-trained language model, and combined with the isovariable attention module for training.

Benefits of technology

The training speed of antibody structure prediction model is improved, the storage space requirement is reduced, the hardware resource requirements are reduced, and hardware limitations are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229974B_ABST
    Figure CN119229974B_ABST
Patent Text Reader

Abstract

The present application discloses a training method for an antibody structure prediction model, an antibody structure prediction method and a device. The training method specifically includes obtaining a high-dimensional vector representation of the amino acid sequence of an antibody sample; extracting features from the high-dimensional vector through a feature extraction module in an initial antibody structure prediction model to obtain two low-dimensional first representation vectors and second representation vectors; training the initial antibody structure prediction model based on the first representation vector and the second representation vector to obtain an antibody structure prediction model. In the present application, the feature extraction module uses a linear unit for feature extraction, reducing the model parameters of the antibody structure prediction model. On the one hand, this can reduce the storage space required by the antibody structure prediction model, and on the other hand, it can reduce the number of parameters to be trained by the antibody structure prediction model, improving the training speed of the antibody structure prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of biopharmaceutical technology, and particularly to a training method for an antibody structure prediction model, an antibody structure prediction method, and a device. Background Art

[0002] Antibody structure prediction is an important research direction in the fields of bioinformatics and computational biology, and is of great significance for understanding antibody function, designing novel therapeutic antibodies, and vaccine development. In recent years, with the development of machine learning technology, especially the application of deep learning methods, significant progress has been made in the field of antibody structure prediction. Existing antibody structure prediction models generally adopt large-scale neural network models, which makes the neural network have a large number of network parameters to be trained, resulting in the need for a large amount of GPU resources and a long training time during the training process, which seriously affects the research efficiency and the model iteration speed. In addition, the trained antibody structure prediction model requires a large amount of storage space to save model parameters and intermediate calculation results, restricting its use by hardware resources.

[0003] Therefore, the existing technology still needs to be improved. Summary of the Invention

[0004] The technical problem to be solved by the present application is to provide a training method for an antibody structure prediction model, an antibody structure prediction method, and a device in view of the deficiencies of the existing technology.

[0005] To solve the above technical problem, a first aspect of the present application provides a training method for an antibody structure prediction model, wherein the training method for the antibody structure prediction model specifically includes:

[0006] Obtain the amino acid coding sequence of the amino acid sequence of an antibody sample, and determine the high-dimensional vector representation of the amino acid sequence of the antibody sample based on the amino acid coding sequence;

[0007] Input the high-dimensional vector representation into the feature extraction module in the initial antibody structure prediction model, and determine the attention coding matrix through the attention unit in the feature extraction module;

[0008] According to the high-dimensional vector representation and the attention coding matrix, determine the first representation vector through the first linear unit in the feature extraction module;

[0009] According to the attention coding matrix and the first representation vector, determine the second representation vector through the second linear unit in the feature extraction module;

[0010] Based on the first representation vector and the second representation vector, train the equivariant attention module in the initial antibody structure prediction model to obtain the antibody structure prediction model.

[0011] The training method of the antibody structure prediction model, wherein the obtaining of the amino acid coding sequence of the antibody sample specifically includes:

[0012] Encoding each amino acid in the amino acid sequence of the antibody sample to obtain an amino acid coding sequence.

[0013] The training method of the antibody structure prediction model, wherein the determining of the high-dimensional vector representation of the amino acid sequence of the antibody sample based on the amino acid coding sequence specifically is:

[0014] Inputting the amino acid coding sequence into the language model in the initial antibody structure prediction model, and outputting the high-dimensional vector representation of the amino acid sequence of the antibody sample through the language model.

[0015] The training method of the antibody structure prediction model, wherein the vector dimension of the first representation vector and the vector dimension of the second representation vector are both smaller than the vector dimension of the high-dimensional vector representation.

[0016] The training method of the antibody structure prediction model, wherein the inputting of the high-dimensional vector representation into the feature extraction module in the initial antibody structure prediction model and the determining of the attention coding matrix through the attention unit in the feature extraction module specifically include:

[0017] Inputting the high-dimensional vector representation into the first attention encoder in the attention unit, and outputting an intermediate attention coding matrix through the first attention encoder;

[0018] Inputting the intermediate attention coding matrix into the second attention encoder in the attention unit, and outputting the attention coding matrix through the second attention encoder.

[0019] The training method of the antibody structure prediction model, wherein the determining of the first representation vector through the first linear unit in the feature extraction module according to the high-dimensional vector representation and the attention coding matrix specifically includes:

[0020] Fusing the high-dimensional vector representation with the attention coding matrix to obtain a first fusion matrix;

[0021] Inputting the first fusion matrix into the first linear layer in the first linear unit, and outputting an intermediate high-dimensional vector representation through the first linear layer;

[0022] Inputting the intermediate high-dimensional vector representation into the second linear layer in the first linear unit, and outputting the first representation vector through the second linear layer.

[0023] The training method of the antibody structure prediction model, wherein, the step of determining the second representation vector through the second linear unit in the feature extraction module according to the attention encoding matrix and the first representation vector specifically includes:

[0024] Fuse the attention encoding matrix and the first representation vector to obtain a second fusion matrix;

[0025] Input the second fusion matrix into the third linear layer in the second linear unit, and output the second representation vector through the third linear layer.

[0026] The training method of the antibody structure prediction model, wherein, the step of training the equivariant attention module in the initial antibody structure prediction model based on the first representation vector and the second representation vector to obtain the antibody structure prediction model specifically includes:

[0027] Input the first representation vector and the second representation vector into the equivariant attention module in the initial antibody structure prediction model, and determine the predicted antibody structure of the antibody sample through the equivariant attention module;

[0028] Train the equivariant attention module in the initial antibody structure prediction model based on the predicted antibody structure and the true antibody structure of the antibody sample to obtain the antibody structure prediction model.

[0029] The training method of the antibody structure prediction model, wherein, the step of inputting the first representation vector and the second representation vector into the equivariant attention module in the initial antibody structure prediction model, and determining the predicted antibody structure of the antibody sample through the equivariant attention module specifically includes:

[0030] Obtain the initial antibody structure of the antibody sample, wherein the initial antibody structure includes the initial three-dimensional coordinates of each amino acid in the amino acid sequence of the antibody sample;

[0031] Input the initial antibody structure, the first representation vector and the second representation vector into the equivariant attention unit in the equivariant attention module, and determine the attention weight matrix through the equivariant attention unit;

[0032] Based on the attention weight matrix and the initial antibody structure, determine the predicted three-dimensional coordinates of each amino acid through the third linear unit in the equivariant attention module to obtain the predicted antibody structure of the antibody sample.

[0033] The training method of the antibody structure prediction model, wherein training the equivariant attention module in the initial antibody structure prediction model based on the predicted antibody structure and the true antibody structure of the antibody sample to obtain the antibody structure prediction model specifically includes:

[0034] Construct a frame-aligned point error loss function based on the predicted antibody structure and the true antibody structure of the antibody sample;

[0035] Train the equivariant attention module in the initial antibody structure prediction model based on the frame-aligned point error loss function to obtain the antibody structure prediction model.

[0036] The second aspect of the present application provides an antibody structure prediction method, which uses the antibody structure prediction model constructed based on the training method of the above antibody structure prediction model. The antibody structure prediction method specifically includes:

[0037] Obtain the amino acid sequence of the antibody to be predicted, and determine the amino acid coding sequence of the amino acid sequence;

[0038] Input the amino acid coding sequence into the antibody structure prediction model, and determine the predicted antibody structure of the antibody to be predicted through the antibody structure prediction model, wherein the predicted antibody structure includes the predicted three-dimensional coordinates of each amino acid in the amino acid sequence of the predicted antibody.

[0039] The third aspect of the present application provides a construction device for an antibody structure prediction model, wherein the construction device for the antibody structure prediction model specifically includes:

[0040] An acquisition module, configured to acquire the amino acid coding sequence of the amino acid sequence of the antibody sample, and determine the high-dimensional vector representation of the amino acid sequence of the antibody sample based on the amino acid coding sequence;

[0041] A control module, configured to input the high-dimensional vector representation into the feature extraction module in the initial antibody structure prediction model, and determine the attention coding matrix through the attention unit in the feature extraction module; determine the first representation vector through the first linear unit in the feature extraction module according to the high-dimensional vector representation and the attention coding matrix; determine the second representation vector through the second linear unit in the feature extraction module according to the attention coding matrix and the first representation vector;

[0042] A training module, configured to train the equivariant attention module in the initial antibody structure prediction model based on the first representation vector and the second representation vector to obtain the antibody structure prediction model.

[0043] A fourth aspect of the present application provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement the steps in the training method of the antibody structure prediction model described above.

[0044] A fifth aspect of the present application provides a terminal device, which includes: a processor and a memory;

[0045] A computer-readable program executable by the processor is stored on the memory;

[0046] When the processor executes the computer-readable program, the steps in the training method of the antibody structure prediction model described above are implemented.

[0047] Beneficial effects:

[0048] (1) In the present application, feature extraction is performed through a linear unit. Since the model parameters of the antibody structure prediction model are reduced, on the one hand, the training speed of the antibody structure prediction model is improved, and on the other hand, the storage space required by the antibody structure prediction model can be reduced.

[0049] (2) In the present application, by converting the high-dimensional vector representation into two low-dimensional first representation vectors and second representation vectors, the computational amount for determining the predicted antibody structure based on the first representation vector and the second representation vector can be reduced, thereby improving the training speed of the antibody structure prediction model.

[0050] (3) In the present application, by reducing the model parameters and the computational amount, the storage space and computational resources required by the antibody structure prediction model in the inference stage are reduced, thereby reducing the requirements of the antibody structure prediction model for hardware resources and avoiding the limitations of hardware resources on the use of the antibody structure prediction model. Description of the drawings

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0052] Figure 1 It is a flowchart of the training method of the antibody structure prediction model provided by the embodiment of the present application.

[0053] Figure 2 It is a principle flowchart of a specific example of the feature extraction module in the antibody structure prediction model provided by the embodiment of the present application.

[0054] Figure 3It is a principle flowchart of a specific example of the training method of the antibody structure prediction model provided by the embodiments of the present application.

[0055] Figure 4 It is a schematic structural diagram of the antibody structure prediction model construction device provided by the embodiments of the present application.

[0056] Figure 5 It is a schematic structural diagram of the terminal device provided by the embodiments of the present application. Detailed implementation manners

[0057] The embodiments of the present application provide a training method of an antibody structure prediction model, an antibody structure prediction method and a device. To make the purpose, technical solutions and effects of the present application clearer and more definite, the following further describes the present application in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0058] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0059] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0060] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not mean the sequence of execution, and the execution sequence of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0061] The following further describes the content of the application by describing the embodiments with reference to the accompanying drawings.

[0062] This embodiment provides a training method for an antibody structure prediction model, as Figure 1 shown, the method includes:

[0063] S10. Obtain the amino acid coding sequence of the amino acid sequence of the antibody sample.

[0064] Specifically, the amino acid coding sequence is obtained by encoding the amino acid sequence of the antibody sample. The sequence length of the amino acid coding sequence is equal to the sequence length of the amino acid sequence, and the amino acid coding in the amino acid coding sequence corresponds one-to-one with the amino acids in the amino acid sequence. That is to say, the amino acid coding sequence is obtained by encoding each amino acid in the amino acid sequence. Among them, each amino acid can be encoded as a unique code representation. For example, a single-letter code is used as the code representation of the amino acid (such as A represents alanine, R represents arginine, etc.). The amino acid coding sequence is a one-dimensional sequence with a length of L, where L is the sequence length of the amino acid sequence. Of course, in practical applications, other ways can also be used to represent the code representation of amino acids. For example, numbers can be used as the code representation, or two-dimensional letter vectors can be used as the code representation of amino acids, etc.

[0065] S20. Determine the high-dimensional vector representation of the amino acid sequence of the antibody sample based on the amino acid coding sequence.

[0066] Specifically, the high-dimensional vector representation is a representation matrix, and it includes the high-dimensional vector representations of each amino acid in the amino acid sequence. The high-dimensional vector representation of each amino acid is a vector. That is to say, the high-dimensional vector representation of the amino acid sequence is formed by arranging the high-dimensional vector representations of each amino acid according to their amino acid positions in the amino acid sequence, so that the high-dimensional vector representation of the amino acid sequence of the antibody sample contains rich information about each amino acid in the context of the entire amino acid sequence, providing rich knowledge information for subsequent antibody structure prediction.

[0067] Exemplarily, determining the high-dimensional vector representation of the amino acid sequence of the antibody sample based on the amino acid coding sequence is specifically:

[0068] Input the amino acid coding sequence into the language model in the initial antibody structure prediction model, and output the high-dimensional vector representation of the amino acid sequence of the antibody sample through the language model.

[0069] Specifically, the initial antibody structure prediction model further includes a language model, which is used to determine the high-dimensional vector representation. The input item of the language model is the amino acid coding sequence, and the output item is the high-dimensional vector representation of the amino acid sequence of the antibody sample. That is, the language model learns the context representation of each amino acid from the amino acid coding sequence to obtain the high-dimensional vector representation of the amino acid sequence of the antibody sample. For example, if the sequence length of the amino acid sequence is L and the dimension of the high-dimensional vector representation of each amino acid is N, then the dimension of the high-dimensional vector representation of the amino acid sequence of the antibody sample is L*N, where both L and N are positive integers. Preferably, N is a positive integer greater than 1. Of course, in practical applications, the language model may not be included in the initial antibody structure prediction model, but may be two independent models with the initial antibody structure prediction model. The high-dimensional vector representation can be obtained first through the language model, and then the high-dimensional vector representation can be used as the input item of the initial antibody structure prediction model.

[0070] Furthermore, the language model is a pre-trained language model, and the language model may include an encoder. The encoder may include multiple layers of self-attention mechanisms and feed-forward neural networks, etc. The language model captures complex patterns and long-range dependencies in the amino acid sequence to generate context-related feature representations (i.e., high-dimensional vector representations) for the amino acid sequence of the antibody sample. For example, the language model may adopt the ESM model or the ProtBERT model, etc. By adopting the pre-trained language model in the embodiments of the present application, the high-dimensional vector representation can be directly extracted through the language model, which can reduce the model parameters to be trained in the initial antibody structure prediction model, thereby improving the training speed of the antibody structure prediction model. Of course, in practical applications, the language model may also adopt a non-pre-trained language model, or other network models may be used to extract the high-dimensional vector representation. For example, a trained deep learning model, etc.

[0071] S30. Input the high-dimensional vector representation into the feature extraction module in the initial antibody structure prediction model. Determine the attention encoding matrix through the attention unit in the feature extraction module. Determine the first representation vector through the first linear unit in the feature extraction module according to the high-dimensional vector representation and the attention encoding matrix, and determine the second representation vector through the second linear unit in the feature extraction module according to the attention encoding matrix and the first representation vector.

[0072] Specifically, the input item of the feature extraction module is the high-dimensional vector representation. When the high-dimensional vector representation is extracted by the language model and the language model is included in the initial antibody structure prediction model, such as Figure 2As shown, the language model is connected to the feature extraction module. The high-dimensional vector representation extracted by the language model is input into the feature extraction module, and the feature extraction module learns the high-dimensional vector representation to obtain the first representation vector and the second representation vector of the amino acid sequence. Among them, both the first representation vector and the second representation vector are representation matrices and both contain the context information of the amino acid in the amino acid sequence.

[0073] As Figure 3 As shown, the feature extraction module includes an attention unit, a first linear unit, and a second linear unit. Among them, the attention unit is respectively connected to the first linear unit and the second linear unit. The input item of the attention unit is the high-dimensional vector representation, and the output item of the attention unit is the attention encoding matrix; the input item of the first linear unit includes the high-dimensional vector representation and the attention encoding matrix, and the output item of the first linear unit is the first representation vector; the input item of the second linear unit includes the attention encoding matrix and the first representation vector, and the output item of the second linear unit is the second representation vector. In the embodiment of the present application, both the first linear unit and the second linear unit have a simple network structure and carry a small number of model parameters, and can well learn the knowledge information carried by the high-dimensional vector representation, reducing the model parameters of the antibody structure prediction model well while ensuring the model performance. At the same time, it can also reduce the computing resources and time required in the training process from the initial antibody structure prediction model to the antibody structure prediction model, and improve the training speed of the antibody structure prediction model.

[0074] Exemplarily, the inputting the high-dimensional vector representation into the feature extraction module of the initial antibody structure prediction model and determining the attention encoding matrix through the attention unit in the feature extraction module specifically includes:

[0075] Input the high-dimensional vector representation into the first attention encoder in the attention unit, and output an intermediate attention encoding matrix through the first attention encoder;

[0076] Input the intermediate attention encoding matrix into the second attention encoder in the attention unit, and output the attention encoding matrix through the second attention encoder.

[0077] Specifically, as Figure 3As shown, the attention unit includes a first attention encoder and a second attention encoder. The first attention encoder and the second attention encoder are cascaded. The first attention encoder is respectively connected to a first linear unit and a second linear unit. Among them, the input item of the first attention encoder is a high-dimensional vector representation, the input item of the second attention encoder is the intermediate attention encoding matrix output by the first attention encoder, and the output item of the second attention encoder is the attention encoding matrix. In this application, by combining the first attention encoder and the second attention encoder, such as the combination used in AlphaFold, the amino acid information contained in the high-dimensional vector representation and the context information of the amino acids in the amino acid sequence can be better learned. Of course, in practical applications, the attention unit may also include only one attention encoder, or may include more than two attention encoders. For example, it includes 3 attention encoders, etc.

[0078] Furthermore, the intermediate attention encoding matrix is obtained by encoding the high-dimensional vector representation through the first attention encoder. Among them, the matrix dimension of the intermediate attention encoding matrix is the same as that of the high-dimensional vector representation. For example, if the matrix dimension of the high-dimensional vector representation is L*N, then the matrix dimension of the intermediate attention encoding matrix is L*N. The attention encoding matrix is obtained by encoding the intermediate attention encoding matrix through the second attention encoder. Among them, the matrix dimension of the attention encoding matrix is smaller than that of the intermediate attention encoding matrix. For example, the matrix dimension of the intermediate attention encoding matrix is L*N, and the matrix dimension of the attention encoding matrix is L*N1, where N > N1, and both N1 and N are positive integers.

[0079] It should be noted that in practical applications, the matrix dimension of the intermediate attention encoding matrix may also be different from that of the high-dimensional vector representation, and the matrix dimension of the attention encoding matrix may also be equal to or greater than that of the intermediate attention encoding matrix. In the implementation of this application, encoding the intermediate attention encoding matrix into an attention encoding matrix with a smaller dimension can reduce the computational complexity of determining the first representation vector and the second representation vector based on the attention encoding matrix, thereby further reducing the computational complexity of the antibody structure prediction model.

[0080] Exemplarily, the specific process of determining the first representation vector through the first linear unit in the feature extraction module according to the high-dimensional vector representation and the attention encoding matrix includes:

[0081] Fuse the high-dimensional vector representation and the attention encoding matrix to obtain a first fusion matrix;

[0082] Input the first fusion matrix into the first linear layer in the first linear unit, and output an intermediate high-dimensional vector representation through the first linear layer;

[0083] Input the intermediate high-dimensional vector representation into the second linear layer in the first linear unit, and output a first representation vector through the second linear layer.

[0084] Specifically, the first fusion matrix is obtained by fusing the high-dimensional vector representation and the attention encoding matrix. The corresponding fusion method of the first fusion matrix can adopt matrix addition, matrix splicing, or matrix weighting, etc. In the embodiments of the present application, the corresponding fusion method of the first fusion matrix adopts the matrix addition method. That is to say, the high-dimensional vector representation and the attention encoding matrix are fused by matrix addition to obtain the first fusion matrix. In addition, in the embodiments of the present application, since the matrix dimension of the high-dimensional vector representation is greater than the matrix dimension of the attention encoding matrix, when fusing by matrix addition, the attention encoding matrix can be filled first so that the matrix dimension of the filled attention encoding matrix is equal to the matrix dimension of the high-dimensional vector representation, and then the attention encoding matrix and the high-dimensional vector representation are fused by matrix addition. The matrix dimension of the fused first fusion matrix is equal to the matrix dimension of the high-dimensional vector representation. Of course, in practical applications, when the matrix dimension of the attention encoding matrix is equal to the high-dimensional vector representation, fusion can be directly performed by matrix addition; when the matrix dimension of the attention encoding matrix is greater than the high-dimensional vector representation, the high-dimensional vector representation can be filled first, and the matrix dimension of the fused first fusion matrix is equal to the matrix dimension of the attention encoding matrix.

[0085] The first linear unit is used to extract features based on the attention encoding matrix, and a first representation vector for reflecting the feature information of the amino acid sequence is extracted through the first linear unit. The first representation vector is a feature matrix, and the matrix dimension of the first representation vector is less than the matrix dimension of the high-dimensional vector representation. For example, the matrix dimension of the high-dimensional vector is represented as L*N, and the matrix dimension of the first representation vector is L*N1, where N1 < N, and both N1 and N are positive integers. Specifically, as Figure 3 shown, the first linear unit may include a first linear layer and a second linear layer. The first linear layer is cascaded with the second linear layer, and the first linear layer is connected to the second attention encoder. Among them, the input item of the first linear layer is the first fusion matrix obtained by fusing the high-dimensional vector representation and the attention encoding matrix, and the input item of the second linear layer is the intermediate high-dimensional vector representation output through the first linear layer. In the embodiments of the present application, the first linear layer and the second linear layer are adopted for feature extraction, reducing the model parameters included in the feature extraction module. On the one hand, this can reduce the storage space required by the antibody structure prediction model, and on the other hand, it can reduce the number of parameters that need to be trained by the antibody structure prediction model, improving the training speed of the antibody structure prediction model.

[0086] It should be noted that the structures of the first linear layer and the second linear layer in the first linear unit are the same, and their parameters can be the same or different. For example, the first linear layer and the second linear layer are two identical linear layers, or the first linear layer and the second linear layer are two linear layers used in combination, such as the combination used in AlexNet, that is, the parameters of the first linear layer and the second linear layer are different. In addition, in practical applications, the first linear unit may also include only 1 layer of linear layer, or may include more than 2 layers of linear layers, for example, 3 layers, 4 layers, etc. In the embodiment of the present application, the first linear unit includes 2 layers of linear layers, and feature extraction is performed through 2 layers of linear layers, so as to minimize model parameters while extracting limited feature information.

[0087] Exemplarily, determining the second representation vector according to the attention encoding matrix and the first representation vector through the second linear unit in the feature extraction module specifically includes:

[0088] Fusing the attention encoding matrix and the first representation vector to obtain a second fusion matrix;

[0089] Inputting the second fusion matrix into the third linear layer in the second linear unit, and outputting the second representation vector through the third linear layer.

[0090] Specifically, the second fusion matrix is obtained by fusing the first representation vector and the attention encoding matrix, and the corresponding fusion method of the second fusion matrix can be matrix addition, matrix splicing, or matrix weighting, etc. In the embodiment of the present application, the corresponding fusion method of the second fusion matrix adopts the matrix addition method, that is, the first representation vector and the attention encoding matrix are fused by matrix addition to obtain the second fusion matrix. In addition, in the embodiment of the present application, the matrix dimension of the first representation vector is equal to the matrix dimension of the attention encoding matrix, both of which are L*N1, and the first representation vector and the attention encoding matrix can be directly fused by matrix addition. Of course, in practical applications, when the matrix dimensions of the first representation vector and the attention encoding matrix are different, padding can be performed first and then matrix addition fusion can be performed on the first representation vector and the attention encoding matrix.

[0091] The second linear unit is used to extract features based on the attention encoding matrix, and a second representation vector for reflecting the feature information of the amino acid sequence is extracted through the second linear unit. Among them, the second representation vector is a feature matrix, and the matrix dimensions of the second representation vector are all smaller than those of the high-dimensional vector representation. For example, the matrix dimension of the high-dimensional vector is represented as L*N, the matrix dimension of the first representation vector is L*N2, N2 < N, and both N2 and N are positive integers. Specifically, the second linear unit may include a third linear layer, and the third linear layer is respectively connected to the second linear layer and the second attention encoder. Among them, the input item of the third linear layer is the second fusion matrix obtained by fusing the attention encoding matrix and the first representation vector. By using the third linear layer to extract features in the embodiments of the present application, the model parameters included in the feature extraction module can be reduced. On the one hand, this can reduce the storage space required by the antibody structure prediction model, and on the other hand, it can reduce the number of parameters that need to be trained in the antibody structure prediction model, thereby improving the training speed of the antibody structure prediction model.

[0092] It should be noted that the third linear layer in the second linear unit may be the same as the first linear layer or the second linear layer in the first linear unit, or may be different from both the first linear layer and the second linear layer in the first linear unit. Among them, "different" means the same structure but different parameters, and "same" means the same structure and parameters. For example, the third linear layer may be AlexNet. In addition, in practical applications, the second linear unit may also include multiple linear layers, such as 2 layers, 3 layers, etc. In the embodiments of the present application, the second linear unit includes 1 layer of linear layer, and 1 layer of linear layer is used to extract features, which can minimize the model parameters while extracting limited feature information.

[0093] S40. Based on the first representation vector and the second representation vector, train the equivariant attention module in the initial antibody structure prediction model to obtain an antibody structure prediction model.

[0094] Specifically, the antibody structure prediction model is obtained by training the equivariant attention module in the initial antibody structure prediction model. The model structure of the antibody structure prediction model is the same as that of the initial antibody structure prediction model. The difference between the two is that the model parameters of the initial antibody structure prediction model are initial parameters, and the model parameters of the antibody structure prediction model are model parameters after training.

[0095] The initial antibody structure prediction model may include a feature extraction module and an equivariant attention module. In addition, when extracting features through the feature extraction module, a high-dimensional vector representation of the amino acid sequence of the antibody sample will be obtained first, and the high-dimensional vector representation can be extracted by a pre-trained language model. For this reason, such as Figure 2As shown, the initial antibody structure prediction model may further include a language model, that is, the initial antibody structure prediction model may include a language model, a feature extraction module, and an equivariant attention module, where the language model is connected to the feature extraction module, and the feature extraction module is connected to the equivariant attention model. Of course, in practical applications, the language model and the initial antibody structure prediction model may be two independent models. When predicting the antibody structure, the language model and the initial antibody structure prediction model can be used jointly, etc. In addition, it should be noted that when obtaining the antibody structure prediction model, in addition to training the equivariant attention module, the feature extraction module can also be trained, and when the initial antibody structure prediction model includes a language model, the language model, the feature extraction module, and the equivariant attention module can be trained.

[0096] The equivariant attention module is used to predict the predicted antibody structure of the antibody sample based on the feature information provided by the first representation vector and the second representation vector. Among them, the equivariant attention module is configured with an equivariant attention mechanism, and the predicted antibody structure includes the predicted three-dimensional coordinates of each amino acid in the amino acid sequence of the antibody sample. That is to say, the equivariant attention module predicts the predicted three-dimensional coordinates of each amino acid in the amino acid sequence of the antibody sample through the equivariant attention mechanism to obtain the predicted antibody structure, and then trains the equivariant attention module based on the predicted antibody structure to obtain the antibody structure prediction model. Among them, the equivariant attention mechanism is a special neural network structure. When using the attention mechanism principle to process data in three-dimensional space, it maintains rotation and translation invariance, which can avoid the influence of the direction and position of the protein structure in space on the predicted antibody structure during the antibody structure prediction process, so as to improve the accuracy of the predicted antibody structure.

[0097] In one implementation manner, the training of the equivariant attention module in the initial antibody structure prediction model based on the first representation vector and the second representation vector to obtain the antibody structure prediction model specifically includes:

[0098] S41: Input the first representation vector and the second representation vector into the equivariant attention module in the initial antibody structure prediction model, and determine the predicted antibody structure of the antibody sample through the equivariant attention module;

[0099] S42: Train the equivariant attention module in the initial antibody structure prediction model based on the predicted antibody structure and the true antibody structure of the antibody sample to obtain the antibody structure prediction model.

[0100] Specifically, in step S41, the equivariant attention module processes the initial antibody structure of the antibody sample through the equivariant attention mechanism and maintains rotational and translational invariance to obtain the predicted antibody structure. That is, the equivariant attention module learns the context representation between each amino acid in the amino acid sequence of the antibody sample from the first representation vector and the second representation vector to determine the predicted antibody structure of the antibody sample. Among them, when determining the predicted antibody structure through the equivariant attention module, the equivariant attention module can first learn the context representation between each amino acid in the amino acid sequence based on the first representation vector and the second representation vector, and then update the initial antibody structure of the antibody sample based on the learned context representation and maintain the equivariant property to obtain the predicted antibody structure.

[0101] Exemplarily, the inputting of the first representation vector and the second representation vector into the equivariant attention module in the initial antibody structure prediction model, and determining the predicted antibody structure of the antibody sample through the equivariant attention module specifically includes:

[0102] S411. Obtain the initial antibody structure of the antibody sample, where the initial antibody structure includes the initial three-dimensional coordinates of each amino acid in the amino acid sequence of the antibody sample;

[0103] S412. Input the initial antibody structure, the first representation vector, and the second representation vector into the equivariant attention unit in the equivariant attention module, and determine the attention weight matrix through the equivariant attention unit;

[0104] S413. Based on the attention weight matrix and the initial antibody structure, determine the predicted three-dimensional coordinates of each amino acid through the third linear unit in the equivariant attention module to obtain the predicted antibody structure of the antibody sample.

[0105] Specifically, the initial antibody structure includes the initial three-dimensional coordinates of each amino acid in the amino acid sequence of the antibody sample, and the initial antibody structure can be obtained by initializing the three-dimensional coordinate representation of each amino acid in the amino acid sequence. Among them, the initialization method can be initialized by using the random coordinate method or the default coordinate method, etc.

[0106] After obtaining the initial antibody structure, the initial antibody structure, the first representation vector, and the second representation vector are input into the equivariant attention mechanism in the equivariant attention module. For each amino acid in the amino acid sequence, the equivariant attention mechanism first calculates the relationship information between the three-dimensional coordinates of the amino acid and the other amino acids in the amino acid sequence except this amino acid. The relationship information includes relative distance and direction information. Then, the attention weights between the amino acid and the other amino acids in the amino acid sequence except this amino acid are calculated based on the calculated relationship information, and the attention weight matrix of the amino acid sequence is obtained according to the attention weights between each amino acid and the other amino acids in the amino acid sequence except this amino acid.

[0107] The third linear unit is used to update the initial antibody structure based on the attention weights to obtain the predicted three-dimensional coordinates of the amino acids. The third linear unit maintains the equivariant property while updating the initial antibody structure, that is, the initial antibody structure is rotated or translated as a whole, and the relative relationship between the amino acids remains unchanged. The third linear unit may include one or more linear layers. In the embodiment of the present application, the third linear unit includes one linear layer, denoted as the fourth linear layer such as AlexNet. The attention weight matrix and the initial antibody structure are input into the fourth linear layer, and the predicted antibody structure is output through the fourth linear layer.

[0108] Further, in step S42, the predicted antibody structure includes the predicted three-dimensional coordinates of each amino acid in the amino acid sequence of the antibody sample, and the true antibody structure is used as the true label of the antibody sample to be used as the supervised data for the training process of the antibody structure prediction model to supervise the training process of the antibody sample prediction model. The true antibody structure includes the true three-dimensional coordinates of each amino acid in the amino acid sequence of the antibody sample. That is, based on the true three-dimensional coordinates of each amino acid, the predicted three-dimensional coordinates of each amino acid in the predicted antibody structure are supervised to optimize the model parameters in the initial antibody structure prediction model, so that the predicted antibody structure determined by the antibody structure prediction model obtained by training the initial antibody structure prediction model can be closer to the true antibody structure.

[0109] Exemplarily, training the initial antibody structure prediction model based on the predicted antibody structure and the true antibody structure of the antibody sample to obtain an antibody structure prediction model specifically includes:

[0110] Construct a frame-aligned point error loss function based on the predicted antibody structure and the true antibody structure of the antibody sample;

[0111] Train the equivariant attention module in the initial antibody structure prediction model based on the frame-aligned point error loss function to obtain an antibody structure prediction model.

[0112] Specifically, the frame alignment point error loss function reflects and measures the error between the predicted antibody structure and the true antibody structure. Supervising the training of the antibody structure prediction model is achieved by controlling the frame alignment point error loss function. That is, when training the equivariant attention module in the initial antibody structure prediction model based on the frame alignment point error loss function, the end condition of the training is the convergence of the frame alignment point error loss function.

[0113] Furthermore, when training the equivariant attention module in the initial antibody structure prediction model based on the frame alignment point error loss function, if the frame alignment point error loss function does not meet the preset convergence condition, the initial antibody structure is re-obtained, and a new predicted antibody structure is re-determined by the equivariant attention module based on the re-obtained initial antibody structure, the first representation vector, and the second representation vector. This process is repeated until the frame alignment point error loss function meets the preset convergence condition, at which point the antibody structure prediction model is obtained. Of course, in practical applications, when the training data set corresponding to the antibody structure prediction model includes multiple antibody samples, the above steps S10 - S40 can be performed for each antibody sample to obtain the antibody structure prediction model.

[0114] In summary, the embodiment of the present application provides a training method for an antibody structure prediction model. The training method for the antibody structure prediction model specifically includes: obtaining the amino acid coding sequence of the amino acid sequence of an antibody sample, and determining the high-dimensional vector representation of the amino acid sequence of the antibody sample based on the amino acid coding sequence; inputting the high-dimensional vector representation into the feature extraction module in the initial antibody structure prediction model, determining the attention encoding matrix through the attention unit in the feature extraction module, determining the first representation vector through the first linear unit in the feature extraction module according to the high-dimensional vector representation and the attention encoding matrix, and determining the second representation vector through the second linear unit in the feature extraction module according to the attention encoding matrix and the first representation vector; training the equivariant attention module in the initial antibody structure prediction model based on the first representation vector and the second representation vector to obtain the antibody structure prediction model. The present application uses linear units for feature extraction to obtain the first representation vector and the second representation vector, reducing the model parameters of the antibody structure prediction model. On the one hand, this can reduce the storage space required by the antibody structure prediction model, and on the other hand, it can reduce the number of parameters that need to be trained in the antibody structure prediction model, thereby improving the training speed of the antibody structure prediction model. At the same time, by reducing the model parameters and computational complexity, the present application reduces the storage space and computational resources required by the antibody structure prediction model during the inference stage, thereby reducing the requirements of the antibody structure prediction model for hardware resources and avoiding the limitations of hardware resources on the use of the antibody structure prediction model.

[0115] Based on the above training method of the antibody structure prediction model, this embodiment provides a device for constructing an antibody structure prediction model, as Figure 4 shown. The device for constructing the antibody structure prediction model specifically includes:

[0116] An acquisition module 100, configured to acquire the amino acid coding sequence of the amino acid sequence of an antibody sample, and determine the high-dimensional vector representation of the amino acid sequence of the antibody sample based on the amino acid coding sequence;

[0117] A control module 200, configured to input the high-dimensional vector representation into the feature extraction module in the initial antibody structure prediction model, and determine the attention coding matrix through the attention unit in the feature extraction module; determine the first representation vector through the first linear unit in the feature extraction module according to the high-dimensional vector representation and the attention coding matrix; determine the second representation vector through the second linear unit in the feature extraction module according to the attention coding matrix and the first representation vector;

[0118] A training module 300, configured to train the equivariant attention module in the initial antibody structure prediction model based on the first representation vector and the second representation vector to obtain an antibody structure prediction model.

[0119] Based on the above training method of the antibody structure prediction model, this embodiment provides an antibody structure prediction method, using the antibody structure prediction model constructed by the training method of the antibody structure prediction model described in the above embodiment. The antibody structure prediction method specifically includes:

[0120] Acquire the amino acid sequence of the antibody to be predicted, and determine the amino acid coding sequence of the amino acid sequence;

[0121] Input the amino acid coding sequence into the antibody structure prediction model, and determine the predicted antibody structure of the antibody to be predicted through the antibody structure prediction model, where the predicted antibody structure includes the predicted coordinates of each amino acid in the amino acid sequence of the predicted antibody.

[0122] Based on the above training method of the antibody structure prediction model, this embodiment provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the training method of the antibody structure prediction model described in the above embodiment.

[0123] Based on the above training method of the antibody structure prediction model, this application also provides a terminal device, as Figure 5As shown in the figure, it includes at least one processor 20; a display screen 21; and a memory 22, and may further include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22 and the communication interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above-mentioned embodiment.

[0124] In addition, when the logical instructions in the above-mentioned memory 22 can be implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium.

[0125] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as the program instructions or modules corresponding to the method in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions or modules stored in the memory 22, that is, to implement the method in the above-mentioned embodiment.

[0126] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc that can store program codes can also be a transient storage medium.

[0127] In addition, the specific processes of loading and executing multiple instructions by the above-mentioned storage medium and the instruction processor in the terminal device have been described in detail in the above method, and will not be repeated here one by one.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A training method for an antibody structure prediction model, characterized in that The training method of the antibody structure prediction model specifically includes: Obtain the amino acid coding sequence of the amino acid sequence of the antibody sample, and determine the high-dimensional vector representation of the amino acid sequence of the antibody sample based on the amino acid coding sequence; Input the high-dimensional vector representation into the feature extraction module in the initial antibody structure prediction model, and determine the attention coding matrix through the attention unit in the feature extraction module; Determine the first representation vector through the first linear unit in the feature extraction module according to the high-dimensional vector representation and the attention coding matrix; Determine the second representation vector through the second linear unit in the feature extraction module according to the attention coding matrix and the first representation vector; Train the equivariant attention module in the initial antibody structure prediction model based on the first representation vector and the second representation vector to obtain the antibody structure prediction model.

2. The training method of the antibody structure prediction model according to claim 1, characterized in that The obtaining of the amino acid coding sequence of the amino acid sequence of the antibody sample specifically includes: Encode each amino acid in the amino acid sequence of the antibody sample to obtain the amino acid coding sequence.

3. The training method of the antibody structure prediction model according to claim 1, wherein The determining of the high-dimensional vector representation of the amino acid sequence of the antibody sample based on the amino acid coding sequence is specifically: Input the amino acid coding sequence into the language model in the initial antibody structure prediction model, and output the high-dimensional vector representation of the amino acid sequence of the antibody sample through the language model.

4. The training method of the antibody structure prediction model according to claim 1, wherein The vector dimension of the first representation vector and the vector dimension of the second representation vector are both smaller than the vector dimension of the high-dimensional vector representation.

5. The training method of the antibody structure prediction model according to any one of claims 1-4, characterized in that, The inputting of the high-dimensional vector representation into the feature extraction module in the initial antibody structure prediction model and the determining of the attention coding matrix through the attention unit in the feature extraction module specifically include: Input the high-dimensional vector representation into the first attention encoder in the attention unit, and output the intermediate attention coding matrix through the first attention encoder; Input the intermediate attention coding matrix into the second attention encoder in the attention unit, and output the attention coding matrix through the second attention encoder.

6. The training method of the antibody structure prediction model according to any one of claims 1-4, characterized in that, The determining of the first representation vector through the first linear unit in the feature extraction module according to the high-dimensional vector representation and the attention coding matrix specifically includes: Fuse the high-dimensional vector representation and the attention coding matrix to obtain the first fusion matrix; Input the first fusion matrix into the first linear layer in the first linear unit, and output the intermediate high-dimensional vector representation through the first linear layer; Input the intermediate high-dimensional vector representation into the second linear layer in the first linear unit, and output the first representation vector through the second linear layer.

7. The training method of the antibody structure prediction model according to any one of claims 1-4, characterized in that, The determining of the second representation vector through the second linear unit in the feature extraction module according to the attention coding matrix and the first representation vector specifically includes: Fuse the attention coding matrix and the first representation vector to obtain the second fusion matrix; Input the second fusion matrix into the third linear layer in the second linear unit, and output the second representation vector through the third linear layer.

8. The training method of the antibody structure prediction model according to any one of claims 1-4, characterized in that, Training the equivariant attention module in the initial antibody structure prediction model based on the first representation vector and the second representation vector to obtain the antibody structure prediction model specifically includes: Inputting the first representation vector and the second representation vector into the equivariant attention module in the initial antibody structure prediction model, and determining the predicted antibody structure of the antibody sample through the equivariant attention module, where the structure of the predicted antibody includes the predicted three-dimensional coordinates of each amino acid in the amino acid sequence of the antibody sample; Training the equivariant attention module in the initial antibody structure prediction model based on the predicted antibody structure and the true antibody structure of the antibody sample to obtain the antibody structure prediction model, where the true antibody structure includes the true three-dimensional coordinates of each amino acid in the amino acid sequence of the antibody sample.

9. The training method of the antibody structure prediction model according to claim 8, characterized in that The step of inputting the first representation vector and the second representation vector into the equivariant attention module in the initial antibody structure prediction model and determining the predicted antibody structure of the antibody sample through the equivariant attention module specifically includes: Obtaining the initial antibody structure of the antibody sample, where the initial antibody structure includes the initial three-dimensional coordinates of each amino acid in the amino acid sequence of the antibody sample; Inputting the initial antibody structure, the first representation vector, and the second representation vector into the equivariant attention unit in the equivariant attention module, and determining the attention weight matrix through the equivariant attention unit; Based on the attention weight matrix and the initial antibody structure, determining the predicted three-dimensional coordinates of each amino acid through the third linear unit in the equivariant attention module to obtain the predicted antibody structure of the antibody sample.

10. The training method of the antibody structure prediction model according to claim 8, wherein Training the equivariant attention module in the initial antibody structure prediction model based on the predicted antibody structure and the true antibody structure of the antibody sample to obtain the antibody structure prediction model specifically includes: Constructing a frame-aligned point error loss function based on the predicted antibody structure and the true antibody structure of the antibody sample; Training the equivariant attention module in the initial antibody structure prediction model based on the frame-aligned point error loss function to obtain the antibody structure prediction model.

11. A method for predicting antibody structure, characterized in that, Using the antibody structure prediction model constructed by the training method of the antibody structure prediction model according to any one of claims 1-9, the antibody structure prediction method specifically includes: Obtaining the amino acid sequence of the antibody to be predicted and determining the amino acid coding sequence of the amino acid sequence; Inputting the amino acid coding sequence into the antibody structure prediction model, and determining the predicted antibody structure of the antibody to be predicted through the antibody structure prediction model, where the predicted antibody structure includes the predicted three-dimensional coordinates of each amino acid in the amino acid sequence of the predicted antibody.

12. An apparatus for constructing an antibody structure prediction model, characterized in that, The apparatus for constructing the antibody structure prediction model specifically includes: An acquisition module, configured to acquire the amino acid coding sequence of the amino acid sequence of the antibody sample and determine the high-dimensional vector representation of the amino acid sequence of the antibody sample based on the amino acid coding sequence; A control module, configured to input the high-dimensional vector representation into a feature extraction module in an initial antibody structure prediction model, and determine an attention encoding matrix through an attention unit in the feature extraction module; determine a first representation vector through a first linear unit in the feature extraction module according to the high-dimensional vector representation and the attention encoding matrix; determine a second representation vector through a second linear unit in the feature extraction module according to the attention encoding matrix and the first representation vector. A training module, configured to train an equivariant attention module in the initial antibody structure prediction model based on the first representation vector and the second representation vector to obtain an antibody structure prediction model.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the training method of the antibody structure prediction model according to any one of claims 1-10, and / or to implement the steps in the prediction method of the antibody structure according to claim 11.

14. A terminal device, characterized in that, Comprising: A processor and a memory; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, it implements the steps in the training method of the antibody structure prediction model according to any one of claims 1-10, and / or implements the steps in the prediction method of the antibody structure according to claim 11.

Citation Information

Patent Citations

  • Ligand generation model, training method thereof and ligand generation method

    CN117457061A

  • Model training method, antibacterial peptide prediction method and system

    CN117875444A