Training method for antigen-antibody affinity prediction model and antibody screening method
By training the antigen antibody affinity prediction model, using three-dimensional feature map extraction and parameter assignment optimization, the problem of high cost and low accuracy in antibody drug screening is solved, and efficient and accurate antibody screening is achieved.
Patent Information
- Application Number
- CN202311323310.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-10-12
AI Technical Summary
The prior art has problems of high cost and low accuracy in antibody drug screening, especially in the generalization of antigen-antibody affinity prediction during large-scale screening.
By training the antigen-antibody affinity prediction model, the parameter assignment of the model and the initial affinity prediction model is extracted using the three-dimensional feature map of the trained target complex, and the model training is performed by combining wild-type and mutant antigen-antibody complex samples to optimize the training effect of the affinity prediction model and improve prediction accuracy and generalization.
It improves the efficiency and accuracy of antigen-antibody affinity prediction, reduces the cost of antibody screening, enhances the accuracy and practicality of antibody screening, and optimizes the antibody screening method.
Smart Images

Figure CN117334247B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, particularly to artificial intelligence fields such as natural language processing and deep learning, and is applied to the field of antibody screening. Background Art
[0002] With the development of technology, large-scale screening of macromolecular antibody drugs has become increasingly important in the field of drug design. Optionally, large-scale antibody drug screening can be performed through high-throughput wet experimental techniques, but the cost is high and the accuracy is poor.
[0003] In related technologies, large-scale screening of macromolecular antibody drugs can also be achieved through the prediction of the affinity between antigens and antibodies. In this scenario, an energy model obtained by learning a small amount of data can be used to evaluate the stability of antigen-antibody complexes, and based on this, the affinity between antigens and antibodies can be predicted and judged, but the generalization is poor. Summary of the Invention
[0004] The present disclosure proposes a training method for an antigen-antibody affinity prediction model and an antibody screening method.
[0005] According to a first aspect of the present disclosure, a training method for an antigen-antibody affinity prediction model is proposed. The method includes: obtaining a trained target complex three-dimensional feature map extraction model and a first model parameter set of the target complex three-dimensional feature map extraction model; obtaining an initial feature map extraction layer of an initial affinity prediction model, and performing parameter assignment on the initial feature map extraction layer based on the first model parameter set to obtain a candidate affinity prediction model to be trained; obtaining a first training sample of the candidate affinity prediction model, where the first training sample includes a first complex sample and a set of mutant complex samples of the first complex sample; inputting the first complex sample and the set of mutant complex samples into the candidate affinity prediction model for model training until the training ends, to obtain a trained target affinity prediction model.
[0006] According to a second aspect of the present disclosure, an antibody screening method is provided. The method includes: obtaining a trained target affinity prediction model, where the target affinity prediction model is obtained based on the training method of the antigen-antibody affinity prediction model proposed in the first aspect above; obtaining a wild-type antigen-antibody complex of a target antigen and a set of candidate mutant antigen-antibody complexes corresponding to the wild-type antigen-antibody complex; inputting the wild-type antigen-antibody complex and the set of candidate mutant antigen-antibody complexes into the target affinity prediction model, and obtaining, through the target affinity prediction model, candidate antigen-antibody affinity change parameters of each candidate mutant antigen-antibody complex in the set of candidate mutant antigen-antibody complexes based on the wild-type antigen-antibody complex; and screening out a target mutant antigen-antibody complex from the set of candidate mutant antigen-antibody complexes according to the candidate antigen-antibody affinity change parameters, so as to obtain a target antibody of the target antigen.
[0007] According to a third aspect of the present disclosure, a training device for an antigen-antibody affinity prediction model is provided. The device includes: a first obtaining module, configured to obtain a trained target three-dimensional complex feature map extraction model and a first model parameter set of the target three-dimensional complex feature map extraction model; an assignment module, configured to obtain an initial feature map extraction layer of an initial affinity prediction model, and perform parameter assignment on the initial feature map extraction layer based on the first model parameter set to obtain a candidate affinity prediction model to be trained; a second obtaining module, configured to obtain a first training sample of the candidate affinity prediction model, where the first training sample includes a first complex sample and a set of mutant complex samples of the first complex sample; and a training module, configured to input the first complex sample and the set of mutant complex samples into the candidate affinity prediction model for model training until the training ends, so as to obtain a trained target affinity prediction model.
[0008] According to a fourth aspect of the present disclosure, an antibody screening device is provided. The device includes: a third acquisition module configured to acquire a trained target affinity prediction model, where the target affinity prediction model is obtained based on the training device for the antigen-antibody affinity prediction model proposed in the above-mentioned third aspect; a fourth acquisition module configured to acquire a wild-type antigen-antibody complex of a target antigen and a set of candidate mutant antigen-antibody complexes corresponding to the wild-type antigen-antibody complex; a prediction module configured to input the wild-type antigen-antibody complex and the set of candidate mutant antigen-antibody complexes into the target affinity prediction model, and obtain candidate antigen-antibody affinity change parameters of each candidate mutant antigen-antibody complex in the set of candidate mutant antigen-antibody complexes based on the wild-type antigen-antibody complex through the target affinity prediction model; and a screening module configured to screen out a target mutant antigen-antibody complex from the set of candidate mutant antigen-antibody complexes according to the candidate antigen-antibody affinity change parameters, so as to obtain a target antibody of the target antigen.
[0009] According to a fifth aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the training method of the antigen-antibody affinity prediction model proposed in the above-mentioned first aspect and / or the antibody screening method proposed in the above-mentioned second aspect.
[0010] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the training method of the antigen-antibody affinity prediction model proposed in the above-mentioned first aspect and / or the antibody screening method proposed in the above-mentioned second aspect.
[0011] According to a seventh aspect of the present disclosure, a computer program product is provided, including a computer program which, when executed by a processor, implements the training method of the antigen-antibody affinity prediction model proposed in the above-mentioned first aspect and / or the antibody screening method proposed in the above-mentioned second aspect.
[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0014] Figure 1Schematic flowchart of a method for training an antigen-antibody affinity prediction model according to an embodiment of the present disclosure;
[0015] Figure 2 Schematic flowchart of a method for training an antigen-antibody affinity prediction model according to another embodiment of the present disclosure;
[0016] Figure 3 Schematic flowchart of a method for training an antigen-antibody affinity prediction model according to another embodiment of the present disclosure;
[0017] Figure 4 Schematic diagram of an antigen-antibody complex according to an embodiment of the present disclosure;
[0018] Figure 5 Schematic flowchart of a method for training an antigen-antibody affinity prediction model according to another embodiment of the present disclosure;
[0019] Figure 6 Schematic flowchart of an antibody screening method according to an embodiment of the present disclosure;
[0020] Figure 7 Schematic flowchart of an antibody screening method according to another embodiment of the present disclosure;
[0021] Figure 8 Schematic diagram of the structure of a training device for an antigen-antibody affinity prediction model according to an embodiment of the present disclosure;
[0022] Figure 9 Schematic diagram of the structure of an antibody screening device according to an embodiment of the present disclosure;
[0023] Figure 10 Schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0024] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0025] Deep learning (DL) is a new research direction in the field of machine learning. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analysis and learning like humans, and to be able to recognize data such as text, images, and sounds.
[0026] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life. So it has a close connection with the research of linguistics, but there are also important differences. Natural language processing does not generally study natural language, but rather aims to develop computer systems that can effectively achieve natural language communication.
[0027] Artificial Intelligence (AI) is a new technical science that studies, develops, and applies theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Since its birth, the theory and technology of artificial intelligence have become increasingly mature, and the application fields have also been continuously expanding. It can be imagined that the future technological products brought by artificial intelligence will be the "containers" of human wisdom. Artificial intelligence can simulate the information process of human consciousness and thinking.
[0028] Figure 1 The flowchart of the training method of the antigen-antibody affinity prediction model according to an embodiment of the present disclosure is shown as Figure 1 shown. The method includes:
[0029] S101, obtain a trained target complex three-dimensional feature map extraction model and a first model parameter set of the target complex three-dimensional feature map extraction model.
[0030] In the embodiments of the present disclosure, the affinity prediction model needs to extract the three-dimensional feature map of the antigen-antibody complex input to the model. In this scenario, a model for extracting the three-dimensional feature map of the antigen-antibody complex to be trained can be obtained and trained to obtain a trained model for extracting the three-dimensional feature map of the antigen-antibody complex, which is marked as the trained target complex three-dimensional feature map extraction model.
[0031] Optionally, an extraction layer for the three-dimensional feature map of the antigen-antibody complex is provided in the affinity prediction model. In this scenario, the model parameters of the trained target complex three-dimensional feature map extraction model can be reused in the extraction layer for the three-dimensional feature map in the affinity prediction model.
[0032] Among them, the set composed of the model parameters of the obtained three-dimensional feature extraction model of the target complex can be marked as the first model parameter set of the three-dimensional feature extraction model of the target complex.
[0033] S102. Obtain the initial feature map extraction layer of the initial affinity prediction model, and perform parameter assignment on the initial feature map extraction layer based on the first model parameter set to obtain a candidate affinity prediction model to be trained.
[0034] In the embodiments of the present disclosure, the initially constructed affinity prediction model can be marked as the initial affinity prediction model, and the feature extraction layer used to extract the three-dimensional feature map of the antigen-antibody complex in the initial affinity prediction model can be marked as the initial feature map extraction layer of the initial affinity prediction model.
[0035] Optionally, after obtaining the first model parameter set of the trained three-dimensional feature extraction model of the target complex, parameter assignment can be performed on the initial feature map extraction layer of the initial affinity prediction model based on the parameters in the first model parameter set. Furthermore, the initial affinity prediction model after parameter assignment to the initial feature map extraction layer can be determined as the candidate affinity prediction model to be trained.
[0036] It can be understood that each parameter in the first model parameter set is reused for the initial feature map extraction layer, and then the candidate affinity prediction model to be trained is obtained.
[0037] S103. Obtain the first training sample of the candidate affinity prediction model, where the first training sample includes the first complex sample and the set of mutant complex samples of the first complex sample.
[0038] In implementation, there is a set of mutant antibodies composed of multiple corresponding mutant antibodies for the wild-type antibody. In this scenario, the antigen-antibody complexes corresponding to the wild-type antibody and the antigen-antibody complexes corresponding to the mutant antibodies can be respectively obtained as the training samples of the candidate affinity prediction model.
[0039] Among them, the antigen-antibody complex corresponding to the wild-type antibody can be marked as the first complex sample.
[0040] And for any wild-type antibody, the antigen-antibody complex of its corresponding mutant antibody can be marked as the mutant complex sample, and then the mutant complex samples of each mutant antibody in the set of mutant antibodies can be obtained.
[0041] As an example, such as Figure 2 shown, Figure 2 In the shown candidate affinity prediction model, it can be based on Figure 2The wild-type antigen-antibody complex, mutant 1 antigen-antibody complex, and mutant 2 antigen-antibody complex shown are used for model training.
[0042] Among them, the mutant antibodies in the mutant 1 antigen-antibody complex and the mutant 2 antigen-antibody complex are the mutant antibodies corresponding to the wild-type antibody in the wild-type antigen-antibody complex.
[0043] In this example, the wild-type antigen-antibody complex can be regarded as Figure 2 the first complex sample of the candidate affinity prediction model shown. Correspondingly, the mutant 1 antigen-antibody complex and the mutant 2 antigen-antibody complex are both regarded as Figure 2 the mutant complex samples of the candidate affinity prediction model shown. Furthermore, the set composed of the mutant 1 antigen-antibody complex and the mutant 2 antigen-antibody complex is labeled as Figure 2 the mutant complex sample set of the candidate affinity prediction model shown.
[0044] S104, input the first complex sample and the mutant complex sample set into the candidate affinity prediction model for model training until the training ends, and obtain the trained target affinity prediction model.
[0045] In the embodiment of the present disclosure, after obtaining the first complex sample and the mutant complex sample set, the first complex sample and the mutant complex sample set corresponding to the first complex sample can be input into the candidate affinity prediction model for model training.
[0046] Among them, the affinity of the antigen-antibody in the first complex sample can be predicted by the candidate affinity prediction model, and the affinity of the antigen-antibody in each mutant complex in the mutant complex sample set can be predicted to obtain the prediction results of the antigen-antibody affinity in the first complex sample and the antigen-antibody affinity of each mutant complex sample in the mutant complex sample set.
[0047] Furthermore, the affinity differences between the antigen-antibody affinities of each mutant complex sample in the mutant complex sample set based on the antigen-antibody affinity of the first complex sample are obtained respectively, and the obtained affinity differences of each mutant complex sample are used as the model output of the candidate affinity prediction model.
[0048] Optionally, obtain the training loss of the candidate affinity prediction model for the current round based on the output result of the candidate affinity prediction model, perform iterative adjustment of the parameters of the candidate affinity prediction model based on the training loss, and return to obtain the next first complex sample and the next set of mutant complex samples of the next first complex sample, and continue to perform model training on the candidate affinity prediction model with adjusted parameters until the training is completed to obtain the trained target affinity prediction model.
[0049] As an example, such as Figure 2 shown, it is possible to Figure 2 input the wild-type antigen-antibody complex, mutant 1 antigen-antibody complex, and mutant 2 antigen-antibody complex shown Figure 2 into the candidate affinity prediction model shown for model training, and obtain the antigen-antibody affinity in the wild-type antigen-antibody complex, the antigen-antibody affinity in the mutant 1 antigen-antibody complex, and the antigen-antibody affinity in the mutant 2 antigen-antibody complex through the Figure 2 candidate affinity prediction model shown.
[0050] Furthermore, respectively obtain the antigen-antibody affinity change parameters of the mutant 1 antigen-antibody complex based on the wild-type antigen-antibody complex and the antigen-antibody affinity change parameters of the mutant 2 antigen-antibody complex based on the wild-type antigen-antibody complex through the Figure 2 candidate affinity prediction model shown, and then obtain the training loss of the candidate affinity prediction model based on the output result of the candidate affinity prediction model.
[0051] Optionally, perform iterative optimization of the model parameters of the Figure 2 candidate affinity prediction model shown based on the training loss until the training is completed to obtain the trained target affinity prediction model.
[0052] The training method of the antigen-antibody affinity prediction model proposed by the present disclosure is to obtain the first model parameter set of the trained target complex three-dimensional feature map extraction model, and assign parameters to the initial feature map extraction layer of the initial affinity prediction model based on the first model parameter set to obtain the candidate affinity prediction model to be trained. Obtain the first training sample of the candidate affinity prediction model, and combine the first complex sample and the mutant complex sample of the first complex sample in the first training sample and input them into the candidate affinity prediction model for model training until the training ends to obtain the trained target affinity prediction model. In the present disclosure, the model parameters of the trained target complex three-dimensional feature map extraction model are reused in the antigen-antibody affinity prediction model, which improves the extraction accuracy and accuracy of the three-dimensional feature map of the antigen-antibody complex by the antigen-antibody affinity prediction model, reduces the training complexity of the antigen-antibody affinity prediction model, and further optimizes the training effect of the antigen-antibody affinity prediction model. The candidate affinity prediction model is trained by a mutant complex sample set composed of the first complex sample and multiple mutant complex samples corresponding to the first complex sample. Compared with the model obtained by learning based on a small amount of data, the generalization of the antigen-antibody affinity prediction model is improved. Based on the trained target affinity prediction model, the antigen-antibody affinity in multiple antigen-antibody complexes is predicted, which improves the efficiency and accuracy of antigen-antibody affinity prediction. In the scenario of large-scale antibody screening, the cost of antibody screening is reduced, the accuracy of antibody screening is improved, the practicability and applicability of the antigen-antibody affinity prediction and antibody screening methods are improved, the prediction method of antigen-antibody affinity is optimized, and the antibody screening method is further optimized.
[0053] In the above embodiment, regarding the acquisition of the trained target complex three-dimensional feature map extraction model, it can be combined with Figure 3 for further understanding. Figure 3 FIG. is a schematic flow chart of the training method of the antigen-antibody affinity prediction model according to another embodiment of the present disclosure. As Figure 3 shown, the method includes:
[0054] S301, obtain a candidate complex three-dimensional feature map extraction model to be trained.
[0055] In the embodiment of the present disclosure, a corresponding three-dimensional feature map extraction model for the antigen-antibody complex can be constructed based on the model construction method in the related art as the candidate complex three-dimensional feature map extraction model to be trained.
[0056] Among them, a candidate complex three-dimensional feature map extraction model can be constructed based on the equivariant graph neural network in the related technology, or a candidate complex three-dimensional feature map extraction model can be constructed based on other neural networks that can achieve feature map extraction. No specific limitation is made here.
[0057] S302, obtain the first atom set in the second complex sample and the first atom information of each first atom in the first atom set. Among them, the first atom set is obtained based on the atom sets of all amino acid residues included in the second complex sample, and the first atom information at least includes the first atom feature and the first atom coordinate of the first atom.
[0058] In the embodiments of the present disclosure, the complex used as the training sample for generating the candidate complex three-dimensional feature map extraction model can be marked as the second complex sample.
[0059] Optionally, the second complex sample is a partial complex in the binding surface region of the sample antigen and the sample antibody in the sample antigen-antibody complex.
[0060] Among them, a sample antigen-antibody complex composed of a sample antigen and a sample antibody can be obtained, and the complex in the antigen-antibody binding surface region composed of the antigenic epitope region of the sample antigen and the Complementarity Determining Region (CDR) of the sample antibody in the complex can be used as the second complex sample.
[0061] As an example, the second complex sample can be as Figure 4 shown, Figure 4 The atoms A1, A2, A3, A4, and A5 shown represent the sample antigen in the second complex sample, and the atoms B1, B2, B3, B4, and B5 represent the sample antibody in the second complex sample. Edges between the sample antigen and the sample antibody need to be represented.
[0062] Among them, the Figure 2 region corresponding to the atoms A1, A2, A3, A4, and A5 shown can be marked as the antigenic epitope region of the sample antigen, and, the Figure 2 region corresponding to the atoms B1, B2, B3, B4, and B5 shown can be marked as the CDR region of the sample antibody. In this scenario, the complex of the antigenic epitope region corresponding to the atoms A1, A2, A3, A4, and A5 and the CDR region corresponding to the atoms B2, B3, B4, and B5 can be determined as the second complex sample.
[0063] In the embodiments of the present disclosure, the sample antigen-antibody complex is a complex with crystal structure analysis data. In this scenario, the analysis data of the second complex sample can be obtained from the crystal structure analysis data of the sample antigen-antibody complex, so as to obtain the atomic information included in each of all the amino acid residues included in the second complex sample.
[0064] Among them, all the atoms included in each of the amino acid residues included in the second complex sample can be marked as the first atoms in the second complex sample, and then the first atom set in the second complex sample is obtained.
[0065] Optionally, there is corresponding first atom information for the first atom. Among them, the first atom information indicates the atomic coordinate information and atomic feature information of the atom including the first atom, and can be respectively marked as the first atomic coordinate and the first atomic feature of the first atom.
[0066] Among them, the first atomic feature at least includes atomic feature information such as the atomic category of the first atom, the residue category to which the atom belongs, the chain where the atom is located, and the relative position where the atom is located, and no specific limitation is made here.
[0067] In addition, the first atomic coordinate can be the spatial coordinate corresponding to the first atom in the second complex sample. Among them, based on the chain where the first atom is located and the position of the first atom on its chain, position coding is performed on the first atom, so as to obtain the coordinate information of the first atom.
[0068] S303. Obtain the preset atomic feature masking ratio, and perform atomic feature masking on part of the first atoms in the first atom set based on the atomic feature masking ratio to obtain the masked third complex sample.
[0069] In the embodiments of the present disclosure, a corresponding masking ratio can be set for the first atom set and marked as the preset atomic feature masking ratio. Among them, the atomic feature masking ratio can be 15%, or other ratios, and no specific limitation is made here.
[0070] Optionally, based on the atomic feature masking ratio, a second atom set to be masked is obtained from the first atom set.
[0071] It can be understood that after obtaining the atomic feature masking ratio, part of the atoms can be obtained from the first atom set based on this ratio, atomic feature masking is performed on this part of the atoms, and the set composed of this part of the atoms is marked as the second atom set in the first atom set.
[0072] Optionally, obtain the masking information of each second atom in the second atom set. The masking information includes at least one of atomic feature masking information and atomic coordinate masking information. For any second atom, obtain the second atom information of the second atom, and mask the second atom information based on the masking information of the second atom to obtain a masked third atom.
[0073] In the embodiments of the present disclosure, for any second atom, the information of the second atom can be obtained and marked as the second atom information. The second atom information includes the second atom coordinate and the second atom feature of the second atom.
[0074] In this scenario, corresponding coordinate masking information can be obtained for the second atom coordinate, and / or corresponding feature masking information can be obtained for the second atom feature as the masking information of the second atom.
[0075] Among them, the atom at the center point of the second complex sample can be obtained, and based on the coordinate of the center point atom, the initial coordinate of the second atom, and random Gaussian noise, the corresponding noise coordinate is generated, and the noise coordinate is used as the masking information of the second atom coordinate of the second atom.
[0076] Correspondingly, the second atom feature of the second atom can be obtained, and noise perturbation is performed on the second atom feature, so that the feature information obtained after the noise perturbation is used as the masking information of the second atom feature of the second atom.
[0077] In this scenario, for any second atom, the second atom information can be masked based on the obtained masking information of the second atom. Among them, the masking information corresponding to the second atom coordinate of the second atom information can be obtained to mask the second atom coordinate, and / or the masking information corresponding to the second atom feature of the second atom information can be obtained to mask the second atom feature, and the second atom after the second atom information is masked is marked as the masked third atom.
[0078] As an example, it is set to mask the second atom coordinate in the second atom information. The second atom coordinate is (x, y, z), and the corresponding masking information is (x*, y*, z*). Then, (x, y, z) can be masked based on (x*, y*, z*), and the masked second atom coordinate is shown as (x*, y*, z*), and the second atom showing the coordinate corresponding to the masking information is marked as the masked third atom.
[0079] Further, the set composed of the third atoms of each second atom is marked as the third atom set.
[0080] Optionally, a masked third complex sample is obtained based on a fourth atomic set in the first atomic set excluding the second atomic set and a third atomic set composed of third atoms of each second atom.
[0081] In the embodiments of the present disclosure, a set composed of the remaining first atoms in the first atomic set excluding the second atomic set may be marked as the fourth atomic set in the first atomic set.
[0082] In this scenario, the third atomic set and the fourth atomic set may be combined, and the data obtained after combination is used as a complex sample for training a candidate complex three-dimensional feature map extraction model and is marked as the third complex sample.
[0083] S304. Model training is performed on the candidate complex three-dimensional feature map extraction model based on the third complex sample until the training is completed, and a trained target complex three-dimensional feature map extraction model is obtained.
[0084] Optionally, through the candidate complex three-dimensional feature map extraction model, the masked information of the third complex sample is restored to obtain a fifth atomic set in the restored third complex sample and the third atomic information of each fifth atom. For any fifth atom, the third atomic information includes the third atomic feature and the third atomic coordinate of the fifth atom.
[0085] In the embodiments of the present disclosure, the third complex sample may be input into the candidate complex three-dimensional feature map extraction model, and the atomic information of the masked atoms included in the third complex sample is restored through the candidate complex three-dimensional feature extraction model.
[0086] Among them, the masked information of the atomic coordinates of the masked atoms may be restored, or the masked information of the atomic features of the masked atoms may be restored, so as to obtain the fifth atomic set included in the third complex sample, and the atomic information of each fifth atom is marked as the third atomic information of each fifth atom.
[0087] For any fifth atom, the third atomic information of the fifth atom may include the atomic feature information and the atomic coordinate information of the fifth atom. Among them, the atomic feature information of the fifth atom may be marked as the third atomic feature, and the atomic coordinate information of the fifth atom may be marked as the third atomic coordinate.
[0088] It should be noted that the fifth atomic set is composed of the atoms not masked in the third complex sample and the atoms after the masked information in the third complex sample is restored.
[0089] Optionally, a first three-dimensional feature map of the third complex sample is obtained based on the third atomic information of each fifth atom in the fifth atomic set.
[0090] For any fifth atom, based on the second atomic coordinates included in the second atomic information of each fifth atom, a set of sixth atoms adjacent to the fifth atom in the set of fifth atoms is obtained.
[0091] In the embodiments of the present disclosure, for any fifth atom, at least one atom adjacent to the fifth atom can be obtained in the set of fifth atoms and marked as a sixth atom, and then a set of sixth atoms composed of at least one sixth atom is obtained.
[0092] Optionally, a preset adjacent atom screening strategy can be obtained. The adjacent atom screening strategy includes one of an adjacent atom screening strategy based on K-nearest neighbors and an adjacent atom screening strategy based on a preset radius. For any fifth atom, through the adjacent atom screening strategy, based on the second atomic coordinates of each fifth atom, at least one sixth atom adjacent to the fifth atom is screened out from the set of fifth atoms to obtain a set of sixth atoms adjacent to the fifth atom.
[0093] In the embodiments of the present disclosure, the set of sixth atoms of each fifth atom can be obtained based on the K-nearest neighbor method, or the set of sixth atoms of each fifth atom can be obtained based on the radius selection method.
[0094] It can be understood that in the scenario of obtaining the set of sixth atoms based on the adjacent atom screening strategy of K-nearest neighbors, for any fifth atom, the K atoms closest to the fifth atom can be obtained, and the K atoms are used as the K sixth atoms adjacent to the fifth atom, and then the set of sixth atoms adjacent to the fifth atom is obtained.
[0095] Correspondingly, in the scenario of obtaining the set of sixth atoms based on the adjacent atom screening strategy of a preset radius, for any fifth atom, the fifth atom can be used as the center of the sphere, and a spherical range corresponding to the fifth atom is obtained based on the preset radius, and at least one atom within the range is used as at least one sixth atom adjacent to the fifth atom, and then the set of sixth atoms adjacent to the fifth atom composed of the at least one sixth atom is obtained.
[0096] Optionally, taking the fifth atom as a point and the connection lines between the fifth atom and each sixth atom in the corresponding set of sixth atoms as edges, a first three-dimensional feature map of the third complex sample extracted by the candidate complex three-dimensional feature map extraction model is obtained. For any sixth atom, the edge is used to represent the spatial distance information between the fifth atom and the sixth atom.
[0097] In an embodiment of the present disclosure, for any fifth atom, the fifth atom and each sixth atom in the corresponding sixth atom set of the fifth atom can be used as points, and the connection lines between the fifth atom and each sixth atom can be used as edges, so as to obtain a three-dimensional feature map corresponding to the third complex sample, which is marked as the first three-dimensional feature map.
[0098] It should be noted that the third atom information of the fifth atom carries the coordinate information of the fifth atom. In this scenario, the edge between the fifth atom and any sixth atom can be used to represent the spatial distance information between the fifth atom and the sixth atom.
[0099] Optionally, obtain the first sample label of the third complex sample, and based on the first three-dimensional feature map and the first sample label, obtain the first training loss of the candidate complex three-dimensional feature map extraction model.
[0100] In an embodiment of the present disclosure, the label information of the third complex sample can be obtained as the first sample label, and the first sample label and the first three-dimensional feature map are processed by an algorithm based on the loss algorithm in the related art, so as to obtain the loss value of the first three-dimensional feature map based on the first sample label, and this loss value is determined as the first training loss of the candidate complex three-dimensional feature map extraction model.
[0101] Optionally, adjust the model parameters of the candidate complex three-dimensional feature map extraction model based on the first training loss, and return to obtain the next third complex sample to continue the model training for the candidate complex three-dimensional feature map extraction model with adjusted parameters until the training is completed, so as to obtain the trained target complex three-dimensional feature map extraction model.
[0102] In an embodiment of the present disclosure, the model parameters of the candidate complex three-dimensional feature map extraction model can be adjusted based on the first training loss, and return to obtain the next third complex sample to continue the model training for the candidate complex three-dimensional feature map extraction model with adjusted parameters until the training is completed, so as to obtain the trained target complex three-dimensional feature map extraction model.
[0103] Among them, the training end condition of the candidate complex three-dimensional feature map extraction model can be set based on the training round. In the scenario where the current training round meets the preset training end condition, the model training of the candidate complex three-dimensional feature map extraction model can be ended, and the candidate complex three-dimensional feature map extraction model output in the last training round is determined as the trained target complex three-dimensional feature map extraction model.
[0104] Correspondingly, the training end condition of the candidate complex three-dimensional feature map extraction model can also be set based on the training output result. In the scenario where the output result of the current training round meets the preset training end condition, the model training of the candidate complex three-dimensional feature map extraction model can be ended, and the candidate complex three-dimensional feature map extraction model output in the last training round is determined as the trained target complex three-dimensional feature map extraction model.
[0105] As an example, the training iteration of the candidate complex three-dimensional feature map extraction model can be further understood in combination with the following content:
[0106] Optionally, the iteration of the candidate complex three-dimensional feature map extraction model can be understood in combination with the following formula group:
[0107] g = (v, ε)
[0108]
[0109] In the above formula, g represents the three-dimensional feature map extracted by the candidate complex three-dimensional feature map extraction model, v represents the nodes in the graph, ε represents the edges in the graph, h represents the feature representation of each node, x represents the coordinates of each node, i / j represents the identifier of the atom, that is, the i-th atom / the j-th atom, l represents the iteration layer number of the candidate complex three-dimensional feature map extraction model, a ij represents the feature of the edge, m i,j represents the feature representation of the node after iteration, and C is a hyperparameter used to control the coordinate update range of the atom.
[0110] The training method of the antigen-antibody affinity prediction model proposed in the present disclosure obtains a candidate complex three-dimensional feature map extraction model to be trained, performs atomic feature masking on the second atom set in the first atom set of the second complex sample to obtain a masked third complex sample, and performs model training on the candidate complex three-dimensional feature map extraction model based on the third complex sample until the training ends, obtaining a trained target complex three-dimensional feature map extraction model. In the present disclosure, by performing model training on the candidate complex three-dimensional feature map extraction model with the third complex sample, the trained target complex three-dimensional feature map extraction model focuses more on the region belonging to the antigen-antibody binding surface, so that the three-dimensional feature map extraction accuracy of the candidate affinity prediction model obtained based on the first model parameters of the trained target complex three-dimensional feature map extraction model is improved, and the training effect of the candidate affinity prediction model is optimized.
[0111] In the above embodiment, regarding the training of the affinity prediction model, it can also be combined with Figure 5 for further understanding. Figure 5The flowchart of the training method of the antigen-antibody affinity prediction model according to another embodiment of the present disclosure is shown as Figure 5 follows. The method includes:
[0112] S501. Obtain the initial feature map extraction layer of the initial affinity prediction model, and assign parameters to the initial feature map extraction layer based on the first model parameter set to obtain a candidate affinity prediction model to be trained.
[0113] Optionally, a second model parameter set of the initial feature map extraction layer can be obtained. For any second model parameter, obtain the third model parameter corresponding to the second model parameter from the first model parameter set, and assign parameters to the second model parameter based on the third model parameter to obtain the fourth model parameter after assignment.
[0114] In the embodiments of the present disclosure, multiple model parameters of the initial feature map extraction layer of the initial affinity prediction model can be obtained and marked as the second model parameter set.
[0115] In this scenario, based on each parameter in the first model parameter set of the trained target complex three-dimensional feature map extraction model, the parameters in the second model parameter set can be adjusted and assigned, so as to realize the reuse of the first model parameter set in the antigen-antibody affinity prediction model.
[0116] Among them, for any second model parameter, the parameter corresponding to the second model parameter can be obtained from the first model parameter set and marked as the third model parameter, and the third model parameter is assigned to the second model parameter, and then the second model parameter after assignment is marked as the fourth model parameter.
[0117] Optionally, based on the fourth model parameters after assignment of each second model parameter, a candidate affinity prediction model to be trained is obtained.
[0118] It can be understood that the initial affinity prediction model is adjusted based on the fourth model parameters obtained after parameter assignment of the second model parameter set based on the first model parameter set, and the adjusted model is determined as the candidate affinity prediction model to be trained.
[0119] S502. Through the candidate affinity prediction model, obtain the first antigen-antibody affinity change parameter of each mutant complex sample in the mutant complex sample set based on the first complex sample.
[0120] Optionally, for any mutant complex sample, through the candidate three-dimensional feature map extraction layer of the candidate affinity prediction model, extract the second three-dimensional feature map of the first complex sample and the third three-dimensional feature map of the mutant complex sample.
[0121] In the embodiments of the present disclosure, through the candidate three-dimensional feature map extraction layer of the candidate affinity prediction model, three-dimensional feature maps can be extracted for the first complex sample input to the candidate affinity prediction model and each complex in the mutant complex sample set.
[0122] Among them, the three-dimensional feature map of the first complex sample extracted by the candidate three-dimensional feature map extraction layer can be marked as the second three-dimensional feature map, and the three-dimensional feature maps of each mutant complex sample in the mutant complex sample set extracted by the candidate three-dimensional feature map extraction layer can be marked as the third three-dimensional feature map of each mutant complex sample.
[0123] Optionally, through the candidate three-dimensional feature map extraction layer of the candidate affinity prediction model, the first sample atom set in the first complex sample and the first mutant sample atom set in the mutant complex sample are extracted.
[0124] In the embodiments of the present disclosure, through the candidate three-dimensional feature map extraction layer, the atom set in the first complex sample can be extracted and marked as the first sample atom set, and the atom sets in each mutant complex sample can be marked as the first mutant sample atom sets.
[0125] Optionally, for any first sample atom, a second sample atom set adjacent to the first sample atom is obtained from the first sample atom set, and the first sample atom is used as a point, and the connection lines between the first sample atom and each second sample atom are used as edges to obtain the second three-dimensional feature map of the first complex sample.
[0126] In the embodiments of the present disclosure, for the first sample atom, at least one adjacent atom of the first sample atom can be obtained from the first sample atom set, and the set composed of the at least one adjacent atom is marked as the second sample atom set adjacent to the first sample atom.
[0127] Among them, the second sample atom set adjacent to the first sample atom can be obtained from the first sample atom set based on the adjacent atom acquisition strategy corresponding to K-nearest neighbors and / or the adjacent atom acquisition strategy corresponding to a preset radius.
[0128] In this scenario, the first sample atom and each second sample atom adjacent to the first sample atom can be used as points, and the connection lines between the first sample atom and each second sample atom are used as edges, thereby obtaining the second three-dimensional feature map of the first complex sample.
[0129] Optionally, for any mutant sample atom, obtain a set of second mutant sample atoms adjacent to the first mutant sample atom from the set of first mutant sample atoms, and use the first mutant sample atom as a point and the connection lines between the first mutant sample atom and each second mutant sample atom as edges to obtain a third three-dimensional feature map of the mutant complex sample.
[0130] In the embodiments of the present disclosure, for the process of obtaining the third three-dimensional feature map of any mutant complex sample, reference may be made to the process of obtaining the second three-dimensional feature map of the first complex sample above, which will not be elaborated here.
[0131] Optionally, through the candidate affinity prediction layer in the candidate affinity prediction model, obtain the first antigen-antibody affinity parameter of the first complex sample based on the second three-dimensional feature map, and obtain the second antigen-antibody affinity parameter of the mutant complex sample based on the third three-dimensional feature map.
[0132] In the embodiments of the present disclosure, the layer used for affinity prediction in the candidate affinity prediction model can be determined as the candidate affinity prediction layer, and the antigen-antibody affinity in the second three-dimensional feature map and each third three-dimensional feature map can be predicted through the candidate affinity prediction layer.
[0133] Among them, the predicted value of the antigen-antibody affinity in the second three-dimensional feature map output by the candidate affinity prediction layer can be determined as the first antigen-antibody affinity parameter of the first complex sample, and the predicted values of the antigen-antibody affinity in the third three-dimensional feature maps of each mutant complex sample output by the candidate affinity prediction layer can be determined as the second antigen-antibody affinity parameters of each mutant complex sample.
[0134] Optionally, obtain the change parameter of the second antigen-antibody affinity parameter based on the first antigen-antibody affinity parameter as the first antigen-antibody affinity change parameter of the mutant complex sample based on the first complex sample.
[0135] In the embodiments of the present disclosure, for the second antigen-antibody affinity parameter of any mutant complex sample, the change value of the second antigen-antibody affinity parameter based on the first antigen-antibody affinity parameter of the first complex sample can be obtained through the candidate affinity prediction model, and this change value can be determined as the affinity change parameter of the mutant complex sample based on the first complex sample and marked as the first antigen-antibody affinity change parameter.
[0136] In this scenario, the change parameters of the antigen-antibody affinity of each mutant complex sample based on the first complex sample can be obtained respectively, so as to obtain the first antigen-antibody affinity change parameters of each mutant complex sample.
[0137] It should be noted that the candidate affinity prediction model can obtain the change parameters of the antigen-antibody affinity of each mutant complex sample set based on the first complex sample, or can obtain the change parameters of the antigen-antibody affinity between any two or more mutant complex samples in the mutant complex sample set, and specific limitations are not made here.
[0138] S503. For any mutant complex sample, obtain the second sample label of the mutant complex sample to obtain the loss value of the first antigen-antibody affinity change parameter of the mutant complex sample based on the second sample label.
[0139] Optionally, for any mutant complex sample, its label information can be marked as the second sample label.
[0140] In this scenario, the first antigen-antibody affinity change parameter of the mutant complex sample output by the candidate affinity prediction model based on the first complex sample can be obtained, and based on the loss algorithm in the related technology, algorithm processing is performed on the second sample label and the first antigen-antibody affinity change parameter, and then the loss value of the mutant complex sample is obtained according to the result of the algorithm processing.
[0141] S504. According to the loss values of each mutant complex sample, obtain the second training loss of the candidate affinity prediction model.
[0142] In the embodiments of the present disclosure, the loss values of each mutant complex sample can be obtained, and the loss values of each mutant complex sample can be integrated, and then the integrated result is determined as the second training loss of the candidate affinity prediction model.
[0143] Optionally, the loss values of each mutant complex sample can be weighted and fused, and then the result of the weighted fusion is determined as the second training loss of the candidate affinity prediction model.
[0144] S505. According to the second training loss, adjust the model parameters of the candidate affinity prediction model, and return to obtain the next first complex sample and multiple next mutant complex samples of the next first complex sample, and continue to train the candidate affinity prediction model with adjusted model parameters until the training ends to obtain the trained target affinity prediction model.
[0145] In the embodiments of the present disclosure, the model parameters of the candidate affinity prediction model can be adjusted according to the second training loss, and return to obtain the next first complex sample and the corresponding set of next mutant complex samples of the next first complex sample to continue the model training of the candidate affinity prediction model with adjusted parameters until the training ends to obtain the trained target affinity prediction model.
[0146] It should be noted that the candidate three-dimensional feature map extraction layer in the candidate affinity prediction model that reuses the model parameters of the target complex three-dimensional feature map extraction model, and its model parameters participate in the iterative optimization of the model parameters of the candidate affinity prediction model.
[0147] The training method of the antigen-antibody affinity prediction model proposed by the present disclosure is to obtain a candidate affinity prediction model, input the first complex sample and the mutant complex sample set into the candidate affinity prediction model to obtain the antigen-antibody affinity prediction parameters in each complex in the first complex sample and the mutant complex sample set, and then obtain the first antigen-antibody affinity change parameters of each mutant complex sample in the mutant complex sample set based on the first complex sample. Further, based on the first antigen-antibody affinity change parameters of each mutant complex sample, the second training loss of the candidate affinity prediction model is obtained, and the candidate affinity prediction model is iteratively optimized according to the second training loss until the training is completed to obtain the trained target affinity prediction model. In the present disclosure, the model parameters of the trained target complex three-dimensional feature map extraction model are reused in the antigen-antibody affinity prediction model, which improves the extraction accuracy and accuracy of the three-dimensional feature map of the antigen-antibody complex by the antigen-antibody affinity prediction model, reduces the training complexity of the antigen-antibody affinity prediction model, and further optimizes the training effect of the antigen-antibody affinity prediction model. By using the mutant complex sample set composed of the first complex sample and multiple mutant complex samples corresponding to the first complex sample to train the candidate affinity prediction model, compared with the model learned based on a small amount of data, the generalization of the antigen-antibody affinity prediction model is improved. Based on the trained target affinity prediction model, the antigen-antibody affinity in multiple antigen-antibody complexes is predicted, which improves the efficiency and accuracy of antigen-antibody affinity prediction. In the scenario of large-scale antibody screening, the cost of antibody screening is reduced, the accuracy of antibody screening is improved, the practicability and applicability of the antigen-antibody affinity prediction and antibody screening methods are improved, the antigen-antibody affinity prediction method is optimized, and thus the antibody screening method is optimized.
[0148] The present disclosure also proposes an antibody screening method, combined with Figure 6 Understanding, Figure 6 is a schematic flowchart of the antibody screening method according to an embodiment of the present disclosure. As Figure 6 shown, the method includes:
[0149] S601, obtaining the trained target affinity prediction model.
[0150] In the embodiment of the present disclosure, the affinity between the antigen-antibody complexes input into the model can be predicted based on the trained target affinity prediction model.
[0151] Among them, the target affinity prediction model is obtained based on the above Figures 1 to 5 training method of the antigen-antibody affinity prediction model proposed in the embodiment.
[0152] S602, obtain the wild-type antigen-antibody complex of the target antigen and the set of candidate mutant antigen-antibody complexes corresponding to the wild-type antigen-antibody complex.
[0153] In the embodiments of the present disclosure, the antigen corresponding to the antibody to be screened can be marked as the target antigen. In this scenario, multiple antibodies to be screened can be obtained, and then the complexes of each of the multiple antibodies with the target antigen can be obtained.
[0154] Optionally, the antibodies to be screened are multiple mutant antibodies obtained based on the wild-type antibody. As an example, if the multiple mutant antibodies corresponding to the wild-type antibody are set to include mutant antibody 1 and mutant antibody 2, then the set composed of mutant antibody 1 and mutant antibody 2 can be regarded as the set of candidate mutant antibodies corresponding to the wild-type antibody.
[0155] In this scenario, the complex of the wild-type antibody and the target antigen can be marked as the wild-type antigen-antibody complex, and further, the complexes of each mutant antibody in the set of mutant antibodies corresponding to the wild-type antibody with the target antigen can be marked as the candidate mutant antigen-antibody complexes of each mutant antibody, so as to obtain the set of candidate mutant antigen-antibody complexes corresponding to the wild-type antigen-antibody complex.
[0156] S603, input the wild-type antigen-antibody complex and the set of candidate mutant antigen-antibody complexes into the target affinity prediction model, and obtain the candidate antigen-antibody affinity change parameters of each candidate mutant antigen-antibody complex in the set of candidate mutant antigen-antibody complexes based on the wild-type antigen-antibody complex through the target affinity prediction model.
[0157] In the embodiments of the present disclosure, the wild-type antigen-antibody complex and the corresponding set of candidate mutant antigen-antibody complexes can be input into the trained target affinity prediction model, and the predicted value of the antigen-antibody affinity in the wild-type antigen-antibody complex and the predicted value of the antigen-antibody affinity in each candidate mutant antigen-antibody complex in the set of candidate mutant antigen-antibody complexes can be obtained respectively through the target affinity prediction model.
[0158] Furthermore, for any candidate mutant antigen-antibody complex, the change parameter of the affinity of the candidate mutant antigen-antibody complex based on the wild-type antigen-antibody complex can be obtained through the target affinity prediction model, and it can be marked as the candidate antigen-antibody affinity change parameter of the candidate mutant antigen-antibody complex.
[0159] S604. Based on the change parameter of the affinity of the candidate antigen-antibody complex, screen out the target mutant antigen-antibody complex from the set of candidate mutant antigen-antibody complexes to obtain the target antibody of the target antigen.
[0160] In the embodiments of the present disclosure, the screening conditions of the antigen-antibody complex preset can be obtained, and based on the screening conditions, screening is performed from the change parameters of the candidate antigen-antibody affinities of each candidate mutant antigen-antibody complex to obtain the target mutant antigen-antibody complex that meets the screening conditions.
[0161] Further, the mutant antibody in the target mutant antigen-antibody complex is determined as the target antibody of the target antigen obtained from the set of candidate mutant antibodies based on the target affinity prediction model.
[0162] For the antibody screening method proposed by the present disclosure, a trained target affinity prediction model is obtained. The wild-type antigen-antibody complex of the target antigen and the corresponding set of candidate mutant antigen-antibody complexes are input into the target affinity prediction model, and the target affinity prediction model screens out the corresponding target antibody for the target antigen from the set of candidate mutant antigen-antibody complexes. In the present disclosure, the target antibody corresponding to the target antigen is obtained from multiple antibodies to be screened based on the trained target affinity prediction model, realizing antibody screening in a large-scale antibody screening scenario, improving the efficiency and accuracy of antibody screening, reducing the cost of antibody screening, improving the practicability and applicability of the antibody screening method, and optimizing the antibody screening method and effect.
[0163] In the above embodiments, regarding the acquisition of the target antibody, it can also be combined with Figure 7 For further understanding, Figure 7 is a schematic flowchart of the antibody screening method according to another embodiment of the present disclosure. As Figure 7 shown, the method includes:
[0164] S701. Obtain the wild-type antigen-antibody complex of the target antigen and the corresponding set of candidate mutant antigen-antibody complexes of the wild-type antigen-antibody complex.
[0165] Optionally, the wild-type antibody of the target antigen and the corresponding set of candidate mutant antibodies of the wild-type antibody can be obtained.
[0166] In the embodiments of the present disclosure, for the target antigen, there is its corresponding wild-type antibody and a set of candidate mutant antibodies composed of multiple mutant antibodies obtained based on the wild-type antibody.
[0167] In this scenario, a complex composed of a wild-type antibody and a target antigen can be obtained, and the complex of the wild-type antibody and the target antigen is determined as the wild-type antigen-antibody complex. Moreover, for any candidate mutant antibody, the complex of the candidate mutant antibody and the target antigen is determined as the candidate mutant antigen-antibody complex of the candidate mutant antibody, thereby obtaining a set of candidate mutant antigen-antibody complexes corresponding to the wild-type antigen-antibody complex.
[0168] S702. Input the wild-type antigen-antibody complex and the set of candidate mutant antigen-antibody complexes into the target affinity prediction model, and obtain the candidate antigen-antibody affinity change parameters of each candidate mutant antigen-antibody complex in the set of candidate mutant antigen-antibody complexes based on the wild-type antigen-antibody complex through the target affinity prediction model.
[0169] In the embodiments of the present disclosure, the relevant information of step S702 can be referred to the detailed content in the above embodiments, and will not be elaborated here.
[0170] S703. Screen out the target mutant antigen-antibody complex from the set of candidate mutant antigen-antibody complexes according to the candidate antigen-antibody affinity change parameters, so as to obtain the target antibody of the target antigen.
[0171] Optionally, based on the order from high to low affinity, sort the candidate antigen-antibody affinity change parameters, obtain the target antigen-antibody affinity parameter ranked first among the candidate antigen-antibody affinity change parameters, obtain the target mutant antigen-antibody complex corresponding to the target antigen-antibody affinity parameter in the candidate mutant antigen-antibody complex, and determine the candidate mutant antibody in the target mutant antigen-antibody complex as the target antibody of the target antigen.
[0172] In the embodiments of the present disclosure, based on the sorting algorithm in the related art, the candidate antigen-antibody affinity change parameters can be sorted according to the order from high to low affinity, and the parameters that meet the preset conditions can be screened out from the sorted candidate antigen-antibody affinity change parameters as the target antigen-antibody affinity change parameters.
[0173] Among them, the candidate antigen-antibody affinity change parameter ranked first among the sorted candidate antigen-antibody affinity change parameters can be obtained, and this parameter is marked as the target antigen-antibody affinity change parameter.
[0174] Furthermore, the corresponding complex of the target antigen-antibody affinity change parameter in the candidate mutant antigen-antibody complex can be obtained and marked as the target mutant antigen-antibody complex, and then the candidate mutant antibody in the target mutant antigen-antibody complex is determined as the target antibody of the target antigen.
[0175] It can be understood that the target antibody obtained based on the target antigen-antibody affinity change parameter has the best affinity with the target antigen compared to the remaining mutant antibodies in the candidate mutant antibody set except this target antibody.
[0176] The antibody screening method proposed by the present disclosure obtains the target antibody corresponding to the target antigen from multiple antibodies to be screened based on the trained target affinity prediction model, realizes antibody screening in a large-scale antibody screening scenario, improves the efficiency and accuracy of antibody screening, reduces the cost of antibody screening, improves the practicability and applicability of the antibody screening method, and optimizes the antibody screening method and effect.
[0177] Corresponding to the training methods of the antigen-antibody affinity prediction model proposed in the above several embodiments, an embodiment of the present disclosure also proposes a training device for the antigen-antibody affinity prediction model. Since the training device for the antigen-antibody affinity prediction model proposed in the embodiments of the present disclosure corresponds to the training methods of the antigen-antibody affinity prediction model proposed in the above several embodiments, the implementation manners of the above antigen-antibody affinity prediction model training methods are also applicable to the training device for the antigen-antibody affinity prediction model proposed in the embodiments of the present disclosure, and will not be described in detail in the following embodiments.
[0178] Figure 8 It is a structural schematic diagram of a training device for an antigen-antibody affinity prediction model according to an embodiment of the present disclosure. As Figure 8 shown, the training device 800 for the antigen-antibody affinity prediction model includes a first acquisition module 81, an assignment module 82, a second acquisition module 83, and a training module 84, wherein:
[0179] The first acquisition module 81 is configured to acquire a trained three-dimensional feature map extraction model of the target complex and a first model parameter set of the three-dimensional feature map extraction model of the target complex.
[0180] The assignment module 82 is configured to acquire an initial feature map extraction layer of the initial affinity prediction model and perform parameter assignment on the initial feature map extraction layer based on the first model parameter set to obtain a candidate affinity prediction model to be trained.
[0181] The second acquisition module 83 is configured to acquire a first training sample of the candidate affinity prediction model, wherein the first training sample includes a first complex sample and a set of mutant complex samples of the first complex sample.
[0182] The training module 84 is configured to input the first complex sample and the set of mutant complex samples into the candidate affinity prediction model for model training until the training is completed to obtain a trained target affinity prediction model.
[0183] In an embodiment of the present disclosure, the first acquisition module 81 is further configured to: acquire a candidate complex three-dimensional feature map extraction model to be trained. Acquire a first atom set in a second complex sample and first atom information of each first atom in the first atom set, where the first atom set is obtained based on the atom sets of all amino acid residues included in the second complex sample, and the first atom information includes at least a first atom feature and a first atom coordinate of the first atom. Acquire a preset atom feature masking ratio, and perform atom feature masking on a part of the first atoms in the first atom set based on the atom feature masking ratio to obtain a masked third complex sample. Train the candidate complex three-dimensional feature map extraction model based on the third complex sample until the training ends to obtain a trained target complex three-dimensional feature map extraction model.
[0184] In an embodiment of the present disclosure, the second complex sample is a partial complex in the binding surface region of the sample antigen and the sample antibody in the sample antigen-antibody complex.
[0185] In an embodiment of the present disclosure, the first acquisition module 81 is further configured to: acquire a second atom set to be masked from the first atom set based on the atom feature masking ratio. Acquire masking information of each second atom in the second atom set, where the masking information includes at least one of atom feature masking information and atom coordinate masking information. For any second atom, acquire the second atom information of the second atom, and perform masking on the second atom information based on the masking information of the second atom to obtain a masked third atom. Based on the fourth atom set other than the second atom set in the first atom set and the third atom set composed of the third atoms of each second atom, obtain a masked third complex sample.
[0186] In an embodiment of the present disclosure, the first acquisition module 81 is further configured to: perform masking information restoration on the third complex sample through the candidate complex three-dimensional feature map extraction model to obtain a fifth atom set in the restored third complex sample and third atom information of each fifth atom, where for any fifth atom, the third atom information includes a third atom feature and a third atom coordinate of the fifth atom. Based on the third atom information of each fifth atom in the fifth atom set, obtain a first three-dimensional feature map of the third complex sample. Acquire a first sample label of the third complex sample, and based on the first three-dimensional feature map and the first sample label, obtain a first training loss of the candidate complex three-dimensional feature map extraction model. Adjust the model parameters of the candidate complex three-dimensional feature map extraction model based on the first training loss, and return to acquire the next third complex sample to continue training the candidate complex three-dimensional feature map extraction model with the adjusted parameters until the training ends to obtain a trained target complex three-dimensional feature map extraction model.
[0187] In an embodiment of the present disclosure, the first acquisition module 81 is further configured to: for any fifth atom, based on the second atomic coordinates included in the third atomic information of each fifth atom, acquire a set of sixth atoms adjacent to the fifth atom in the set of fifth atoms. Taking the fifth atom as a point and the connection lines between the fifth atom and each sixth atom in the corresponding set of sixth atoms as edges, a first three-dimensional feature map of the third complex sample extracted by the candidate complex three-dimensional feature map extraction model is obtained, where for any sixth atom, the edge is used to represent the spatial distance information between the fifth atom and the sixth atom.
[0188] In an embodiment of the present disclosure, the first acquisition module 81 is further configured to: acquire a pre-set adjacent atom screening strategy, where the adjacent atom screening strategy includes one of an adjacent atom screening strategy based on K-nearest neighbors and an adjacent atom screening strategy based on a pre-set radius. For any fifth atom, through the adjacent atom screening strategy, based on the second atomic coordinates of each fifth atom, at least one sixth atom adjacent to the fifth atom is screened out from the set of fifth atoms to obtain a set of sixth atoms adjacent to the fifth atom.
[0189] In an embodiment of the present disclosure, the assignment module 82 is further configured to: acquire a second set of model parameters of the initial feature map extraction layer. For any second model parameter, acquire a third model parameter corresponding to the second model parameter from the first set of model parameters, and perform parameter assignment on the second model parameter based on the third model parameter to obtain an assigned fourth model parameter. Based on the fourth model parameters obtained by assigning values to each second model parameter, a candidate affinity prediction model to be trained is obtained.
[0190] In an embodiment of the present disclosure, the training module 84 is further configured to: through the candidate affinity prediction model, acquire the first antigen-antibody affinity change parameters of each mutant complex sample in the set of mutant complex samples based on the first complex sample. For any mutant complex sample, acquire the second sample label of the mutant complex sample to obtain the loss value of the first antigen-antibody affinity change parameter of the mutant complex sample based on the second sample label. According to the loss values of each mutant complex sample, acquire the second training loss of the candidate affinity prediction model. Adjust the model parameters of the candidate affinity prediction model according to the second training loss, and return to acquire the next first complex sample and multiple next mutant complex samples of the next first complex sample, and continue to train the candidate affinity prediction model with adjusted model parameters until the training ends to obtain a trained target affinity prediction model.
[0191] In the embodiment of the present disclosure, the training module 84 is further configured to: for any mutant complex sample, extract a second three-dimensional feature map of the first complex sample and a third three-dimensional feature map of the mutant complex sample through the candidate three-dimensional feature map extraction layer of the candidate affinity prediction model. Through the candidate affinity prediction layer in the candidate affinity prediction model, obtain a first antigen-antibody affinity parameter of the first complex sample based on the second three-dimensional feature map, and obtain a second antigen-antibody affinity parameter of the mutant complex sample based on the third three-dimensional feature map. Obtain a change parameter of the second antigen-antibody affinity parameter based on the first antigen-antibody affinity parameter, and use it as the first antigen-antibody affinity change parameter of the mutant complex sample based on the first complex sample.
[0192] In the embodiment of the present disclosure, the training module 84 is further configured to: through the candidate three-dimensional feature map extraction layer of the candidate affinity prediction model, extract a first sample atom set in the first complex sample and a first mutant sample atom set in the mutant complex sample. For any first sample atom, obtain a second sample atom set adjacent to the first sample atom from the first sample atom set, and use the first sample atom as a point and the connection line between the first sample atom and each second sample atom as an edge to obtain a second three-dimensional feature map of the first complex sample. For any mutant sample atom, obtain a second mutant sample atom set adjacent to the first mutant sample atom from the first mutant sample atom set, and use the first mutant sample atom as a point and the connection line between the first mutant sample atom and each second mutant sample atom as an edge to obtain a third three-dimensional feature map of the mutant complex sample.
[0193] The training device for the antigen-antibody affinity prediction model proposed by the present disclosure obtains the first set of model parameters of the trained target complex three-dimensional feature map extraction model, and assigns parameters to the initial feature map extraction layer of the initial affinity prediction model based on the first set of model parameters to obtain a candidate affinity prediction model to be trained. The first training sample of the candidate affinity prediction model is obtained, and the first complex sample and the mutant complex sample of the first complex sample in the first training sample are combined and input into the candidate affinity prediction model for model training until the training is completed, and a trained target affinity prediction model is obtained. In the present disclosure, the model parameters of the trained target complex three-dimensional feature map extraction model are reused in the antigen-antibody affinity prediction model, which improves the extraction accuracy and accuracy of the three-dimensional feature map of the antigen-antibody complex by the antigen-antibody affinity prediction model, reduces the training complexity of the antigen-antibody affinity prediction model, and further optimizes the training effect of the antigen-antibody affinity prediction model. By using the mutant complex sample set composed of the first complex sample and multiple mutant complex samples corresponding to the first complex sample to train the candidate affinity prediction model, the generalization of the antigen-antibody affinity prediction model is improved compared with the model learned based on a small amount of data. Based on the trained target affinity prediction model, the antigen-antibody affinity in multiple antigen-antibody complexes is predicted, which improves the efficiency and accuracy of antigen-antibody affinity prediction, reduces the cost of antibody screening in the scenario of large-scale antibody screening, improves the accuracy of antibody screening, improves the practicality and applicability of the antigen-antibody affinity prediction and antibody screening methods, optimizes the antigen-antibody affinity prediction method, and further realizes the optimization of the antibody screening method.
[0194] Corresponding to the antibody screening methods proposed in the above several embodiments, an embodiment of the present disclosure also proposes an antibody screening device. Since the antibody screening device proposed in the embodiment of the present disclosure corresponds to the antibody screening methods proposed in the above several embodiments, the implementation manners of the above antibody screening methods are also applicable to the antibody screening device proposed in the embodiment of the present disclosure and will not be described in detail in the following embodiments.
[0195] Figure 9 It is a schematic structural diagram of an antibody screening device according to an embodiment of the present disclosure. As Figure 9 shown, the antibody screening device 900 includes a third acquisition module 91, a fourth acquisition module 92, a prediction module 93, and a screening module 94, where:
[0196] The third acquisition module 91 is configured to acquire a trained target affinity prediction model, where the target affinity prediction model is obtained based on the training device of the antigen-antibody affinity prediction model according to any one of claims 15-25 above.
[0197] The fourth acquisition module 92 is configured to acquire a wild-type antigen-antibody complex of a target antigen and a set of candidate mutant antigen-antibody complexes corresponding to the wild-type antigen-antibody complex.
[0198] The prediction module 93 is configured to input the wild-type antigen-antibody complex and the set of candidate mutant antigen-antibody complexes into a target affinity prediction model, and obtain candidate antigen-antibody affinity change parameters of each candidate mutant antigen-antibody complex in the set of candidate mutant antigen-antibody complexes based on the wild-type antigen-antibody complex through the target affinity prediction model.
[0199] The screening module 94 is configured to screen out a target mutant antigen-antibody complex from the set of candidate mutant antigen-antibody complexes according to the candidate antigen-antibody affinity change parameters, so as to obtain a target antibody of the target antigen.
[0200] In an embodiment of the present disclosure, the fourth acquisition module 92 is further configured to: acquire a wild-type antibody of the target antigen and a set of candidate mutant antibodies corresponding to the wild-type antibody. Determine the complex of the wild-type antibody and the target antigen as the wild-type antigen-antibody complex. For any candidate mutant antibody, determine the complex of the candidate mutant antibody and the target antigen as the candidate mutant antigen-antibody complex of the candidate mutant antibody.
[0201] In an embodiment of the present disclosure, the screening module 94 is further configured to: sort the candidate antigen-antibody affinity change parameters in descending order of affinity, and obtain a target antigen-antibody affinity parameter that ranks first among the candidate antigen-antibody affinity change parameters. Obtain the target mutant antigen-antibody complex corresponding to the target antigen-antibody affinity parameter in the set of candidate mutant antigen-antibody complexes, and determine the candidate mutant antibody in the target mutant antigen-antibody complex as the target antibody of the target antigen.
[0202] The antibody screening device proposed by the present disclosure acquires a trained target affinity prediction model, inputs the wild-type antigen-antibody complex of the target antigen and the corresponding set of candidate mutant antigen-antibody complexes into the target affinity prediction model, and screens out the corresponding target antibody for the target antigen from the set of candidate mutant antigen-antibody complexes through the target affinity prediction model. In the present disclosure, the target antibody corresponding to the target antigen is obtained from a plurality of antibodies to be screened based on the trained target affinity prediction model, realizing antibody screening in a large-scale antibody screening scenario, improving the efficiency and accuracy of antibody screening, reducing the cost of antibody screening, improving the practicability and applicability of the antibody screening method, and optimizing the antibody screening method and effect.
[0203] According to an embodiment of the present disclosure, the present disclosure also proposes an electronic device, a readable storage medium, and a computer program product.
[0204] Figure 10 FIG. 1 is a schematic block diagram of an exemplary electronic device 100 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0205] As Figure 10 shown, the device 100 includes a computing unit 101 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 102 or a computer program loaded from a storage unit 108 into a random access memory (RAM) 103. In the RAM 103, various programs and data required for the operation of the device 100 can also be stored. The computing unit 101, the ROM 102, and the RAM 103 are connected to each other via a bus 104. An input / output (I / O) interface 105 is also connected to the bus 104.
[0206] A plurality of components in the device 100 are connected to the I / O interface 105, including: an input unit 106, such as, for example, a keyboard, a mouse, etc.; an output unit 106, such as, for example, various types of displays, speakers, etc.; a storage unit 108, such as, for example, a magnetic disk, an optical disk, etc.; and a communication unit 109, such as, for example, a network card, a modem, a wireless communication transceiver, etc. The communication unit 109 allows the device 100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0207] The computing unit 101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 101 executes the various methods and processes described above, such as the training method of the antigen-antibody affinity prediction model and / or the antibody screening method. For example, in some embodiments, the training method of the antigen-antibody affinity prediction model and / or the antibody screening method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 100 via the ROM 102 and / or the communication unit 109. When the computer program is loaded into the RAM 103 and executed by the computing unit 101, one or more steps of the training method of the antigen-antibody affinity prediction model and / or the antibody screening method described above can be executed. Alternatively, in other embodiments, the computing unit 101 can be configured to execute the training method of the antigen-antibody affinity prediction model and / or the antibody screening method by any other suitable means (such as, by means of firmware).
[0208] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0209] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0210] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0211] In order to provide an interaction with a user account, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user account; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user account can provide input to the computer. Other kinds of devices may also be used to provide an interaction with the user account; for example, the feedback provided to the user account may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user account may be in any form (including acoustic input, voice input, or tactile input).
[0212] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user account computer having a graphical user account interface or a web browser through which the user account can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0213] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0214] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0215] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A training method for an antigen-antibody affinity prediction model, wherein, The method includes: Obtaining a trained target complex three-dimensional feature map extraction model and a first model parameter set of the target complex three-dimensional feature map extraction model; Obtaining an initial feature map extraction layer of an initial affinity prediction model, and performing parameter assignment on the initial feature map extraction layer based on the first model parameter set to obtain a candidate affinity prediction model to be trained; Obtaining a first training sample of the candidate affinity prediction model, where the first training sample includes a first complex sample and a set of mutant complex samples of the first complex sample; Inputting the first complex sample and the set of mutant complex samples into the candidate affinity prediction model for model training until the training ends to obtain a trained target affinity prediction model.
2. The method according to claim 1, wherein, The obtaining of the trained target complex three-dimensional feature map extraction model includes: Obtaining a candidate complex three-dimensional feature map extraction model to be trained; Obtaining a first atom set in a second complex sample and first atom information of each first atom in the first atom set, where the first atom set is obtained based on atom sets of all amino acid residues included in the second complex sample, and the first atom information includes at least a first atom feature and a first atom coordinate of the first atom; Obtaining a preset atom feature masking ratio, and performing atom feature masking on part of the first atoms in the first atom set based on the atom feature masking ratio to obtain a masked third complex sample; Performing model training on the candidate complex three-dimensional feature map extraction model based on the third complex sample until the training ends to obtain the trained target complex three-dimensional feature map extraction model.
3. The method according to claim 2, wherein The second complex sample is a partial complex of the binding surface region of the sample antigen and the sample antibody in the sample antigen-antibody complex.
4. The method according to claim 2, wherein, The obtaining of the preset atom feature masking ratio and performing atom feature masking on part of the first atoms in the first atom set based on the atom feature masking ratio to obtain a masked third complex sample includes: Based on the atom feature masking ratio, obtaining a second atom set to be masked from the first atom set; Obtaining masking information of each second atom in the second atom set, where the masking information includes at least one of atom feature masking information and atom coordinate masking information; For any second atom, obtaining second atom information of the second atom, and masking the second atom information based on the masking information of the second atom to obtain a masked third atom; Based on a fourth atom set in the first atom set other than the second atom set and a third atom set composed of the third atoms of each second atom, obtaining the masked third complex sample.
5. The method according to claim 2, wherein, The performing of model training on the candidate complex three-dimensional feature map extraction model based on the third complex sample until the training ends to obtain the trained target complex three-dimensional feature map extraction model includes: Through the candidate complex three-dimensional feature map extraction model, the masked information of the third complex sample is restored to obtain the fifth atomic set in the restored third complex sample and the third atomic information of each fifth atom. For any fifth atom, the third atomic information includes the third atomic feature and the third atomic coordinate of the fifth atom; Based on the third atomic information of each fifth atom in the fifth atomic set, the first three-dimensional feature map of the third complex sample is obtained; Obtain the first sample label of the third complex sample, and based on the first three-dimensional feature map and the first sample label, obtain the first training loss of the candidate complex three-dimensional feature map extraction model; Based on the first training loss, adjust the model parameters of the candidate complex three-dimensional feature map extraction model, and return to obtain the next third complex sample to continue the model training for the candidate complex three-dimensional feature map extraction model with adjusted parameters until the training ends, and obtain the trained target complex three-dimensional feature map extraction model.
6. The method according to claim 5, wherein The obtaining the first three-dimensional feature map of the third complex sample based on the third atomic information of each fifth atom in the fifth atomic set includes: For any fifth atom, based on the second atomic coordinates included in the third atomic information of each fifth atom, obtain the sixth atomic set adjacent to the fifth atom in the fifth atomic set; Taking the fifth atom as a point and the connection lines between the fifth atom and each sixth atom in the corresponding sixth atomic set as edges, obtain the first three-dimensional feature map of the third complex sample extracted by the candidate complex three-dimensional feature map extraction model. For any sixth atom, the edge is used to represent the spatial distance information between the fifth atom and the sixth atom.
7. The method according to claim 6, wherein, The obtaining the sixth atomic set adjacent to the fifth atom in the fifth atomic set based on the second atomic coordinates included in the third atomic information of each fifth atom includes: Obtain a preset adjacent atom screening strategy, where the adjacent atom screening strategy includes one of an adjacent atom screening strategy based on K-nearest neighbors and an adjacent atom screening strategy based on a preset radius; For any of the fifth atoms, through the adjacent atom screening strategy, based on the second atomic coordinates of each fifth atom, screen out at least one sixth atom adjacent to the fifth atom from the fifth atomic set to obtain the sixth atomic set adjacent to the fifth atom.
8. The method according to claim 1, wherein, The obtaining the initial feature map extraction layer of the initial affinity prediction model and performing parameter assignment on the initial feature map extraction layer based on the first model parameter set to obtain the candidate affinity prediction model to be trained includes: Obtain the second model parameter set of the initial feature map extraction layer; For any second model parameter, obtain the third model parameter corresponding to the second model parameter from the first model parameter set, and perform parameter assignment on the second model parameter based on the third model parameter to obtain the fourth model parameter after assignment; Based on the fourth model parameters after the assignment of each second model parameter, the candidate affinity prediction model to be trained is obtained.
9. The method according to claim 1, wherein The step of inputting the first complex sample and the mutant complex sample set into the candidate affinity prediction model for model training until the training is completed to obtain the trained target affinity prediction model includes: Through the candidate affinity prediction model, obtaining the first antigen-antibody affinity change parameters of each mutant complex sample in the mutant complex sample set based on the first complex sample; For any mutant complex sample, obtaining the second sample label of the mutant complex sample to obtain the loss value of the first antigen-antibody affinity change parameter of the mutant complex sample based on the second sample label; According to the loss values of each mutant complex sample, obtaining the second training loss of the candidate affinity prediction model; Adjusting the model parameters of the candidate affinity prediction model according to the second training loss, and returning to obtain the next first complex sample and multiple next mutant complex samples of the next first complex sample, and continuing to train the candidate affinity prediction model with adjusted model parameters until the training is completed to obtain the trained target affinity prediction model.
10. The method according to claim 9, wherein, The step of, for any mutant complex sample, obtaining the first antigen-antibody affinity change parameters of the mutant complex sample based on the first complex sample through the candidate affinity prediction model includes: For any mutant complex sample, through the candidate three-dimensional feature map extraction layer of the candidate affinity prediction model, extracting the second three-dimensional feature map of the first complex sample and the third three-dimensional feature map of the mutant complex sample; Through the candidate affinity prediction layer in the candidate affinity prediction model, obtaining the first antigen-antibody affinity parameter of the first complex sample based on the second three-dimensional feature map, and obtaining the second antigen-antibody affinity parameter of the mutant complex sample based on the third three-dimensional feature map; Obtaining the change parameter of the second antigen-antibody affinity parameter based on the first antigen-antibody affinity parameter as the first antigen-antibody affinity change parameter of the mutant complex sample based on the first complex sample.
11. The method according to claim 10, wherein, The step of, for any mutant complex sample, through the candidate three-dimensional feature map extraction layer of the candidate affinity prediction model, extracting the second three-dimensional feature map of the first complex sample and the third three-dimensional feature map of the mutant complex sample includes: Through the candidate three-dimensional feature map extraction layer of the candidate affinity prediction model, extracting the first sample atom set in the first complex sample and the first mutant sample atom set in the mutant complex sample; For any first sample atom, obtaining the second sample atom set adjacent to the first sample atom from the first sample atom set, and using the first sample atom as a point and the connection line between the first sample atom and each second sample atom as an edge to obtain the second three-dimensional feature map of the first complex sample; For any mutant sample atom, obtain a set of second mutant sample atoms adjacent to the first mutant sample atom from the set of first mutant sample atoms, and use the first mutant sample atom as a point and the connection lines between the first mutant sample atom and each second mutant sample atom as edges to obtain the third three-dimensional feature map of the mutant complex sample.
12. An antibody screening method, wherein, The method includes: Obtain a trained target affinity prediction model, where the target affinity prediction model is obtained based on the training method of the affinity prediction model according to any one of claims 1-11 above; Obtain the wild-type antigen-antibody complex of the target antigen and a set of candidate mutant antigen-antibody complexes corresponding to the wild-type antigen-antibody complex; Input the wild-type antigen-antibody complex and the set of candidate mutant antigen-antibody complexes into the target affinity prediction model, and obtain the candidate antigen-antibody affinity change parameters of each candidate mutant antigen-antibody complex in the set of candidate mutant antigen-antibody complexes based on the wild-type antigen-antibody complex through the target affinity prediction model; According to the candidate antigen-antibody affinity change parameters, screen out the target mutant antigen-antibody complex from the set of candidate mutant antigen-antibody complexes to obtain the target antibody of the target antigen.
13. The method according to claim 12, wherein, The obtaining of the wild-type antigen-antibody complex of the target antigen and the set of candidate mutant antigen-antibody complexes corresponding to the wild-type antigen-antibody complex includes: Obtain the wild-type antibody of the target antigen and a set of candidate mutant antibodies corresponding to the wild-type antibody; Determine the complex of the wild-type antibody and the target antigen as the wild-type antigen-antibody complex; For any candidate mutant antibody, determine the complex of the candidate mutant antibody and the target antigen as the candidate mutant antigen-antibody complex of the candidate mutant antibody.
14. The method according to claim 12, wherein The screening out of the target mutant antigen-antibody complex from the set of candidate mutant antigen-antibody complexes according to the candidate antigen-antibody affinity change parameters to obtain the target antibody of the target antigen includes: Sort the candidate antigen-antibody affinity change parameters in descending order of affinity, and obtain the target antigen-antibody affinity parameter at the top of each candidate antigen-antibody affinity change parameter; Obtain the target mutant antigen-antibody complex corresponding to the target antigen-antibody affinity parameter in the set of candidate mutant antigen-antibody complexes, and determine the candidate mutant antibody in the target mutant antigen-antibody complex as the target antibody of the target antigen.
15. A training device for an antigen-antibody affinity prediction model, wherein, The device includes: A first acquisition module for obtaining a trained target complex three-dimensional feature map extraction model and a first set of model parameters of the target complex three-dimensional feature map extraction model; An assignment module for obtaining the initial feature map extraction layer of the initial affinity prediction model and performing parameter assignment on the initial feature map extraction layer based on the first set of model parameters to obtain a candidate affinity prediction model to be trained; A second acquisition module, configured to acquire a first training sample of the candidate affinity prediction model, where the first training sample includes a first complex sample and a set of mutant complex samples of the first complex sample; A training module, configured to input the first complex sample and the set of mutant complex samples into the candidate affinity prediction model for model training until the training ends, thereby obtaining a trained target affinity prediction model.
16. The device according to claim 15, wherein, The first acquisition module is further configured to: Acquire a candidate complex three-dimensional feature map extraction model to be trained; Acquire a first atom set in a second complex sample and first atom information of each first atom in the first atom set, where the first atom set is obtained based on atom sets of all amino acid residues included in the second complex sample, and the first atom information includes at least a first atom feature and a first atom coordinate of the first atom; Acquire a preset atom feature masking ratio, and based on the atom feature masking ratio, perform atom feature masking on some first atoms in the first atom set to obtain a masked third complex sample; Based on the third complex sample, perform model training on the candidate complex three-dimensional feature map extraction model until the training ends, thereby obtaining the trained target complex three-dimensional feature map extraction model.
17. The apparatus according to claim 16, wherein, The second complex sample is a partial complex of the binding surface region between a sample antigen and a sample antibody in a sample antigen-antibody complex.
18. The device according to claim 16, wherein, The first acquisition module is further configured to: Based on the atom feature masking ratio, acquire a second atom set to be masked from the first atom set; Acquire masking information of each second atom in the second atom set, where the masking information includes at least one of atom feature masking information and atom coordinate masking information; For any second atom, acquire second atom information of the second atom, and based on the masking information of the second atom, mask the second atom information to obtain a masked third atom; Based on a fourth atom set in the first atom set other than the second atom set and a third atom set composed of third atoms of each second atom, obtain the masked third complex sample.
19. The apparatus according to claim 16, wherein, The first acquisition module is further configured to: Through the candidate complex three-dimensional feature map extraction model, perform masking information restoration on the third complex sample to obtain a fifth atom set in the restored third complex sample and third atom information of each fifth atom, where for any fifth atom, the third atom information includes a third atom feature and a third atom coordinate of the fifth atom; Based on the third atom information of each fifth atom in the fifth atom set, obtain a first three-dimensional feature map of the third complex sample; Acquire a first sample label of the third complex sample, and based on the first three-dimensional feature map and the first sample label, obtain a first training loss of the candidate complex three-dimensional feature map extraction model. Adjust the model parameters of the candidate complex three-dimensional feature map extraction model based on the first training loss, and return to obtain the next third complex sample pair. Continue to perform model training on the candidate complex three-dimensional feature map extraction model with adjusted parameters until the training ends, and obtain the trained target complex three-dimensional feature map extraction model.
20. The apparatus according to claim 19, wherein, The first acquisition module is further configured to: For any fifth atom, based on the second atomic coordinates included in the third atomic information of each fifth atom, obtain a set of sixth atoms adjacent to the fifth atom in the set of fifth atoms; Use the fifth atom as a point and the connection lines between the fifth atom and each sixth atom in the corresponding set of sixth atoms as edges to obtain the first three-dimensional feature map of the third complex sample extracted by the candidate complex three-dimensional feature map extraction model. For any sixth atom, the edge is used to represent the spatial distance information between the fifth atom and the sixth atom.
21. The device according to claim 20, wherein, The first acquisition module is further configured to: Obtain a pre-set adjacent atom screening strategy, where the adjacent atom screening strategy includes one of an adjacent atom screening strategy based on K-nearest neighbors and an adjacent atom screening strategy based on a preset radius; For any of the fifth atoms, through the adjacent atom screening strategy, based on the second atomic coordinates of each fifth atom, screen out at least one sixth atom adjacent to the fifth atom from the set of fifth atoms to obtain the set of sixth atoms adjacent to the fifth atom.
22. The device according to claim 15, wherein, The assignment module is further configured to: Obtain the second model parameter set of the initial feature map extraction layer; For any second model parameter, obtain the third model parameter corresponding to the second model parameter from the first model parameter set, and perform parameter assignment on the second model parameter based on the third model parameter to obtain the fourth model parameter after assignment; Based on the fourth model parameters after assignment of each second model parameter, obtain the candidate affinity prediction model to be trained.
23. The apparatus according to claim 15, wherein, The training module is further configured to: Through the candidate affinity prediction model, obtain the first antigen-antibody affinity change parameters of each mutant complex sample in the mutant complex sample set based on the first complex sample; For any mutant complex sample, obtain the second sample label of the mutant complex sample to obtain the loss value of the first antigen-antibody affinity change parameter of the mutant complex sample based on the second sample label; According to the loss values of each mutant complex sample, obtain the second training loss of the candidate affinity prediction model; Adjust the model parameters of the candidate affinity prediction model according to the second training loss, and return to obtain the next first complex sample and multiple next mutant complex samples of the next first complex sample. Continue to train the candidate affinity prediction model with adjusted model parameters until the training ends, and obtain the trained target affinity prediction model.
24. The apparatus according to claim 23, wherein, The training module is further configured to: For any mutant complex sample, through the candidate three-dimensional feature map extraction layer of the candidate affinity prediction model, extract the second three-dimensional feature map of the first complex sample and the third three-dimensional feature map of the mutant complex sample; Through the candidate affinity prediction layer in the candidate affinity prediction model, based on the second three-dimensional feature map, obtain the first antigen-antibody affinity parameter of the first complex sample, and based on the third three-dimensional feature map, obtain the second antigen-antibody affinity parameter of the mutant complex sample; Obtain the change parameter of the second antigen-antibody affinity parameter based on the first antigen-antibody affinity parameter as the first antigen-antibody affinity change parameter of the mutant complex sample based on the first complex sample.
25. The apparatus according to claim 24, wherein, The training module is further configured to: Through the candidate three-dimensional feature map extraction layer of the candidate affinity prediction model, extract the first sample atom set in the first complex sample and the first mutant sample atom set in the mutant complex sample; For any first sample atom, obtain the second sample atom set adjacent to the first sample atom from the first sample atom set, and use the first sample atom as a point and the connection lines between the first sample atom and each second sample atom as edges to obtain the second three-dimensional feature map of the first complex sample; For any mutant sample atom, obtain the second mutant sample atom set adjacent to the first mutant sample atom from the first mutant sample atom set, and use the first mutant sample atom as a point and the connection lines between the first mutant sample atom and each second mutant sample atom as edges to obtain the third three-dimensional feature map of the mutant complex sample.
26. An antibody screening device, wherein, The device includes: A third acquisition module, configured to acquire a trained target affinity prediction model, where the target affinity prediction model is obtained based on the training device of the affinity prediction model according to any one of claims 15-25 above; A fourth acquisition module, configured to acquire a wild-type antigen-antibody complex of a target antigen and a set of candidate mutant antigen-antibody complexes corresponding to the wild-type antigen-antibody complex; A prediction module, configured to input the wild-type antigen-antibody complex and the set of candidate mutant antigen-antibody complexes into the target affinity prediction model, and obtain candidate antigen-antibody affinity change parameters of each candidate mutant antigen-antibody complex in the set of candidate mutant antigen-antibody complexes based on the wild-type antigen-antibody complex through the target affinity prediction model; A screening module, configured to screen out a target mutant antigen-antibody complex from the set of candidate mutant antigen-antibody complexes according to the candidate antigen-antibody affinity change parameters to obtain a target antibody of the target antigen.
27. The apparatus according to claim 26, wherein The fourth acquisition module is further configured to: Acquire a wild-type antibody of the target antigen and a set of candidate mutant antibodies corresponding to the wild-type antibody; Determine the complex of the wild-type antibody and the target antigen as the wild-type antigen-antibody complex; For any candidate mutant antibody, the complex of the candidate mutant antibody and the target antigen is determined as the candidate mutant antigen-antibody complex of the candidate mutant antibody.
28. The apparatus according to claim 26, wherein, The screening module is further configured to: Sort the candidate antigen-antibody affinity change parameters in descending order of affinity, and obtain the target antigen-antibody affinity parameter that ranks first among the candidate antigen-antibody affinity change parameters; Obtain the target mutant antigen-antibody complex corresponding to the target antigen-antibody affinity parameter in the candidate mutant antigen-antibody complex, and determine the candidate mutant antibody in the target mutant antigen-antibody complex as the target antibody of the target antigen.
29. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-11 and / or claims 12-14.
30. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11 and / or claims 12-14.
31. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-11 and / or claims 12-14.
Citation Information
Patent Citations
Method and device for model training, drug screening and affinity prediction
CN114333986A
Method and device for training antibody-protein binding affinity prediction model
CN115206415A