Antibody multi-site mutation evolution method based on multi-agent reinforcement learning

Optimizing the antibody multi-site mutation strategy through multi-agent reinforcement learning has solved the problem of high computational cost in the prior art, and achieved efficient antibody multi-site mutations to generate antibodies that can bind multiple viral variants at the same time.

CN120260676AActive Publication Date: 2025-07-04BEIHANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510704718.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-04
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing antibody evolution methods are cost-effective when performing multi-site mutations, and relying on the order of mutations, it is impossible to efficiently explore the multi-site mutation space of antibody.

Method used

Multi-agent reinforcement learning method is adopted, and multi-agent structure extraction module, residue agent mutation decision-making module and residue population value evaluation module are used to optimize the antibody multi-site mutation strategy, and multi-agent collaboration is used to perform multi-site mutations.

Benefits of technology

Significantly reduce the amount of calculation, improve the evolution efficiency of antibodies, and make full use of the antibody mutation space to achieve simultaneous binding of multiple viral variants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260676A_ABST
    Figure CN120260676A_ABST
Patent Text Reader

Abstract

The invention discloses an antibody multi-site mutation evolution method based on multi-agent reinforcement learning, which is characterized by comprising the following steps: 1) performing multi-agent reinforcement learning; the method comprises the following steps: inputting a sequence and a structure of an antibody-antigen protein compound to be evolved into a protein multi-level structure extraction module to obtain sequence relation characteristics of the antibody-antigen protein compound to be evolved and an initial state vector # imgabs1 # of a selected amino acid residue # imgabs0 # in the antibody-antigen protein compound to be evolved; inputting the initial state vector # imgabs2 # and the sequence relation characteristics into a trained residue agent mutation decision module to obtain an optimal mutation action # imgabs3 #, a value function # imgabs4 # and a global characteristic # imgabs5 #; and inputting the optimal mutation action # imgabs6 #, the value function # imgabs7 # and the global feature # imgabs8 # into a trained residue group value evaluation module to obtain an optimal global mutation strategy. The efficiency of antibody evolution design can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of antibody evolution, and more specifically, to an antibody multi-site mutation evolution method based on multi-agent reinforcement learning. Background Art

[0002] Antibodies play a crucial role in the adaptive immune system due to their precise specificity and high affinity. Traditional antibody development methods are labor-intensive and rely on animal experiments or hybridoma technology. Recent advances in artificial intelligence have paved the way for fully computational antibody engineering, greatly reducing the burden of wet laboratory experiments. Currently popular computational antibody design methods mainly fall into two approaches: de novo antibody design and evolution of existing antibodies. Although de novo design can generate novel molecules, its therapeutic assurance and potential immunogenicity are not yet clear. In contrast, evolution of existing antibodies makes minimal changes to the existing antibody structure by introducing stabilizing mutations, which can more convincingly ensure its biological function.

[0003] The current state-of-the-art antibody evolution algorithms mainly rely on mutation effect predictors to estimate the change in binding energy before and after mutation. The mutant antibody-antigen protein complex with the most stable predicted energy is selected as the final evolutionary result. However, these prediction-based evolution algorithms are based on single-residue mutations. In biology, the ability to mutate is usually achieved through simultaneous changes in multiple residues, a process called multi-site mutation, which contributes significantly to the genetic diversity required for biological evolution. For example, antibodies generate multi-site mutations in the third complementarity-determining region (CDRH3) of the antibody heavy chain through the variable region-diversity-joining rearrangement mechanism, thereby generating up to potential antibody candidates to neutralize the virus. Although existing antibody evolution methods can apply single-site mutation effect predictors to multi-site mutation situations, due to the need to perform large-scale mutation effect estimation in a vast multi-site evolution space, the computational cost is high. Therefore, these prediction-based methods can only simulate multi-residue evolution by sequentially performing single-site mutations, which essentially depends on the mutation order and usually requires wet laboratory assistance after each round of single-residue evolution.

[0004] Therefore, how to provide an efficient antibody multi-site mutation evolution method is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide an antibody multi-site mutation evolution method based on multi-agent reinforcement learning.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] An antibody multi-site mutation evolution method based on multi-agent reinforcement learning, comprising the following steps;

[0008] S1: Input the sequence and structure of the antibody-antigen protein complex to be evolved into the protein multi-level structure extraction module to obtain the sequence relationship features of the antibody-antigen protein complex to be evolved and the initial state vector of the selected amino acid residues in the antibody-antigen protein complex to be evolved. ; where ; N represents the number of selected amino acid residues;

[0009] S2: Input the initial state vector and the sequence relationship features into the trained residue agent mutation decision module to obtain the optimal mutation action of the selected amino acid residues, the value function corresponding to the optimal mutation action, and the global features of the antibody-antigen protein complex to be evolved. ;

[0010] S3: Input the optimal mutation action, the value function, and the global features into the trained residue population value evaluation module to obtain the optimal global mutation strategy of the antibody-antigen protein complex to be evolved.

[0011] Preferably, S1 specifically includes:

[0012] Perform one-hot encoding on the selected amino acid residues to obtain the one-hot features of the selected amino acid residues; ;

[0013] Use the atoms in the selected amino acid residues as the central nodes to establish a local coordinate system for the selected amino acid residues, and calculate the local coordinates of all atoms in the selected amino acid residues using the local coordinate system; ;

[0014] Concatenate the one-hot features and the local coordinates and input them into the MLP network to obtain the comprehensive feature vector; ;

[0015] The comprehensive feature vector and the learnable mutation embedding feature are subjected to feature splicing to obtain the initial state vector .

[0016] Preferably, the local coordinates include the local coordinates of the atom and the local coordinates of other atoms except the atom ;

[0017] The local coordinates are obtained based on the following formula:

[0018] ;

[0019] ;

[0020] ;

[0021] ;

[0022] wherein, successively represent the unit vectors in the x-axis, y-axis, and z-axis directions in the local coordinate system ; represents the transformation matrix of the selected amino acid residue ; represents the global coordinates of the atom in the selected amino acid residue ; represents the global coordinates of other atoms except the atom represents Gram-Schmidt orthogonalization; represents the vector pointing from the atom to the atom in the selected amino acid residue ; represents the vector pointing from the atom to the atom in the selected amino acid residue ;

[0023] Preferably, S2 specifically includes the following steps:

[0024] The local state vector and the sequence relationship feature are successively input into three invariant point attention modules to obtain the global state vector ; wherein, the local state vector is equal to the initial state vector ;

[0025] Input the local state vector and the sequence relationship feature into the first geometric attention module to obtain the local state vector ;

[0026] Utilize the global state vector and the local state vector to obtain the local state vector ; where ; and represent learnable weights shared by residues of the same amino acid type;

[0027] Input the local state vector and the sequence relationship feature into the second geometric attention module to obtain the local state vector ;

[0028] Utilize the global state vector and the local state vector to obtain the local state vector ; where ; and represent learnable weights shared by residues of the same amino acid type;

[0029] Input the local state vector and the sequence relationship feature into the third geometric attention module to obtain the local state vector ;

[0030] Maximize the mutation action - value function to obtain the optimal mutation action of the selected amino acid residue and the value function corresponding to the optimal mutation action ;

[0031] Add the local state sequence vector and the global state sequence vector to obtain the global feature ; where , .

[0032] Preferably, the calculation formula of the mutation action - value function is:

[0033] ;

[0034] The optimal mutation action and the value function The expression of

[0035] ;

[0036] ;

[0037] wherein, represents the mutation action of the selected amino acid residue , specifically representing a mutation into one of the 20 possible amino acid types; A represents the set composed of all types of mutation actions; represents estimating the value network of the expected mutation effect after performing the mutation action on the selected amino acid residue under the local state vector .

[0038] Preferably, S3 specifically includes:

[0039] Selecting from N optimal mutation actions for simultaneous multi-site mutation to obtain multi-site evolutionary strategies; wherein, ;

[0040] Inputting the optimal mutation actions included in each multi-site evolutionary strategy and the corresponding value function into the hypernetwork with parameters and to obtain the total value function corresponding to each multi-site evolutionary strategy; wherein, the multi-site evolutionary strategy corresponding to the minimum total value function is the optimal global mutation strategy.

[0041] Preferably, the parameters and are obtained based on the following formula;

[0042] ;

[0043] wherein, represents the global feature of the antibody-antigen protein complex to be evolved; represents the global feature of the antibody-antigen protein complex mutated by the multi-site evolutionary strategy; the sequence and structure of the antibody-antigen protein complex mutated by the multi-site evolutionary strategy are sequentially input into the protein multi-level structure extraction module and the residue agent mutation decision module to obtain the global feature ; represents the value mixing network.

[0044] Preferably, it further includes optimizing the network parameters of the residue agent mutation decision module and the residue population value evaluation module to obtain a trained residue agent mutation decision module and a trained residue agent mutation decision module;

[0045] Specifically, it includes the following steps:

[0046] Using the protein complex mutation effect standard data set and minimizing the loss function Training the protein multi-level structure extraction module and the residue agent mutation decision module to obtain the preliminary network parameters of the protein multi-level structure extraction module and the residue agent mutation decision module; where, represents the change in free binding energy from the wild-type antibody-antigen protein complex to the mutant antibody-antigen protein complex, and the protein complex mutation effect standard data set includes the wild-type antibody-antigen protein complex and the mutant antibody-antigen protein complex; represents the total value function obtained by the residue population value evaluation module when the optimal global mutation strategy output by the residue agent mutation decision module mutates the wild-type antibody-antigen protein complex into the mutant antibody-antigen protein complex when the input of the protein multi-level structure extraction module is the wild-type antibody-antigen protein complex.

[0047] Preferably, minimizing the loss function to further update the preliminary network parameters to obtain the final network parameters of the protein multi-level structure extraction module and the residue agent mutation decision module; where, represents the total value function when the antibody-antigen protein complex to be evolved adopts the optimal global mutation strategy, and r represents the reward function for the antibody-antigen protein complex to be evolved to mutate into the optimal antibody-antigen protein complex; the optimal antibody-antigen protein complex is the antibody-antigen protein complex mutated from the antibody-antigen protein complex to be evolved by adopting the optimal global mutation strategy.

[0048] Preferably, the calculation formula of the reward function r is:

[0049] ;

[0050] where, represents the change in binding energy; represents the binding energy of the optimal antibody-antigen protein complex; represents the binding energy of the antibody-antigen protein complex to be evolved; represents the reward scaling hyperparameter.

[0051] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses an antibody multi-site mutation evolution method based on multi-agent reinforcement learning, which can significantly reduce the computational amount and improve the efficiency of antibody evolution. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.

[0053] Figure 1 It is a flowchart of the antibody multi-site mutation evolution method based on multi-agent reinforcement learning provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0055] As Figure 1 shown, the embodiments of the present invention disclose an antibody multi-site mutation evolution method based on multi-agent reinforcement learning, including the following steps;

[0056] S1: Input the sequence and structure of the antibody-antigen protein complex to be evolved into the protein multi-level structure extraction module to obtain the sequence relationship features of the antibody-antigen protein complex to be evolved and the initial state vectors of the selected amino acid residues in the antibody-antigen protein complex to be evolved ; where ; N represents the number of selected amino acid residues;

[0057] It can be understood that:

[0058] The present invention regards the amino acid residues in the antibody-antigen protein complex to be evolved as an agent, and regards the antibody multi-site mutation evolution task as a multi-agent cooperation task;

[0059] The selected amino acid residues are the amino acid residues in the third complementarity-determining region (CDRH3) of the antibody heavy chain. CDRH includes 20 amino acid residues, that is, N = 20 in the present invention;

[0060] The antibody-antigen protein complex to be evolved consists of 128 amino acid residues in total. The 128 amino acid residues include 20 amino acid residues in CDRH3 and 108 amino acid residues adjacent to CDRH3.

[0061] The sequence relationship feature is represented as a feature matrix, where each entry corresponds to a learnable embedding vector for encoding the relative sequence distance between two amino acid residues.

[0062] Furthermore, S1 specifically includes:

[0063] Perform one-hot encoding on the selected amino acid residues to obtain the one-hot feature of the selected amino acid residues ;

[0064] It can be understood that: the one-hot feature is a binary vector of length 20, and each entry represents a specific residue type;

[0065] Using the atoms in the selected amino acid residues as the central nodes, establish a local coordinate system for the selected amino acid residues and calculate the local coordinates of all atoms in the selected amino acid residues using the local coordinate system ;

[0066] Furthermore, the local coordinates include the local coordinates of the atoms and the local coordinates of other atoms (such as N, C, O...) except the atoms ;

[0067] The local coordinates are obtained based on the following formula:

[0068] ;

[0069] ;

[0070] ;

[0071] ;

[0072] where successively represent the unit vectors in the x-axis, y-axis, and z-axis directions of the local coordinate system ; represents the selected amino acid residues​​​​ Transformation matrix; Representing the selected amino acid residue in Global coordinates of the atom; Representing Global coordinates of other atoms except the atom; Representing Gram - Schmidt orthogonalization; Representing the selected amino acid residue in which from atom pointing to atom vector; Representing the selected amino acid residue in which from atom pointing to atom vector.

[0073] Concatenate the one - hot feature and the local coordinates and input them into the MLP network to obtain a comprehensive feature vector ;

[0074] Concatenate the comprehensive feature vector and the learnable mutation embedding feature to obtain the initial state vector .

[0075] It can be understood that: the learnable mutation embedding feature is a learnable feature vector.

[0076] The present invention presets two learnable feature vectors indicating whether mutation can occur, and selects the corresponding learnable feature vector for concatenation according to whether the current amino acid residue can mutate.

[0077] S2: Input the initial state vector and the sequence relationship feature into the trained residue agent mutation decision module to obtain the optimal mutation action of the selected amino acid residue , the value function corresponding to the optimal mutation action and the global feature of the antibody - antigen protein complex to be evolved;

[0078] Furthermore, S2 specifically includes the following steps:

[0079] Input the local state vector and the sequence relationship feature into three invariant point attention (IPA) modules in sequence to obtain the global state vector ; where the local state vector is equal to the initial state vector ;

[0080] Input the local state vector and the sequence relationship feature into the first geometric attention (GA) module to obtain the local state vector ;

[0081] Utilize the global state vector and the local state vector to obtain the local state vector ; where

[0082] ; and represent learnable weights shared by residues of the same amino acid type;

[0083] Input the local state vector and the sequence relationship feature into the second geometric attention (GA) module to obtain the local state vector ;

[0084] Utilize the global state vector and the local state vector to obtain the local state vector ; where ; and represent learnable weights shared by residues of the same amino acid type;

[0085] Input the local state vector and the sequence relationship feature into the third geometric attention (GA) module to obtain the local state vector ;

[0086] Maximize the mutation action - value function to obtain the optimal mutation action of the selected amino acid residue and the value function corresponding to the optimal mutation action ;

[0087] Furthermore, the calculation formula of the mutation action - value function is:

[0088] ;

[0089] The expressions of the optimal mutation action and the value function are:

[0090] ;

[0091] ;

[0092] Among them, represents the mutation action of the selected amino acid residue , specifically representing a mutation to one of the 20 possible amino acid types; A represents the set of all types of mutation actions; represents estimating the selected amino acid residue under the local state vector and performing the mutation action The value network of the expected mutation effect after that.

[0093] Add the local state sequence vector and the global state sequence vector to obtain the global feature ; Among them, , .

[0094] S3: Input the optimal mutation action , the value function and the global feature into the trained residue population value evaluation module to obtain the optimal global mutation strategy of the antibody-antigen protein complex to be evolved.

[0095] Furthermore, S3 specifically includes:

[0096] Select from N optimal mutation actions for simultaneous multi-site mutation to obtain multi-site evolution strategies; Among them, ;

[0097] Input the optimal mutation actions included in each multi-site evolution strategy and the corresponding value function into the hypernetwork with parameters and to obtain the total value function corresponding to each multi-site evolution strategy; Among them, the multi-site evolution strategy corresponding to the minimum total value function is the optimal global mutation strategy.

[0098] Furthermore, obtain the parameters and based on the following formula;

[0099] ;

[0100] Among them, Represents the global features of the antibody-antigen protein complex to be evolved; Represents the global features of the antibody-antigen protein complex mutated by the multi-site evolution strategy; The sequence and structure of the antibody-antigen protein complex mutated by the multi-site evolution strategy are input into the protein multi-level structure extraction module and the residue agent mutation decision module in sequence to obtain the global features ; Represents the value mixing network.

[0101] It can be understood that: the global features and are obtained in the same way, except for the input. When the input is the antibody-antigen protein complex to be evolved, the residue agent mutation decision module outputs ; When the input is the antibody-antigen protein complex mutated by the multi-site evolution strategy, the residue agent mutation decision module outputs .

[0102] The structure of the antibody-antigen protein complex mutated by the multi-site evolution strategy is generated by the state transition function, and the state transition function is specifically the Markov transition function, which is implemented by the protein mutation structure generation platform Jackal.

[0103] Furthermore, the present invention also includes optimizing the network parameters of the residue agent mutation decision module and the residue population value evaluation module to obtain a trained residue agent mutation decision module and a trained residue agent mutation decision module;

[0104] Specifically, it includes the following steps:

[0105] Using the standard dataset of protein complex mutation effects and minimizing the loss function Train the protein multi-level structure extraction module and the residue agent mutation decision module to obtain the preliminary network parameters of the protein multi-level structure extraction module and the residue agent mutation decision module; wherein, Represents the change in free binding energy from the wild-type antibody-antigen protein complex to the mutant antibody-antigen protein complex, and the standard dataset of protein complex mutation effects includes the wild-type antibody-antigen protein complex and the mutant antibody-antigen protein complex; Represents the total value function obtained by the residue population value evaluation module when the optimal global mutation strategy output by the residue agent mutation decision module mutates the wild-type antibody-antigen protein complex into the mutant antibody-antigen protein complex when the input of the protein multi-level structure extraction module is the wild-type antibody-antigen protein complex.

[0106] It is understandable that: in the pre-training stage, the present invention uses the data in the standard dataset SKEMPIv2.0 of protein complex mutation effects, inputs the wild-type antibody-antigen protein complexes and mutant antibody-antigen protein complexes (including multiple) in the dataset, and after passing through the protein multi-level structure feature extraction module, residue agent mutation decision module, and residue group value evaluation module, obtains the evaluation of the affinity change under specific mutations of the complex (that is ); is the known data in the dataset, minimizing the loss function

[0107] Training the protein multi-level structure extraction module and the residue agent mutation decision module can obtain the initial network parameters of the protein multi-level structure extraction module and the residue agent mutation decision module.

[0108] Furthermore, minimizing the loss function to further update the initial network parameters and obtain the final network parameters of the protein multi-level structure extraction module and the residue agent mutation decision module; where represents the total value function when the antibody-antigen protein complex to be evolved adopts the optimal global mutation strategy, and r represents the reward function for the antibody-antigen protein complex to be evolved to mutate into the optimal antibody-antigen protein complex; the optimal antibody-antigen protein complex is the antibody-antigen protein complex mutated from the antibody-antigen protein complex to be evolved by adopting the optimal global mutation strategy.

[0109] Furthermore, the calculation formula of the reward function r is:

[0110] ;

[0111] where represents the change in binding energy; represents the binding energy of the optimal antibody-antigen protein complex; represents the binding energy of the antibody-antigen protein complex to be evolved; represents the reward scaling hyperparameter, and the present invention ;

[0112] It should be noted that the reward function adopted by the present invention is not unique. The embodiments of the present invention only give one form of the reward function,

[0113] After pre-training, the prediction ability of MERF (the antibody multi-site mutation evolution method provided by the present invention is called MERF) was benchmarked on the AB-Bind dataset, with a focus on evaluating its performance on previously "unseen" antibody-antigen protein complexes.

[0114] The present invention compared MERF with four structure-based methods: an empirical energy method FoldX, a machine learning method GeoPPI using self-supervised representation learning, an end-to-end deep learning method DDG-pred, and a newly proposed pre-trained deep learning method Gearbind.

[0115] Since FoldX is designed for single-site mutations and predicts mutation effects only under the sequential residue-by-residue mutation setting. Except for FoldX, the other methods predicted mutation effects under both direct multi-site mutation prediction and sequential residue-by-residue mutation settings.

[0116] In the sequential residue-by-residue mutation setting, the mutation order was determined based on the arrangement of the dataset to avoid introducing additional biases. The predicted mutation effects were compared with the true values using Pearson correlation coefficient (Pearson-r), Spearman correlation coefficient (Spearman-r), mean absolute error (MAE), and mean squared error (MSE).

[0117] The final results showed that MERF exhibited superior performance and was comparable to the state-of-the-art method Gearbind. In addition, MERF could stably achieve accurate predictions under different antibody-antigen protein complexes, while the results of other methods fluctuated greatly.

[0118] Benefiting from the efficient exploration of the antibody multi-site mutation space, this method can make more full use of the antibody mutation space and achieve functions that are difficult to achieve with previous single-site mutation methods: for example, optimizing a single antibody to bind multiple virus variants simultaneously. Taking the design of an antibody that binds three antigens (i.e., three virus variants) as an example, according to the method of calculating the reward function of the present invention, three reward functions of a single antibody relative to three antigens can be obtained. , by minimizing the loss function further optimizing the network parameters, the antibody generated by MERF finally obtained can bind three virus variants simultaneously. Among them, is a hyperparameter set artificially to measure the importance of different antigens, and generally, it can be defaulted to 1.

[0119] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For related parts, reference can be made to the description in the method part.

[0120] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An antibody multi-site mutation evolution method based on multi-agent reinforcement learning, characterized in that including the following steps; S1: Input the sequence and structure of the antibody-antigen protein complex to be evolved into the protein multi-level structure extraction module to obtain the sequence relationship features of the antibody-antigen protein complex to be evolved and the initial state vectors of the selected amino acid residues in the antibody-antigen protein complex to be evolved ; where, ; N represents the number of selected amino acid residues; ​ S2: Input the initial state vector and the sequence relationship feature into the trained residue intelligent agent mutation decision module to obtain the optimal mutation action of the selected amino acid residue , the value function corresponding to the optimal mutation action , and the global feature of the antibody-antigen protein complex to be evolved ; ;​ S3: Input the optimal mutation action , the value function and the global feature into the trained residue population value evaluation module to obtain the optimal global mutation strategy of the antibody-antigen protein complex to be evolved.

2. The antibody multi-site mutation evolution method based on multi-agent reinforcement learning according to claim 1, characterized in that S1 specifically includes: Perform one-hot encoding on the selected amino acid residue to obtain the one-hot feature of the selected amino acid residue ;​ With respect to the selected amino acid residue among using the atom as the central node, establish a local coordinate system for the selected amino acid residue and calculate the local coordinates of all atoms in the selected amino acid residue by using the local coordinate system ; ​​ Concatenate the one-hot feature and the local coordinates and input the concatenated features into the MLP network to obtain a comprehensive feature vector ; Concatenate the comprehensive feature vector and the learnable mutation embedding feature to obtain the initial state vector .

3. The antibody multi-site mutation evolution method based on multi-agent reinforcement learning according to claim 2, characterized in that The local coordinates include the local coordinates of atoms and the local coordinates of other atoms except ; The local coordinates are obtained based on the following formula: ; ; ; ; Among them, successively represent the unit vectors in the directions of the x-axis, y-axis, and z-axis in the local coordinate system ; represents the transformation matrix of the selected amino acid residue ; represents the global coordinates of the atom in the selected amino acid residue ; represents the global coordinates of other atoms except the atom; represents Gram - Schmidt orthogonalization; represents the vector from the atom to the atom in the selected amino acid residue ; represents the vector from the atom to the atom in the selected amino acid residue ; 4. A method for antibody multi-site mutation evolution based on multi-agent reinforcement learning according to claim 2, characterized in that S2 specifically includes the following steps: Input the local state vector and the sequence relationship features into three fixed-point attention modules in sequence to obtain the global state vector ; wherein, the local state vector is equal to the initial state vector ; Input the local state vector and the sequence relationship feature into the first geometric attention module to obtain the local state vector ; Using the global state vector and the local state vector , obtain the local state vector ; wherein ; and represent learnable weights shared by residues of the same amino acid type; Input the local state vector and the sequence relationship feature into the second geometric attention module to obtain the local state vector ; Using the global state vector and the local state vector , obtain the local state vector ; wherein ; and represent learnable weights shared by the same amino acid type residues; Input the local state vector and the sequence relationship feature into the third geometric attention module to obtain the local state vector ; Maximize the mutant action-value function to obtain the selected amino acid residue with the optimal mutant action and the value function corresponding to the optimal mutant action ; Add the local state sequence vector and the global state sequence vector to obtain the global feature ; where , .

5. A method for antibody multi-site mutation evolution based on multi-agent reinforcement learning according to claim 4, characterized in that: The mutated action-value function has the following calculation formula: ; The optimal mutation action and the value function are expressed as follows: ; ; Among them, represents the mutation action of the selected amino acid residue , specifically representing a mutation to one of the 20 possible amino acid types; A represents the set composed of all types of mutation actions; represents the estimation of the selected amino acid residue when the local state vector is and the value network of the expected mutation effect after executing the mutation action .

6. The antibody multi-site mutation evolution method based on multi-agent reinforcement learning according to claim 1, wherein S3 specifically includes: Select from N optimal mutant actions and perform multi-site simultaneous mutagenesis to obtain a multi-site evolutionary strategy; where ; Input the optimal mutation actions and the corresponding value functions included in each multi-locus evolution strategy into the hypernetwork with parameters and to obtain the total value function corresponding to each multi-locus evolution strategy; among them, the multi-locus evolution strategy corresponding to the minimum of the total value function is the optimal global mutation strategy.

7. A method for antibody multi-site mutation evolution based on multi-agent reinforcement learning according to claim 6, characterized in that: The parameter is obtained based on the following formula and ; ; Among them, represents the global feature of the antibody-antigen protein complex to be evolved; represents the global feature of the antibody-antigen protein complex mutated by the multi-site evolution strategy; the sequence and structure of the antibody-antigen protein complex mutated by the multi-site evolution strategy are input into the protein multi-level structure extraction module and the residue agent mutation decision module in sequence to obtain the global feature ; represents the value mixing network.

8. A method for antibody multi-site mutation evolution based on multi-agent reinforcement learning according to claim 1, characterized in that, It also includes optimizing the network parameters of the residue agent mutation decision module and the residue population value evaluation module to obtain a trained residue agent mutation decision module and a trained residue agent mutation decision module; Specifically includes the following steps: Utilize the standard dataset of protein complex mutation effects and minimize the loss function Train the protein multi-level structure extraction module and the residue agent mutation decision module to obtain the initial network parameters of the protein multi-level structure extraction module and the residue agent mutation decision module; wherein, represents the change in free binding energy from the wild-type antibody-antigen protein complex to the mutant antibody-antigen protein complex, and the standard dataset of protein complex mutation effects includes the wild-type antibody-antigen protein complex and the mutant antibody-antigen protein complex; represents the total value function obtained by the residue group value evaluation module when the optimal global mutation strategy output by the residue agent mutation decision module mutates the wild-type antibody-antigen protein complex into the mutant antibody-antigen protein complex when the input of the protein multi-level structure extraction module is the wild-type antibody-antigen protein complex.

9. A method for antibody multi-site mutation evolution based on multi-agent reinforcement learning according to claim 8, characterized in that Minimize the loss function to further update the preliminary network parameters and obtain the final network parameters of the protein multi-level structure extraction module and the residue agent mutation decision module; wherein, represents the total value function when the antibody-antigen protein complex to be evolved adopts the optimal global mutation strategy, and r represents the reward function for the antibody-antigen protein complex to be evolved to mutate into the optimal antibody-antigen protein complex; the optimal antibody-antigen protein complex is the antibody-antigen protein complex obtained by mutating the antibody-antigen protein complex to be evolved using the optimal global mutation strategy.

10. A method for antibody multi-site mutation evolution based on multi-agent reinforcement learning according to claim 9, characterized in that, The calculation formula of the reward function r is: ; Among them, represents the binding energy change; represents the binding energy of the optimal antibody-antigen protein complex; represents the binding energy of the antibody-antigen protein complex to be evolved; represents the reward scaling hyperparameter.

Citation Information

Patent Citations

  • Computer antibody combination mutation evolution system and method, information data processing terminal

    CN109086568A

  • Collaborative design method for antibody sequence structure based on flow model

    CN114360636A

  • Deep neural network strategy network model interpretability conversion algorithm

    CN117350362A

  • Antibacterial peptide sequence design framework based on protein language model supervised fine tuning and human feedback reinforcement learning

    CN120048356A

  • Predicting protein structures using protein graphs

    US20230410938A1