Potential energy difference maximization-based antigen-antibody interaction potential energy matrix optimization method

By optimizing the antigen-antibody interaction potential energy matrix based on the maximization of potential energy difference, and utilizing the average energy value and residue frequency statistics of the natural complexes of the training samples, combined with hydrophilicity and hydrophobicity parameter constraints, the problem of insufficient specificity in the prediction of the MJ potential energy matrix is ​​solved, and the prediction accuracy of antigen-antibody docking posture is improved.

CN120913656AActive Publication Date: 2025-11-07BEIJING ANBAISHENG DIAGNOSTIC TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511429423.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2025-11-07
Estimated Expiration
2045-10-09

AI Technical Summary

Technical Problem

The existing MJ potential energy matrix, being a general potential energy function, cannot effectively predict the specificity of antigen-antibody binding, resulting in reduced accuracy in predicting antigen-antibody docking posture.

Method used

By optimizing the antigen-antibody interaction potential energy matrix using a potential energy difference maximization method, the potential energy matrix is ​​optimized by utilizing the average energy value of the natural complex of the training samples, the energy calculation of the non-binding conformation, residue frequency statistics, and hydrophilicity/hydrophobicity parameter constraints to improve prediction accuracy.

Benefits of technology

This improved the accuracy of predicting antigen-antibody docking attitudes, ensured that the optimization process conformed to biophysical laws, avoided overfitting and physically unreasonable results, and enhanced the specificity and reliability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913656A_ABST
    Figure CN120913656A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an antigen-antibody interaction potential energy matrix optimization method based on potential energy difference maximization. A specific embodiment of the method comprises the steps of determining a natural compound average energy value of each training sample in a training sample set based on an initialized potential energy matrix; performing structure adjustment on the antigen-antibody complex crystal structure in each training sample in the training sample set to generate a non-binding conformation set; performing energy calculation on each non-combined conformation in the non-combined conformation set to generate a background average energy value; determining an energy difference between the natural composite average energy value and the background average energy value as a potential energy difference; performing residue pair contact frequency statistics on each non-binding conformation in the non-binding conformation set to generate a residue frequency matrix; and optimizing the initialized potential energy matrix to obtain an optimized potential energy matrix. According to the embodiment, the accuracy of antigen-antibody docking posture prediction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the technical field of target optimization, the technical field of antigen-antibody docking, and the technical field of computer, in particular to an antigen-antibody interaction potential matrix optimization method based on potential energy difference maximization. BACKGROUND

[0002] Antigen-antibody docking is a core technology in immunobiology, drug development and vaccine development. Its goal is to predict the binding mode (conformation) of antigen and antibody through computational simulation, providing a basis for understanding immune recognition mechanisms and designing efficient antibody molecules. The core process of docking software can be summarized as follows: first, a large number of potential antigen-antibody binding poses (docking conformations) are generated through docking algorithms (such as shape complementation, fragment assembly or sampling optimization algorithms), and then these poses are scored and ranked by relying on knowledge-based energy functions, and finally the lowest energy and most likely physiological binding mode is selected. The knowledge-based potential energy function (also known as the statistical potential energy function) derives its parameters from the statistical rules of known protein structures, and usually takes the spatial distribution characteristics (such as distance, angle) of amino acid residue pairs (or atom pairs) as the core consideration object, and derives energy parameters through statistical preference. For example, MJ potential (Miyazawa-Jernigan, Estimation of Effective Interresidue Contact Energies from Protein Crystal Structures: Quasi-Chemical Approximation).

[0003] However, the MJ potential matrix is constructed by statistically analyzing the contact frequency of residue pairs in protein structures, and is a general potential energy function. However, antigen-antibody binding has its own characteristics, which makes the MJ potential matrix insufficient in predicting the specificity of antigen-antibody docking poses, thereby reducing the accuracy of predicting antigen-antibody docking poses. SUMMARY

[0004] This part of the disclosure is used to introduce the concepts in a brief form, which will be described in detail in the specific embodiments part. This part of the disclosure is not intended to identify the key features or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0005] Some embodiments of the present disclosure propose an antigen-antibody interaction potential matrix optimization method based on potential energy difference maximization to solve the technical problems mentioned in the background part.

[0006] In a first aspect, some embodiments of the present disclosure provide an antigen-antibody interaction potential matrix optimization method based on potential energy difference maximization, the method comprising: determining a native complex average energy value of each training sample in a training sample set based on an initialization potential matrix, wherein the initialization potential matrix is composed of initial potential energy differences between each amino acid in a preset amino acid sequence, and the training sample comprises an antigen-antibody complex crystal structure; performing structure adjustment on the antigen-antibody complex crystal structure in each training sample in the training sample set to generate a set of unbound conformations; performing energy calculation on each unbound conformation in the set of unbound conformations based on the initialization potential matrix to generate a background average energy value; determining an energy difference between the native complex average energy value and the background average energy value as a potential energy difference; performing residue contact frequency statistics on each unbound conformation in the set of unbound conformations to generate a residue frequency matrix; optimizing the initialization potential matrix according to the potential energy difference, the residue frequency matrix, and a hydrophilic-hydrophobic parameter constraint to obtain an optimized potential matrix, wherein the hydrophilic-hydrophobic parameter constraint is used to limit the optimization limits of the energy values in the initialization potential matrix; in response to determining that the potential energy difference and the current training round do not satisfy a preset convergence condition, using the optimized potential matrix as the initialization potential matrix to perform initialization potential matrix optimization again.

[0007] In a second aspect, some embodiments of the present disclosure provide an antigen-antibody interaction potential energy matrix optimization device based on potential energy difference maximization, the device comprising: a first determination unit configured to determine a natural complex average energy value of each training sample in a training sample set based on an initialization potential energy matrix, wherein the initialization potential energy matrix is composed of initial potential energy differences between each amino acid in a preset amino acid sequence, and the training sample comprises an antigen-antibody complex crystal structure; a structure adjustment unit configured to perform structure adjustment on the antigen-antibody complex crystal structure in each training sample in the training sample set to generate a set of unbound conformations; an energy calculation unit configured to perform energy calculation on each unbound conformation in the set of unbound conformations based on the initialization potential energy matrix to generate a background average energy value; a second determination unit configured to determine an energy difference between the natural complex average energy value and the background average energy value as a potential energy difference; a residue pair contact frequency statistical unit configured to perform residue pair contact frequency statistics on each unbound conformation in the set of unbound conformations to generate a residue frequency matrix; an optimization unit configured to optimize the initialization potential energy matrix according to the potential energy difference, the residue frequency matrix, and a hydrophilic-hydrophobic parameter constraint to obtain an optimized potential energy matrix, wherein the hydrophilic-hydrophobic parameter constraint is used to limit the optimization limit of the energy value in the initialization potential energy matrix; and a convergence judgment unit configured to, in response to determining that the potential energy difference and the current training round do not satisfy a preset convergence condition, use the optimized potential energy matrix as the initialization potential energy matrix and perform initialization potential energy matrix optimization again.

[0008] In a third aspect, some embodiments of the present disclosure provide an electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the implementations of the first aspect.

[0009] In a fourth aspect, some embodiments of the present disclosure provide a computer readable medium having a computer program stored thereon, wherein the program is executed by a processor to implement the method described in any of the implementations of the first aspect.

[0010] The above various embodiments of the present disclosure have the following beneficial effects: the antigen-antibody interaction potential matrix optimization method based on potential energy difference maximization of some embodiments of the present disclosure can improve the accuracy of predicting the antigen-antibody docking pose. Specifically, the reason for the decrease in the accuracy of predicting the antigen-antibody docking pose is that the MJ potential energy matrix is constructed by counting the contact frequency of residue pairs in the protein structure, which is a general potential energy function, while the antigen-antibody binding has particularity, which makes the MJ potential energy matrix insufficient in specificity for predicting the antigen-antibody docking pose. Based on this, the antigen-antibody interaction potential matrix optimization method based on potential energy difference maximization of some embodiments of the present disclosure first determines the natural complex average energy value of each training sample in the training sample set based on the initialization potential matrix, wherein the initialization potential matrix is composed of the initial potential energy difference between each amino acid in the preset amino acid sequence, and the training sample includes an antigen-antibody complex crystal structure. Here, by determining the natural complex average energy value of the training set, it can be used as a benchmark for optimization to establish a benchmark target for subsequent optimization and ensure that the optimization direction is to reduce the energy of the natural structure. Then, the antigen-antibody complex crystal structure in each training sample in the training sample set is adjusted to generate a set of non-binding conformations. Here, by generating non-binding conformations, the potential energy function provides contrast data for learning how to distinguish between correct and incorrect binding modes. Then, based on the initialization potential matrix, the energy of each non-binding conformation in the non-binding conformation set is calculated to generate a background average energy value. Here, by quantifying the average energy level of negative samples, a key reference benchmark is provided for calculating the energy gap. Then, the energy difference between the natural complex average energy value and the background average energy value is determined as the potential energy difference. Here, because the potential energy difference is determined, it can be used as the core optimization target to optimize the initialization potential matrix. In addition, the contact frequency of residue pairs in each non-binding conformation in the non-binding conformation set is counted to generate a residue frequency matrix. Here, by counting the frequency of non-natural interactions from negative samples, data-driven basis can be provided for updating the initialization potential matrix. Then, the initialization potential matrix is optimized according to the potential energy difference, the residue frequency matrix, and the hydrophilic-hydrophobic parameter constraint to obtain an optimized potential energy matrix, wherein the hydrophilic-hydrophobic parameter constraint is used to limit the optimization range of the energy value in the initialization potential matrix. Here, the hydrophilic-hydrophobic parameter constraint is introduced according to the chemical properties of amino acids, which can be used as a physical constraint to prevent the optimization process from deviating purely under mathematical driving, and to ensure that the final potential energy parameters meet the basic biophysical laws, thereby greatly improving the prediction ability and reliability of the model. Finally, in response to determining that the potential energy difference and the current training round do not satisfy the preset convergence condition, the optimized potential energy matrix is used as the initialization potential matrix to perform initialization potential matrix optimization again.Here, the performance (discrimination ability) of the potential energy matrix is continuously converged to an optimal solution through cyclic iteration, ensuring the stability and reliability of the final result. Thus, the optimized potential energy matrix can have good specificity when predicting the pose of the antigen-antibody docking, thereby improving the accuracy of the prediction of the pose of the antigen-antibody docking. BRIEF DESCRIPTION OF DRAWINGS

[0011] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the attached drawings in which:

[0012] Figure 1 is a flowchart of some embodiments of the antigen-antibody interaction potential energy matrix optimization method based on potential energy difference maximization according to the present disclosure; Figure 2 is an initialization potential energy matrix heat map; Figure 3 is a distance weighting function curve; Figure 4 is an optimized potential energy matrix heat map; Figure 5 is a structural schematic diagram of some embodiments of the antigen-antibody interaction potential energy matrix optimization device based on potential energy difference maximization according to the present disclosure; Figure 6 is a structural schematic diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0013] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only, and are not intended to limit the scope of protection of the present disclosure.

[0014] It should also be noted that, for ease of description, only parts related to the present invention are shown in the drawings. The embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0015] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0016] It should be noted that the modification of "one", "a plurality of" mentioned in the present disclosure is illustrative but not restrictive, and those skilled in the art should understand that unless otherwise explicitly indicated in the context, it should be understood as "one or more".

[0017] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information. The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0018] Figure 1 Flow 100 of some embodiments of the antigen-antibody interaction potential energy matrix optimization method based on potential energy difference maximization according to the present disclosure is shown. The antigen-antibody interaction potential energy matrix optimization method based on potential energy difference maximization includes the following steps: Step 101, based on the initialization potential energy matrix, determining the natural complex average energy value of each training sample in the training sample set.

[0019] In some embodiments, the execution subject (for example, a computing device) of the antigen-antibody interaction potential energy matrix optimization method based on potential energy difference maximization can determine the natural complex average energy value of each training sample in the training sample set based on the initialization potential energy matrix. Wherein the above initialization potential energy matrix is composed of the initial potential energy difference between each amino acid in the preset amino acid sequence, and the training sample includes an antigen-antibody complex crystal structure.

[0020] As an example, see Figure 2The initialization potential energy matrix heat map is shown. The initialization potential energy matrix heat map can be a 20x20 initialization potential energy matrix heat map (MJ Matrix Heatmap) composed of the above-mentioned amino acid sequences (i.e., 20 amino acids). The data of the initialization potential energy matrix is the initial potential energy value between each group of amino acid pairs. Here, the 20 amino acids (Amino Acids) are: CYS (C) - Cysteine, MET (M) - Methionine, PHE (F) - Phenylalanine, ILE (I) - Isoleucine, LEU (L) - Leucine, VAL (V) - Valine, TRP (W) - Tryptophan, TYR (Y) - Tyrosine, ALA (A) - Alanine, GLY (G) - Glycine, THR (T) - Threonine, SER (S) - Serine, ASN (N) - Asparagine, GLN (Q) - Glutamine, ASP (D) - Aspartic acid, GLU (E) - Glutamic acid, HIS (H) - Histidine, ARG (R) - Arginine, LYS (K) - Lysine, PRO (P) - Proline. In addition, the smaller the value in the heat map, the bluer the color, indicating that the interaction tendency of the two amino acids is stronger, and their interaction is more stable; the larger the value, the redder the color, indicating that the interaction of the two amino acids is less stable.

[0021] It should be noted that the above-mentioned computing device can be hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or as a single server or single terminal device. When the computing device is software, it can be installed in the above-mentioned hardware devices listed. It can be implemented as, for example, multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitation is made here.

[0022] In some optional implementations of some embodiments, the subject performer determines the natural complex average energy value of each training sample in the training sample set based on the initialized potential energy matrix, including: Step S1, selecting an amino acid pair that meets the pairing condition from the antigen-antibody complex crystal structure in each training sample in the training sample set to obtain a sample amino acid pair group sequence. For the antigen-antibody complex crystal structure in each training sample, the side chain atoms of the residues with a distance less than or equal to a preset distance (e.g., 6.5 angstroms) can be selected as an amino acid pair from the antigen-antibody complex crystal structure to obtain a sample amino acid pair group.

[0023] Step S2, determining the potential energy value corresponding to each sample amino acid pair group in the sample amino acid pair group sequence to obtain a potential energy value sequence. The potential energy value (energy value) corresponding to each sample amino acid pair can be selected from the initialized potential energy matrix to obtain the potential energy value sequence.

[0024] Step S3, determining the average value of each potential energy value in the potential energy value sequence as the natural complex average energy value.

[0025] Step 102, performing structure adjustment on the antigen-antibody complex crystal structure in each training sample in the training sample set to generate a set of unbound conformations.

[0026] In some embodiments, the subject performer can perform structure adjustment on the antigen-antibody complex crystal structure in each training sample in the training sample set to generate a set of unbound conformations.

[0027] In some optional implementations of some embodiments, the subject performer performs structure adjustment on the antigen-antibody complex crystal structure in each training sample in the training sample set to generate a set of unbound conformations, including: Step S1, for the antigen-antibody complex crystal structure included in each training sample in the training sample set, performing the following structure adjustment steps: First, determine the antigen-antibody center of mass coordinates in the antigen-antibody complex crystal structure included in the training sample. The geometric center of the atomic coordinates of each alpha carbon atom in the antigen-antibody complex crystal structure can be determined as the antigen-antibody center of mass coordinates. Second, moving the antigen-antibody complex crystal structure in space with the antigen-antibody center of mass coordinates as the center to record the temporary conformations after each movement to obtain a set of temporary conformations. The antigen-antibody complex crystal structure can be rotated around the X-Y-Z axis in order with the antigen-antibody center of mass coordinates as the rotation center. The temporary conformations after each rotation (e.g., every 2 degrees) are recorded to obtain a set of temporary conformations.

[0028] In the third step, the above-mentioned temporary conformation set is screened to obtain a non-binding conformation group. Among them, first, the side chain atoms of the residues with a distance less than or equal to the preset distance in each temporary conformation can be determined as a temporary amino acid pair to obtain a temporary amino acid pair group. Second, the temporary conformation corresponding to the temporary amino acid pair group with a number greater than or equal to the preset number of temporary amino acid pairs is determined as a first screening conformation to obtain a first screening conformation group. Then, the root mean square deviation of each first screening conformation in the above-mentioned first screening conformation group is determined to obtain a root mean square deviation group. Then, the above-mentioned root mean square deviation group is clustered to obtain a cluster set. Finally, the first screening conformation with the minimum total potential energy is selected from each cluster in the above-mentioned cluster set as a non-binding conformation set. Here, the total potential energy is the sum of the potential energy values of the temporary amino acid pairs in the first screening conformation.

[0029] In practice, by screening the amino acid pairs, a conformation with a stable interface can be selected at a coarse granularity. Second, by clustering the root mean square deviation, first screening conformations with small structural differences can be divided into the same cluster. Also, the number of first screening conformations with too small structural differences is too redundant for optimization of the potential energy matrix. Therefore, the first screening conformation with the minimum total potential energy is selected as a representative in each cluster to obtain a diversified and high-quality non-binding conformation group. Thus, a conformation more suitable for potential energy matrix optimization can be selected. This facilitates improving the speed and accuracy of potential energy matrix optimization.

[0030] In step S2, each non-binding conformation group obtained is determined as a non-binding conformation set. Among them, each non-binding conformation in each non-binding conformation group can be combined into a non-binding conformation set.

[0031] In step 103, based on the initialization potential energy matrix, energy calculation is performed on each non-binding conformation in the non-binding conformation set to generate a background average energy value.

[0032] In some embodiments, the above-mentioned execution subject can perform energy calculation on each non-binding conformation in the above-mentioned non-binding conformation set based on the above-mentioned initialization potential energy matrix to generate a background average energy value.

[0033] In some optional implementations of some embodiments, the above-mentioned execution subject performs energy calculation on each non-binding conformation in the above-mentioned non-binding conformation set based on the above-mentioned initialization potential energy matrix to generate a background average energy value, including: In step S1, an amino acid pair sequence set is obtained by screening an amino acid pair that meets a pairing condition from each non-binding conformation in the above-mentioned non-binding conformation set. Among them, for each non-binding conformation, the side chain atoms of the residues with a distance less than or equal to the preset distance can be selected as an amino acid pair from the non-binding conformation to obtain an amino acid pair sequence.

[0034] Step S2, determine the background potential energy value corresponding to each amino acid pair sequence in the sequence set, to obtain a background potential energy value sequence. Wherein, the potential energy value corresponding to each amino acid pair can be selected from the initialized potential energy matrix to obtain the background potential energy value sequence.

[0035] Step S3, determine the average value of each background potential energy value in the above background potential energy value sequence as the background average energy value.

[0036] Step 104, determine the energy difference between the natural complex average energy value and the background average energy value as the potential energy difference.

[0037] In some embodiments, the above subject can determine the energy difference between the natural complex average energy value and the background average energy value as the potential energy difference.

[0038] Step 105, perform residue pair contact frequency statistics on each non-binding conformation in the non-binding conformation set to generate a residue frequency matrix.

[0039] In some embodiments, the above subject can perform residue pair contact frequency statistics on each non-binding conformation in the above non-binding conformation set to generate a residue frequency matrix.

[0040] In some optional implementations of some embodiments, the above subject performs residue pair contact frequency statistics on each non-binding conformation in the above non-binding conformation set to generate a residue frequency matrix, comprising: Step S1, determine the Gaussian distance weight of each amino acid in each non-binding conformation in the above non-binding conformation set to obtain a Gaussian distance weight set. Wherein, a Gaussian-type distance weighting function can be used to enhance the weight of short-range interaction to generate a Gaussian distance weight. For example, for each amino acid, if the distance between the two amino acids in the corresponding amino acid pair is less than or equal to a first preset distance (e.g., 4 angstroms), the contact frequency is increased by a weight (e.g., 0.95). If the distance between the two amino acids in the amino acid pair is greater than the first preset distance and less than or equal to a second preset distance (e.g., 5 angstroms), the contact frequency is increased by a weight (e.g., 0.5). If the distance between the two amino acids in the amino acid pair is greater than the second preset distance and less than or equal to a third preset distance (e.g., 6.5 angstroms), the contact frequency is increased by a weight (e.g., 0.1). Thus, the comprehensive distance weight of each amino acid is calculated as the Gaussian distance weight.

[0041] As an example, as Figure 3The distance weighting function curve is shown. Wherein, x represents the distance of two amino acids in the amino acid pair (residue pair distance). y represents the weight value. Specifically, the peak value in the figure (x = 4.0 Å, y ≈ 1.0): it can represent the optimal distance of van der Waals interaction, the lowest energy, and the strongest interaction. Therefore, it is given the highest weight (close to 1). The distance becomes shorter (x < 4.0 Å): when the atomic distance is less than 4.0 Å, strong van der Waals repulsion will be generated, and the interaction becomes unfavorable. Therefore, the weight y starts to decrease from 1. The distance becomes longer (x ≥ 4.0 Å): as the distance increases, the interaction strength decays rapidly. The weight y starts to decrease smoothly from 1, and gradually approaches 0.

[0042] Step S2, determine the co-occurrence matrix corresponding to each Gaussian distance weight in the above Gaussian distance weight set. Wherein, first, arrange each Gaussian distance weight according to the arrangement order of the amino acid sequence into a (20x20) co-occurrence matrix.

[0043] Step S3, normalize the above co-occurrence matrix to obtain a residue frequency matrix.

[0044] Step 106, according to the potential energy difference, the residue frequency matrix and the hydrophilic-hydrophobic parameter constraint, the initialized potential energy matrix is optimized to obtain an optimized potential energy matrix.

[0045] In some embodiments, the execution subject can optimize the initialized potential energy matrix according to the potential energy difference, the residue frequency matrix and the hydrophilic-hydrophobic parameter constraint to obtain an optimized potential energy matrix. Wherein, the hydrophilic-hydrophobic parameter constraint is used to limit the optimization limit of the energy value in the initialized potential energy matrix.

[0046] In some optional implementations of some embodiments, the execution subject optimizes the initialized potential energy matrix according to the potential energy difference, the residue frequency matrix and the hydrophilic-hydrophobic parameter constraint to obtain an optimized potential energy matrix, including: The initialized potential energy matrix is iterated using the residue frequency matrix, the potential energy difference and the objective function to obtain an optimized potential energy matrix. Wherein, the objective function takes the potential energy difference as the maximum optimization target, and updates the initialized potential energy matrix by gradient descent algorithm to obtain the optimized potential energy matrix. Here, the learning rate can be 0.1. The hydrophilic-hydrophobic parameter constraint can be established according to the amino acid pair corresponding to each initialized potential energy value in the initialized potential energy matrix.

[0047] In practice, to prevent excessive deviations from the optimization objective during gradient descent, hydrophilicity / hydrophobicity parameter constraints are introduced. Specifically, for each residue (amino acid) type, parameter constraints can be applied as follows: For hydrophilic residues, a maximum attractive force constraint is applied (e.g., the potential energy value is adjusted upwards (positively) to a range less than or equal to -0.5): This prevents the attractive force between hydrophilic residues (such as ARG, LYS, ASP, GLU, ASN, GLN, SER, THR) from being adjusted too excessively, violating the physical principle that "hydrophilic residues tend to interact with water rather than with each other." For hydrophobic residues, a minimum repulsive force constraint is applied (e.g., the potential energy value is adjusted downwards (negatively) to a range greater than or equal to 0.2): This ensures that the interactions between hydrophobic residues (such as ALA, VAL, ILE, LEU, PHE, TYR, TRP, MET) are always attractive and not incorrectly optimized into repulsion. This ensures that the optimization process is conducted within a reasonable range for antigen-antibody binding, guaranteeing that the final potential energy matrix is ​​not only highly discriminative but also has clear physical meaning.

[0048] As an example, such as Figure 4 The optimized potential energy matrix heatmap shown is shown. Figure 4 In the optimized potential energy matrix (heatmap), the horizontal axis represents the antigen amino acid sequence, and the vertical axis represents the antibody amino acid sequence. Furthermore, the colors in the heatmap characterize the optimized interaction strength between antigen and antibody amino acids. Notably, the interaction strength of the LEU amino acid pair (-19.08) is the strongest, as can be seen from the graph.

[0049] Step 107: In response to the determination that the potential difference and the current training round do not meet the preset convergence conditions, the optimized potential matrix is ​​used as the initial potential matrix, and the initial potential matrix optimization is performed again.

[0050] In some embodiments, the execution subject can, in response to determining that the potential energy difference and the current training round do not satisfy the preset convergence condition, take the optimization potential energy matrix as the initial potential energy matrix and perform initialization potential energy matrix optimization again. The iteration round is preset (e.g., 50 times). The preset convergence condition can be that the current training round reaches the iteration round, and the number of times that the potential energy difference change between adjacent iteration rounds is less than or equal to the preset change threshold is greater than the preset termination number (e.g., 5 times). Here, the change between the two potential energy differences can be determined after each two iterations. If multiple changes are less than or equal to the preset change threshold, it is determined that the iteration converges. Thus, it can be determined that the initialization potential energy matrix training is completed. The optimization potential energy matrix of the current iteration round is stored. In addition, the initialization potential energy matrix optimization can be performed again by taking the optimization potential energy matrix as the initial potential energy matrix, performing the steps 101-106 again to obtain a new optimized potential energy matrix. In this way, the training iteration of the initial potential energy matrix is realized.

[0051] In addition, the optimization potential energy matrix can also be tested to show that the true positive rate is increased by 15% and the false positive rate is reduced by 20%. Here, the true positive rate is the probability of successful prediction. False positive: the probability of failed prediction. Thus, the test steps are as follows: for each natural complex structure in the preset test set, use molecular docking software or perturbation algorithm to generate multiple non-binding conformations. Use the optimized potential energy as the scoring function to calculate the binding interface energy of each test complex and all its corresponding non-binding conformations. Then, if the natural complex conformation is ranked in the top N (e.g., top 10), it is determined that the prediction is successful.

[0052] In practice, the scheme discards the general potential energy function and constructs an iterative optimization framework with the goal of maximizing the energy gap. Through the gradient descent algorithm, the system can automatically learn and adjust the potential energy matrix parameters, so that its prediction ability is significantly enhanced and more specific. Then, the scheme ingeniously combines data-driven and physical constraints. On the one hand, it learns the interaction rules from large-scale antigen-antibody complex crystal structure data; on the other hand, it introduces physical constraints based on the hydrophobicity of amino acids to ensure that the optimization process conforms to the basic principles of biophysics, effectively avoiding overfitting and physically unreasonable results that may be produced by pure data-driven methods. Thus, through the collaborative optimization of machine learning and physical principles, not only is the recognition accuracy and discrimination ability of the potential energy function for the natural conformation greatly improved, but also the physical interpretability of the generated potential energy matrix is ensured, providing an accurate and reliable prediction for antigen-antibody interaction prediction.

[0053] Optionally, in response to determining that the potential energy difference and / or the current training round satisfies the preset convergence condition, a conformation iteration is performed on the target antigen-antibody binder file based on the optimized potential energy matrix to generate an optimal binding conformation. As shown in the test step above, the optimized potential energy matrix is used as a scoring function for conformation prediction of the target antigen-antibody binder file to obtain the optimal binding conformation.

[0054] The above various embodiments of the present disclosure have the following beneficial effects: the antigen-antibody interaction potential matrix optimization method based on potential energy difference maximization of some embodiments of the present disclosure can improve the accuracy of predicting the antigen-antibody docking pose. Specifically, the reason for the decrease in the accuracy of predicting the antigen-antibody docking pose is that the MJ potential energy matrix is constructed by counting the contact frequency of residue pairs in the protein structure, which is a general potential energy function, while the antigen-antibody binding has particularity, which makes the MJ potential energy matrix insufficient in specificity for predicting the antigen-antibody docking pose. Based on this, the antigen-antibody interaction potential matrix optimization method based on potential energy difference maximization of some embodiments of the present disclosure first determines the natural complex average energy value of each training sample in the training sample set based on the initialization potential matrix, wherein the initialization potential matrix is composed of the initial potential energy difference between each amino acid in the preset amino acid sequence, and the training sample includes an antigen-antibody complex crystal structure. Here, by determining the natural complex average energy value of the training set, it can be used as a benchmark for optimization to establish a benchmark target for subsequent optimization and ensure that the optimization direction is to reduce the energy of the natural structure. Then, the antigen-antibody complex crystal structure in each training sample in the training sample set is adjusted to generate a set of non-binding conformations. Here, by generating non-binding conformations, the potential energy function provides contrast data for learning how to distinguish between correct and incorrect binding modes. Then, based on the initialization potential matrix, the energy of each non-binding conformation in the non-binding conformation set is calculated to generate a background average energy value. Here, by quantifying the average energy level of negative samples, a key reference benchmark is provided for calculating the energy gap. Then, the energy difference between the natural complex average energy value and the background average energy value is determined as the potential energy difference. Here, because the potential energy difference is determined, it can be used as the core optimization target to optimize the initialization potential matrix. In addition, the contact frequency of residue pairs in each non-binding conformation in the non-binding conformation set is counted to generate a residue frequency matrix. Here, by counting the frequency of non-natural interactions from negative samples, data-driven basis can be provided for updating the initialization potential matrix. Then, the initialization potential matrix is optimized according to the potential energy difference, the residue frequency matrix, and the hydrophilic-hydrophobic parameter constraint to obtain an optimized potential energy matrix, wherein the hydrophilic-hydrophobic parameter constraint is used to limit the optimization range of the energy value in the initialization potential matrix. Here, the hydrophilic-hydrophobic parameter constraint is introduced according to the chemical properties of amino acids, which can be used as a physical constraint to prevent the optimization process from deviating purely under mathematical driving, and to ensure that the final potential energy parameters meet the basic biophysical laws, thereby greatly improving the prediction ability and reliability of the model. Finally, in response to determining that the potential energy difference and the current training round do not satisfy the preset convergence condition, the optimized potential energy matrix is used as the initialization potential matrix to perform initialization potential matrix optimization again.Here, the performance (discrimination ability) of the potential energy matrix is continuously converged to an optimal solution through iterative circulation, ensuring the stability and reliability of the final result. Thus, the optimized potential energy matrix can have good specificity when predicting the pose of the antigen-antibody docking. Furthermore, the accuracy of the antigen-antibody docking pose prediction is improved. Further referring to Figure 5 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a potential energy difference maximization-based antigen-antibody interaction potential energy matrix optimization device, which corresponds to the method embodiments shown in Figure 1 , and the potential energy difference maximization-based antigen-antibody interaction potential energy matrix optimization device can be applied in various electronic devices.

[0055] As shown in Figure 5 , the potential energy difference maximization-based antigen-antibody interaction potential energy matrix optimization device 500 of some embodiments includes a first determination unit 501, a structure adjustment unit 502, an energy calculation unit 503, a second determination unit 504, a residue docking contact frequency statistical unit 505, an optimization unit 506, and a convergence judgment unit 507. The first determination unit 501 is configured to determine the natural complex average energy value of each training sample in the training sample set based on the initialization potential energy matrix, wherein the initialization potential energy matrix is composed of the initial potential energy difference between each amino acid in the preset amino acid sequence, and the training sample includes an antigen-antibody complex crystal structure. The structure adjustment unit 502 is configured to perform structure adjustment on the antigen-antibody complex crystal structure in each training sample in the training sample set to generate a set of non-binding conformations. The energy calculation unit 503 is configured to perform energy calculation on each non-binding conformation in the set of non-binding conformations based on the initialization potential energy matrix to generate a background average energy value. The second determination unit 504 is configured to determine the energy difference between the natural complex average energy value and the background average energy value as the potential energy difference. The residue docking contact frequency statistical unit 505 is configured to perform residue docking contact frequency statistics on each non-binding conformation in the set of non-binding conformations to generate a residue frequency matrix. The optimization unit 506 is configured to optimize the initialization potential energy matrix according to the potential energy difference, the residue frequency matrix, and the hydrophilic-hydrophobic parameter constraint to obtain an optimized potential energy matrix, wherein the hydrophilic-hydrophobic parameter constraint is used to limit the optimization limit of the energy value in the initialization potential energy matrix. The convergence judgment unit 507 is configured to, in response to determining that the potential energy difference and the current training round do not satisfy the preset convergence condition, use the optimized potential energy matrix as the initialization potential energy matrix and perform initialization potential energy matrix optimization again.

[0056] It can be understood that the units described in the antigen-antibody interaction potential matrix optimization device 500 based on the maximum potential energy difference correspond to the units described in the method Figure 1 The individual steps in the described method correspond. Thus, the operations, features and resulting advantages described above for the method also apply to the antigen-antibody interaction potential matrix optimization device 500 based on the maximum potential energy difference and the units contained therein, which will not be described again here. Reference is made below to Figure 6 which shows a structural schematic diagram of an electronic device (such as a computing device) suitable for implementing some embodiments of the present disclosure. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure. As Figure 6 As shown, the computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory can include a non-volatile storage medium and an internal memory. The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions which, when executed, can cause the processor to perform any of the above methods. The processor is used to provide computing and control capabilities to support the operation of the entire computer device. The internal memory provides an environment for the computer program in the non-volatile storage medium, which, when executed by the processor, can cause the processor to perform any of the above methods. The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present disclosure, and does not constitute a limitation on the computer device to which the scheme of the present disclosure is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0057] It should be understood that the processor can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0058] The processor is configured to determine natural complex average energy values of each training sample in a training sample set based on an initial potential energy matrix, wherein the initial potential energy matrix is composed of initial potential energy differences between each amino acid in a preset amino acid sequence, and the training sample set includes antigen-antibody complex crystal structures; perform structure adjustment on the antigen-antibody complex crystal structures in each training sample in the training sample set to generate a set of unbound conformations; perform energy calculation on each unbound conformation in the set of unbound conformations based on the initial potential energy matrix to generate background average energy values; determine energy differences between the natural complex average energy values and the background average energy values as potential energy differences; perform residue contact frequency statistics on each unbound conformation in the set of unbound conformations to generate a residue frequency matrix; perform optimization on the initial potential energy matrix according to the potential energy differences, the residue frequency matrix, and hydrophilic-hydrophobic parameter constraints to obtain an optimized potential energy matrix, wherein the hydrophilic-hydrophobic parameter constraints are used to limit the optimization boundaries of the energy values in the initial potential energy matrix; and in response to determining that the potential energy differences and the current training round do not satisfy a preset convergence condition, taking the optimized potential energy matrix as the initial potential energy matrix to perform the initial potential energy matrix optimization again.

[0059] The present disclosure also provides a computer readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed, the method can refer to any of the embodiments of the method of the present disclosure.

[0060] The computer readable storage medium can be an internal storage unit of the computer device, such as a hard disk or a memory of the computer device. The computer readable storage medium can also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0061] It should be noted that, in this document, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that processes, methods, articles, or systems that comprise a list of elements do not include only those elements, but also other elements that are not expressly listed, or other elements that are inherent in such processes, methods, articles, or systems. Without more limitations, an element defined by the statement "comprising a" does not exclude the existence of additional identical elements in the process, method, article, or system that includes the element.

[0062] The above description is merely exemplary of some preferred embodiments of the present disclosure and of the principles thereof. It is to be understood that the present disclosure is not limited to the specific technical features described above, and that the scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, but also covers other technical solutions formed by the combinations of the above technical features or equivalent features thereof without departing from the above inventive concept. For example, the above technical features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form technical solutions.

Claims

1. A method for optimizing the potential energy matrix of antigen-antibody interaction based on the maximization of potential energy difference, characterized in that, The method comprises the following steps: determining the natural complex average energy value of each training sample in the training sample set based on the initial potential energy matrix, wherein the initial potential energy matrix is composed of the initial potential energy difference between each amino acid in the preset amino acid sequence, and the training sample includes an antigen-antibody complex crystal structure; adjusting the structure of the antigen-antibody complex crystal structure in each training sample in the training sample set to generate a non-binding conformation set; based on the initial potential energy matrix, energy calculation is performed on each non-binding conformation in the non-binding conformation set to generate a background average energy value; determining the energy difference between the natural complex average energy value and the background average energy value as a potential energy difference; statistical analysis of the residue docking contact frequency of each non-binding conformation in the non-binding conformation set to generate a residue frequency matrix; According to the potential energy difference, the residue frequency matrix and the hydrophilic-hydrophobic parameter constraint, the initial potential energy matrix is optimized to obtain an optimized potential energy matrix, wherein the hydrophilic-hydrophobic parameter constraint is used to limit the optimization limit of the energy value in the initial potential energy matrix; In response to determining that the potential energy difference and the current training round do not satisfy the preset convergence condition, the optimized potential energy matrix is used as the initial potential energy matrix, and the initial potential energy matrix optimization is performed again.

2. The method of claim 1, wherein, The method further comprises: In response to determining that the potential energy difference and / or the current training round satisfies the preset convergence condition, performing conformation prediction on a target antigen-antibody binder file based on the optimized potential energy matrix to generate an optimal binding conformation.

3. The method of claim 2, wherein, The method further comprises: selecting an amino acid pair that satisfies a pairing condition from the antigen-antibody complex crystal structure in each training sample in the training sample set to obtain a sample amino acid pair group sequence; determining the potential energy value corresponding to each sample amino acid pair group in the sample amino acid pair group sequence to obtain a potential energy value sequence; determining the average value of each potential energy value in the potential energy value sequence as the natural complex average energy value.

4. The method of claim 3, wherein, The method further comprises: For the antigen-antibody complex crystal structure included in each training sample in the training sample set, the following structure adjustment steps are performed: determining the antigen-antibody centroid coordinates in the antigen-antibody complex crystal structure included in the training sample; spatially moving the antigen-antibody complex crystal structure included in the training sample with the antigen-antibody centroid coordinates as the center to record the temporary conformations after each movement to obtain a temporary conformation set; screening the temporary conformation set to obtain a non-binding conformation group; determining each obtained non-binding conformation group as a non-binding conformation set.

5. The method of claim 4, wherein, The method further comprises: screening the non-binding conformations in the non-binding conformation set to obtain an amino acid pair sequence set that satisfies a pairing condition; determine a background potential energy value corresponding to each amino acid pair sequence in the set of amino acid pair sequences, to obtain a sequence of background potential energy values; determine an average value of each background potential energy value in the sequence of background potential energy values as a background average energy value.

6. The method of claim 5, wherein, the residue pair contact frequency statistics unit is configured to perform residue pair contact frequency statistics on each of the set of non-binding conformations to generate a residue frequency matrix, including: determine a Gaussian distance weight of each amino acid in each of the set of non-binding conformations to obtain a set of Gaussian distance weights; determine a co-occurrence matrix corresponding to each Gaussian distance weight in the set of Gaussian distance weights; normalize the co-occurrence matrix to obtain a residue frequency matrix.

7. The method of claim 6, wherein, the optimization unit is configured to optimize the initial potential energy matrix according to the potential energy difference, the residue frequency matrix, and a hydrophilic-hydrophobic parameter constraint to obtain an optimized potential energy matrix, wherein the hydrophilic-hydrophobic parameter constraint is used to limit the optimization range of the energy values in the initial potential energy matrix. the optimization unit is configured to optimize the initial potential energy matrix according to the potential energy difference, the residue frequency matrix, and a hydrophilic-hydrophobic parameter constraint to obtain an optimized potential energy matrix, wherein the hydrophilic-hydrophobic parameter constraint is used to limit the optimization range of the energy values in the initial potential energy matrix.

8. An apparatus for optimizing a potential energy matrix of an antigen-antibody interaction based on maximization of potential energy difference, characterized in that, including: the first determination unit is configured to determine a natural complex average energy value of each training sample in the training sample set based on the initial potential energy matrix, wherein the initial potential energy matrix is composed of initial potential energy differences between each amino acid in a preset amino acid sequence, and the training sample includes an antigen-antibody complex crystal structure; the structure adjustment unit is configured to perform structure adjustment on the antigen-antibody complex crystal structure in each training sample in the training sample set to generate a set of non-binding conformations; the energy calculation unit is configured to perform energy calculation on each of the set of non-binding conformations based on the initial potential energy matrix to generate a background average energy value; the second determination unit is configured to determine an energy difference between the natural complex average energy value and the background average energy value as a potential energy difference; the residue pair contact frequency statistics unit is configured to perform residue pair contact frequency statistics on each of the set of non-binding conformations to generate a residue frequency matrix; the optimization unit is configured to optimize the initial potential energy matrix according to the potential energy difference, the residue frequency matrix, and a hydrophilic-hydrophobic parameter constraint to obtain an optimized potential energy matrix, wherein the hydrophilic-hydrophobic parameter constraint is used to limit the optimization range of the energy values in the initial potential energy matrix; the convergence judgment unit is configured to, in response to determining that the potential energy difference and the current training round do not satisfy a preset convergence condition, take the optimized potential energy matrix as the initial potential energy matrix and perform initial potential energy matrix optimization again.

9. An electronic device, comprising: including: one or more processors; a storage device having one or more programs stored thereon, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1-7.

10. A computer readable medium characterized by a computer program is stored thereon, wherein the program is executed by a processor to implement the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Method and system for predicting antigen-antibody binding affinity based on ensemble learning

    CN117935925A

  • Design method of protein variant structure based on molecular dynamics

    CN120600106A

  • Device and method for evaluating candidate structure of antigen antibody complex

    JP2012145349A

  • Method for binding site identification

    US20050114035A1

  • Method and system for binding affinity prediction and method of generating a candidate protein-binding peptide

    US20220208301A1