A protein-ligand binding affinity prediction method, system, device and medium based on Wasserstein hidden distribution alignment
Patent Information
- Application Number
- CN202610706812.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-08-21
AI Technical Summary
当测试样本在蛋白质家族、配体骨架或结合口袋几何结构等方面与训练数据存在显著差异时,模型对于蛋白质-配体的结合亲和力预测性能泛化不足
本发明中将蛋白质-配体复合物的特征编码为Wassertein空间中的高斯分布,并通过基于Wasserstein重心投影,将蛋白质-配体特征分布对齐到参考分布,实现特征分布的对齐,从而有效减少不同结构复合物之间的表征差异,以提升表征的泛化性。本发明能够在不改变现有预测模型结构的前提下实现蛋白质-配体复合物隐表征对齐到参考分布,从而有效降低训练分布与实际应用数据分布之间的差异,提升在结构多样性较高及分布外复合物样本上的预测准确性与泛化能力。
Smart Images

Figure CN122619092A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer biotechnology, specifically relating to a method, system, device, and medium for predicting protein-ligand binding affinity based on Wasserstein hidden distribution alignment. Background Technology
[0002] Structure-based drug design (SBDD) aims to develop small molecule drugs that bind with high affinity to specific protein targets, and is one of the important methods in modern drug development. With the rapid development of computational methods and deep neural networks, SBDD is gradually entering an intelligent stage, bringing new possibilities to computer-aided drug design (CADD). For example, molecular docking algorithms can embed candidate drugs into the binding sites of target proteins and predict the binding posture; emerging protein folding models (such as AlphaFold3) can consider small molecule ligands to predict the binding conformation. However, although these technologies provide diverse tools for drug design, their ability to truly advance the key processes of drug development still largely depends on the accurate prediction of protein-ligand binding affinity. In practical applications, researchers generally use scoring functions to quantitatively evaluate the binding stability of candidate complexes. Traditional physics- and experience-driven scoring functions (such as AutoDock Vina and GOLD), while computationally efficient, have limitations in predictive accuracy.
[0003] In recent years, deep learning-based scoring functions have gradually become the mainstream approach. These models typically utilize 3D convolutional neural networks or graph neural networks to model the spatial structure and interactions of protein-ligand complexes, achieving significant performance improvements on multiple benchmark datasets. However, subsequent research has revealed that these models suffer from insufficient generalization in real-world drug discovery scenarios, and their predictive capabilities struggle to consistently transfer to new protein-ligand complexes. Further research indicates that this performance gap primarily stems from a significant distributional difference between training data and real-world application scenarios. On one hand, data leakage may exist between training and testing data, leading to an overestimation of model performance; on the other hand, even with rigorous data cleaning and partitioning strategies, deep learning models are still prone to learning dataset-specific statistical characteristics rather than truly universal protein-ligand interaction patterns. To alleviate these issues, previous research has proposed more stringent data partitioning and evaluation strategies, such as PDBbind CleanSplit, and has built more robust scoring frameworks (such as Graph neural network for Efficient Molecular Scoring, GEMS) based on these strategies, improving model performance on independent test sets to some extent. While rigorous data cleaning can reduce overlap between training and testing samples, significantly mitigating evaluation bias caused by sample overlap at the data level, it fails to fundamentally overcome the adaptability problem of deep learning models when facing out-of-distribution (OOD) samples during the inference phase. Given the near-infinite possibilities of chemical and protein structural spaces, models built with limited training data cannot cover all structural patterns. Therefore, when models encounter new protein families, ligand backbones, or binding pocket structures, their predictive performance may still face significant degradation.
[0004] Although existing research has made some progress in mitigating data breaches and improving model generalization capabilities, the following prominent problems still exist with current technologies: (1) The generalization ability of the model in out-of-distribution scenarios remains limited: After adopting a more stringent data partitioning strategy, many models that previously performed well on benchmark data showed a significant performance decline. This indicates that existing models often rely on the structural similarity between the training set and the test set, and fail to learn transferable protein-ligand interaction rules. When the test samples differ significantly from the training data in terms of protein family, ligand backbone, or binding pocket geometry, the model's generalization performance for predicting protein-ligand binding affinity is insufficient.
[0005] (2) Lack of modeling mechanism for feature distribution shift: Existing deep learning scoring functions usually only focus on minimizing prediction error on training data, which may make the model too dependent on the statistical features of the training set. When the model encounters a novel protein-ligand complex with a large difference in structure or chemical properties from the training data during the inference stage, its representation may deviate from the training distribution, resulting in unstable prediction results.
[0006] In summary, while existing methods have improved the predictive performance of scoring models to some extent, their generalization ability remains limited when faced with novel protein-ligand complexes with high structural diversity. Therefore, how to effectively alleviate the feature distribution shift problem at the model level, thereby improving the generalization ability of deep learning scoring functions in out-of-distribution scenarios, has become a crucial problem that urgently needs to be solved in the current CADD field. Summary of the Invention
[0007] This invention addresses the limitation of generalization ability in existing technologies when dealing with novel protein-ligand complexes with high structural diversity. It proposes a protein-ligand affinity prediction method, system, device, and medium based on Wasserstein Latent Distribution Alignment (WLDA). This method encodes the features of protein-ligand complexes into a Gaussian distribution in Wasserstein space and aligns the feature distributions through Wasserstein centroid projection, thereby effectively reducing characterization differences between complexes with different structures and improving stability and generalization ability on out-of-distribution samples.
[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows: The first aspect of this invention provides a method for predicting protein-ligand binding affinity based on Wasserstein latent distribution alignment, comprising the following steps: The protein-ligand complex is encoded as a Gaussian distribution in Wasserstein space, and the protein-ligand feature distribution is obtained. Establish a reference distribution in Wasserstein space for protein-ligand characteristic distributions; The protein-ligand feature distribution is aligned to the reference distribution in Wasserstein space based on Wasserstein barycentric projection, and the aligned protein-ligand complex characterization is obtained. The aligned protein-ligand complex characterization is input into the optimized binding affinity prediction module for prediction, and the protein-ligand binding affinity is obtained.
[0009] Furthermore, the protein-ligand complex is encoded as a Gaussian distribution in Wasserstein space to obtain the protein-ligand feature distribution, including the following steps: The protein-ligand complex is represented by an interaction diagram; The latent representation set of each node in the complex is obtained by extracting features from the graph structure of the interaction graph through the encoder. The protein-ligand feature distribution is obtained based on the hidden representation set of each node in the complex.
[0010] Furthermore, protein-ligand characteristic distribution Calculated using the following formula:
[0011] ,
[0012] in, The mean vector representing the characteristics of the complex; The variance vector representing the feature; Denotes the diagonal covariance matrix. It follows a Gaussian distribution. For the first i The total number of nodes in a complex i This is the complex number. j This represents the node number within the complex. For the first The embedding vector of each node. represents the parameters of a Gaussian distribution.
[0013] Furthermore, the reference distribution of the Wasserstein space is established by the following equation:
[0014] in, For reference distribution, For the reference distribution One portion, The mean, For variance, Parameters for the reference distribution It follows a Gaussian distribution. It is a diagonal covariance matrix. The component index in the reference distribution. This represents the total number of components in the reference distribution.
[0015] Furthermore, based on the Wasserstein barycentric projection, the protein-ligand feature distribution is aligned to a reference distribution in Wasserstein space to obtain the aligned protein-ligand complex characterization, including the following steps: Based on the training sample distribution and the reference distribution of the Wasserstein space, the objective function containing the entropy regularization term is obtained. Solving the objective function containing the entropy regularization term yields the optimal transfer matrix; Based on the optimal transfer matrix, the protein-ligand complex feature distribution is mapped to the reference distribution through Wasserstein centroid projection using the Gaussian reference centroid alignment mechanism to obtain the protein-ligand feature distribution parameters. Based on the protein-ligand feature distribution parameters, the protein-ligand node features are mapped to the aligned distribution space to obtain the aligned protein-ligand complex characterization.
[0016] Furthermore, the objective function containing the entropy regularization term is:
[0017] satisfy:
[0018] in, This is the complex number. The component index in the reference distribution. For the training sample distribution, For the reference distribution One portion, For the training sample distribution and the reference distribution, the th Transmission weights between components The squared 2-Wasserstein distance Here is the entropy regularization coefficient. B Batch size; The training sample distribution is as follows:
[0019] in, It follows a Gaussian distribution. The mean vector representing the characteristics of the complex; The variance vector representing the feature; Let represent the diagonal covariance matrix.
[0020] Furthermore, the optimal transfer matrix is:
[0021] in, For the optimal transfer matrix, This indicates the index of the reference distribution component used in the denominator normalization term. For reference distribution components; The loss function for the affinity prediction module is:
[0022] in, For loss function, To incorporate the affinity regression loss function, As a reference for distributed complexity control terms, The features of the complex after being encoded by the encoder. This is an interaction diagram of the complex. These are weight hyperparameters; The reference distributed complexity control terms are:
[0023] in, It is the numerical stability constant. Indicates the first The average transmission quality received in the current batch by a reference distribution.
[0024] A second aspect of the present invention provides a protein-ligand binding affinity prediction system based on Wasserstein hidden distribution alignment, comprising: The encoding module is used to encode the protein-ligand complex as a Gaussian distribution in Wasserstein space to obtain the protein-ligand feature distribution. The Wasserstein space reference distribution establishment module is used to establish a reference distribution in Wasserstein space for protein-ligand feature distributions. The alignment module is used to align the protein-ligand feature distribution to a reference distribution in Wasserstein space based on Wasserstein barycentric projection, so as to obtain the aligned protein-ligand complex characterization. The prediction module is used to input the aligned protein-ligand complex characterization into the optimized binding affinity prediction module for prediction, so as to obtain the protein-ligand binding affinity.
[0025] A third aspect of the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the protein-ligand binding affinity prediction method based on Wasserstein latent distribution alignment.
[0026] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment.
[0027] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention encodes the features of protein-ligand complexes as a Gaussian distribution in Wasserstein space. By using Wasserstein centroid projection, the protein-ligand feature distribution is aligned to a reference distribution, effectively reducing the representational differences between complexes with different structures and improving the generalization ability of the representation. This invention can align the latent representation of protein-ligand complexes to a reference distribution without changing the structure of existing prediction models, thereby effectively reducing the difference between the training distribution and the actual application data distribution, and improving the prediction accuracy and generalization ability on complex samples with high structural diversity and out-of-distribution features.
[0028] Furthermore, by introducing a Gaussian reference centroid alignment mechanism to adjust the effective utilization of the reference distribution, the affinity prediction accuracy is improved while ensuring the stability of the distribution alignment, thereby enhancing the overall stability and generalization ability.
[0029] Furthermore, this invention introduces a reference distribution composed of multiple learnable Gaussian distributions into the latent space, and uses Wasserstein optimal transport to calculate the mapping relationship between the training sample distribution and the reference distribution. By using centroid projection to align the distribution of complex features, unknown complexes can be mapped to a stable reference distribution space, thereby significantly improving the prediction stability and generalization ability on out-of-distribution complexes. This overcomes the problem that existing deep learning affinity prediction models usually rely on the training data distribution for learning, and the model prediction performance tends to decline when encountering protein-ligand complexes with a large difference from the training set distribution.
[0030] Furthermore, this invention represents the node embedding of protein-ligand complexes as a Gaussian probability distribution and performs unified modeling in Wasserstein space. This provides a unified description of the characteristic structures of complexes of different scales at the distribution level, achieving a stable expression of the overall statistical characteristics of the complexes. This enhances the module's ability to characterize structurally diverse complexes and overcomes the problem that protein-ligand affinity prediction methods typically learn complex features directly in Euclidean space. However, due to significant differences in structural scale and interaction patterns among different complexes, their node numbers and feature distributions are often inconsistent, thus affecting the model's stable characterization ability for complex structures.
[0031] Furthermore, this invention introduces a reference distribution complexity control term to constrain the usage pattern of the reference distribution during training, concentrating optimal transmission quality on a small number of effective reference distributions, thereby reducing the complexity of the reference space and improving alignment stability. This mechanism can enhance the stability of module training while ensuring affinity prediction accuracy, and further improve the overall generalization ability, overcoming the problem that if the reference distribution structure is too complex during the latent distribution alignment process, it may lead to model training instability and affect prediction performance. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of the protein-ligand affinity prediction method of the present invention; Figure 2 The flowchart shows the protein-ligand binding affinity prediction method based on Wasserstein latent distribution alignment. Figure 3 The results of hyperparameter analysis using the WLDA method are shown, where (a) represents the influence of the temperature coefficient λ on the four evaluation indicators of the model; and (b) represents the influence of the number of learnable reference distributions K on the four evaluation indicators of the model. Figure 4 This is a schematic diagram of a protein-ligand binding affinity prediction system based on Wasserstein latent distribution alignment. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0035] See Figure 1 and Figure 2 This invention proposes a protein-ligand binding affinity prediction method based on Wasserstein Latent Distribution Alignment (WLDA) to improve the prediction stability and generalization ability of deep learning models on protein-ligand complexes with high structural diversity and out-of-distribution features. This method alleviates the offset problem between the training data distribution and the actual application data distribution by uniformly modeling and aligning the distribution of protein-ligand complexes in the feature space. The method includes four key steps: 1) Gaussian distribution representation of protein-ligand complexes in Wasserstein space: Specifically, the training samples in the protein-ligand complex dataset are encoded into a Gaussian distribution in Wasserstein space by an encoder to obtain the protein-ligand feature distribution, thereby achieving a unified latent representation of complexes of different structural sizes. 2) Establish a reference distribution in Wasserstein space for the protein-ligand feature distribution to provide a stable reference distribution for training and testing samples in the protein-ligand complex dataset; 3) Align the protein-ligand feature distribution to the reference distribution in Wasserstein space based on Wasserstein barycentric projection to obtain the aligned protein-ligand complex characterization. Specifically, based on semi-relaxed optimal transport and Wasserstein barycentric projection, align the protein-ligand feature distribution characterization of the training samples to the reference distribution composed of multiple learnable Gaussian distributions. 4) The aligned protein-ligand complex characterization is input into the binding affinity prediction module for prediction. The encoder, binding affinity prediction module and reference distribution parameters are optimized by loss function to ultimately improve the accuracy of protein-ligand binding affinity prediction.
[0036] Through the above steps, the present invention can align the hidden representation of protein-ligand complexes to the reference distribution without changing the original prediction model structure, thereby effectively reducing the difference between the training distribution and the actual application data distribution, and improving the prediction accuracy and generalization ability of the model on complex samples with high structural diversity and out-of-distribution.
[0037] The specific process of step 1) is as follows: This step extracts structural and chemical features from training samples in a large-scale known protein-ligand complex dataset, and processes the complex using a GNN (Graph Neural Networks) encoder. x i Representation learning is performed to obtain node features that can describe the interaction environment of the complex.
[0038] Because different protein-ligand complexes have varying structural sizes, the number of nodes in their interaction graphs is not fixed, making it difficult to uniformly represent them in a fixed-dimensional Euclidean space. Therefore, this method does not perform simple vector-level alignment of node representations. Instead, it elevates the implicit node representations of the complexes to probabilistic measures and further maps them to Wasserstein space, performing unified modeling of different complexes at the distribution level. This achieves a unified representation of structures of different sizes and provides a foundation for subsequent distribution alignment and modeling. Specifically, each complex... x i It was first represented as an interaction diagram. , For nodes in the interaction graph, These are the edges of an interaction graph that integrates the interactions between ligand atoms, protein residues, and their spatial relationships. (This is achieved through an encoder.) Feature extraction is performed on the graph structure of the interaction graph. Given the parameters of a Gaussian distribution, we obtain the set of hidden representations for each node in the complex. ,in , indicating the first Hidden representation of an atom or residue, for 3D real vector space, To represent the dimension implicitly, For the first i The total number of nodes in a complex i This is the complex number. j This represents the node number within the complex.
[0039] To characterize the statistical features of the complex in the latent space from a holistic perspective, this method first represents the node embedding vector as an empirical measure. :
[0040] in, Indicates the first The number of nodes in a complex Indicates the first The embedding vector of each node. Indicates the first Embedding vectors of nodes The Dirac measure is centered. Therefore, each protein-ligand complex is represented as a single probability measure in Wasserstein space, regardless of its size.
[0041] To improve the stability and computational efficiency during model training, a diagonal Gaussian distribution is used to approximate the aforementioned empirical measure. This distribution serves as an alternative representation of the empirical measure in Wasserstein space, namely the training sample distribution. :
[0042] in, The distribution is Gaussian, and the mean and variance of the Gaussian distribution are obtained through the first... The embedding vectors of each node are statistically calculated to obtain: ,
[0043] in, The mean vector representing the characteristics of the complex; The variance vector representing the feature; Let represent the diagonal covariance matrix. This provides a tractable representation of each protein-ligand complex in Wasserstein space.
[0044] Through the above processing, the set of embedding vectors for each node of the complex can be transformed into a Gaussian probability distribution in the Wasserstein latent space, thus forming the latent distribution set of the training data. This latent distribution set not only reflects the overall characteristics of the complex structure and chemical environment, but also provides a basic representation for the subsequent construction of the reference distribution space.
[0045] The specific process of step 2) is as follows: To achieve a unified representation of protein-ligand feature distributions in Wasserstein space and reduce distribution differences between different protein-ligand complexes, this step introduces a learnable reference distribution in the latent space to construct a stable reference distribution structure and align the protein-ligand feature distributions.
[0046] in, For the total number of components in the reference distribution, the nth component in each reference distribution Each component It is a measure in Wasserstein space. For reference distribution, k The component index in the reference distribution. For nodes The Dirac measure centered on [the target group].
[0047] To achieve efficient computation and stable optimization, each reference distribution component is parameterized as a diagonal Gaussian distribution, i.e.:
[0048] in, The mean, For variance and mean and variance For learnable parameters, Parameters for the reference distribution This is a diagonal covariance matrix, where the elements on the diagonal are the variances. The corresponding component.
[0049] This step represents the embedding vectors of protein-ligand complex nodes as Gaussian probability distributions and performs unified modeling in Wasserstein space, thereby achieving a unified description of the characteristic distributions of complexes of different structural sizes.
[0050] The specific process of step 3) is as follows: After obtaining the protein-ligand feature distribution and the reference distribution, the protein-ligand feature distribution is aligned to the reference distribution to improve the generalization of the representation. Specifically, this step first involves adjusting the distribution based on the training sample distribution. Compared with the reference distribution The latent distribution of training samples is calculated using semi-relaxed optimal transport. Compared with the reference distribution, the first Each component The alignment relationship between them, i.e., the objective function:
[0051] And satisfy the constraints:
[0052] in, For the training sample distribution, For the reference distribution One portion, Represents the distribution of training samples Compared with the reference distribution, the first Each component Transmission weights between B For batch size, Represents the squared 2-Wasserstein distance:
[0053] in, The standard deviation of the training sample distribution. The standard deviation of the reference distribution component.
[0054] To improve numerical stability, an entropy regularization term is added to the objective function:
[0055] satisfy:
[0056] in, Here, represents the entropy regularization coefficient. The entropy regularization term is used to improve the numerical stability of the transfer matrix solution. Under the entropy regularization condition, the above objective function has a closed-form solution:
[0057] in, For the optimal transfer matrix, To indicate the index of the reference distribution component used in the denominator normalization term, The optimal transfer matrix is obtained by using the reference distribution components. It can be used to construct a mapping relationship between protein-ligand characteristic distributions and reference distributions.
[0058] In obtaining the optimal transmission matrix Subsequently, the feature distribution of each protein-ligand complex is mapped to the reference distribution through a Gaussian reference centroid alignment mechanism, thereby aligning the node features:
[0059] in, To perform Gaussian reference barycentric alignment mapping on the training sample distribution, It follows a Gaussian distribution. The mean of the training sample distribution is obtained by mapping it through a Gaussian reference centroid alignment mechanism. The variance of the training sample distribution is obtained by mapping it through the Gaussian reference centroid alignment mechanism. The mean and variance of the aligned node features are specifically calculated by weighted combination of the components in the reference distribution:
[0060]
[0061] in, For weights, specifically , For reference distribution number The variance of each component. After obtaining the aligned protein-ligand feature distribution parameters, this method further maps the protein-ligand node features to the aligned protein-ligand complex characterization. Specific mapping function. Defined as:
[0062] in, This represents element-wise multiplication. The standard deviation of the training sample distribution is obtained by mapping it using the Gaussian reference centroid alignment mechanism. Through the Wasserstein centroid projection described above, the nodal features of the protein-ligand complex can be projected onto the reference distribution, providing a generalized characterization for subsequent affinity prediction.
[0063] This step, based on the complex feature alignment method using a learnable reference Gaussian distribution, constructs a reference distribution in the latent representation space consisting of multiple learnable Gaussian measures. The latent distribution of the complex is aligned to the reference space through Wasserstein optimal transport and centroid projection, thereby improving the model's adaptability to samples with distribution differences.
[0064] The specific process of step 4) is as follows: After aligning the distribution of protein-ligand node features, binding affinity is predicted using the aligned complex characterization, and model parameters are optimized by calculating the training loss. Specifically, the aligned protein-ligand node features are input into the binding affinity prediction module. Combining affinity prediction module Directly outputs the binding affinity of the complex. :
[0065] in, This is a Gaussian reference centroid alignment mapping function used to align protein-ligand feature distributions.
[0066] loss function Defined as:
[0067] in, This indicates that the affinity regression loss function is used, specifically the root mean square error (RMSE). As a reference for distributed complexity control terms, The features of the complex after being encoded by the encoder. This is an interaction diagram of the complex. Here are the weight hyperparameters. To describe the frequency of use of each reference distribution in the batch of training samples, we first define the average transmission quality:
[0068] in, Indicates the first The average transmission quality received by each reference distribution in the current batch. Further, the reference distribution complexity control term... Defined as:
[0069] in, As a numerical stability constant, by minimizing the above-mentioned binding affinity regression loss function, the overall usage pattern of the reference distribution can be adjusted while ensuring the accuracy of affinity prediction, so that the transmission quality is concentrated on a small number of effective reference distributions, thereby reducing the complexity of the reference distribution space and improving the distribution alignment capability of cross-structure complexes.
[0070] In this step, by introducing a reference distribution complexity control mechanism into the training objective, the affinity prediction accuracy is improved while ensuring the stability of distribution alignment, thereby enhancing the overall stability and generalization ability of the model.
[0071] Therefore, this method can not only achieve effective alignment of the latent distribution of protein-ligand complexes during the training phase, but also adaptively adjust its characterization according to the characteristic distribution of protein-ligand complexes during the inference phase, thus maintaining a stable and accurate binding affinity prediction capability when faced with unseen structures or chemical types.
[0072] The following is an experimental verification of the method of the present invention: This invention provides a comprehensive evaluation of the proposed WLDA-based protein-ligand binding affinity prediction method. First, the datasets used for training and testing, along with the corresponding evaluation protocols, are introduced to ensure the fairness and reproducibility of the experimental results. Then, the generalization ability of WLDA on unknown protein-small molecule complexes is validated, and its binding affinity prediction performance on out-of-distribution (OOD) samples is examined. Finally, this method is systematically compared with models such as SIGN, EHIGN, GraphLambda, CurvAGN, DEAttentionDTA, MFE, and GEMS. Experimental results show that WLDA has significant advantages in prediction accuracy and generalization ability.
[0073] 1.1 Dataset (1) Training dataset This invention uses the PDBbind v2020 dataset as the training set, which contains 19,443 protein-ligand complexes with experimentally determined binding affinity. To eliminate structural leakage between the PDBbind dataset and the CASF benchmark, this invention follows a filtering protocol based on the GEMS method to remove complexes in PDBbind v2020 that are structurally similar to the CASF benchmark, ultimately obtaining a leakage-free training set of 16,491 complexes. All experimentally determined values (Ki, Kd, IC50) are uniformly converted to pK values (pK = ...). log10(value) ensures the numerical consistency of the training data.
[0074] (2) Test dataset The CASF-2013 benchmark dataset consists of 195 high-quality protein-ligand complexes, derived from the core set of PDBbind v2013. This dataset includes experimentally determined binding affinity (pK) values, covers diverse protein families and small molecule types, and is commonly used to assess the generalization ability of scoring functions. Compared to CASF-2007, CASF-2013 offers improvements in both sample size and chemical diversity, making it more suitable as an independent test set.
[0075] The CASF-2016 benchmark set is currently the most widely used version, constructed from a core subset of PDBbind v2016, and contains 285 high-resolution protein-ligand complexes. This version boasts a larger sample size, covers a wide range of protein types and ligand chemistry, and is equipped with high-quality experimental binding affinity data, making it a standard test set for evaluating model generalization ability and task adaptability.
[0076] CASF-2016-indep: This independent subset of CASF-2016, containing 155 protein-ligand complexes, was selected from the full benchmark set through a more rigorous structural similarity filtering algorithm. This subset provides a more challenging testing task to evaluate the model's ability to generalize to novel, unseen protein-ligand interactions.
[0077] CASF-2013-indep: The CASF-2013 independent subset was also obtained through the rigorous filtering process described above, consisting of 89 complexes completely unrelated to the PDBbind training set. As an important supplement to the CASF-2016 independent subset, it is used to evaluate the model's generalization ability to novel, unseen protein-ligand interactions.
[0078] Overall, both CASF-2013 and CASF-2016 serve as out-of-distribution (OOD) evaluation benchmarks, with stricter similarity thresholds applied to their independent subsets. CASF-2016-indep and CASF-2013-indep, which are completely independent of the PDBbindv2020 training set, provide a fair test of the model's generalization to predictions of unknown protein-ligand binding affinity.
[0079] (3) Data preprocessing For each protein-ligand complex, this invention integrates an interaction graph of ligand atoms, protein residues, and their spatial contact information. Ligands are represented at the atomic level, with chemical bonds forming undirected edges. Physicochemical descriptors of the ligands (e.g., hybridization type, aromaticity, and bond order) are calculated using RDKit, and self-loop edges are added to enrich the graph's structural information. To enhance the semantic features of the ligands, ChemBERTa-2 (77M-MLM) embedding vectors generated by SMILES are also appended to the graph as global features. Proteins are modeled at the residue level. Residues within 5 Å of any ligand atom have their embedding vectors extracted using an allocation language model, including concatenated embedding vectors of ESM-2 (T6 8M) and ANKH, while one-hot amino acid vectors are appended as node features. Transmolecular edges are added between ligand atoms and their contacting residues (≤5 Å). Each transmolecular edge contains a four-dimensional geometric descriptor, defined by the main chain atoms (N, C) of that residue. α , C) and virtual C β The distance between the atom and the ligand atom constitutes the equation.
[0080] (4) Model Implementation and Training WLDA is implemented based on the GEMS architecture, a graph neural network (GNN) architecture for predicting protein-ligand binding affinity. In WLDA, a Gaussian reference centroid alignment (GRBA) module is inserted after the GEMS encoder. The aligned node embedding vectors are then used to output the predicted values through the binding affinity prediction head. WLDA is implemented in PyTorch with the hyperparameter β set to 0.01. All experiments were performed on a single NVIDIA Tesla V100 GPU. Following the GEMS protocol, five-fold cross-validation was used. In each fold, a model was trained for a total of 600 training epochs, using SGD as the optimizer with a learning rate of 1 × 10⁻⁶. - ³, with a batch size of 256. The final prediction result is obtained by averaging the outputs of the five models, each trained on a different training subset and evaluated on a corresponding validation subset.
[0081] (5) Evaluation indicators Model performance was evaluated using four standard regression metrics: root mean square error (RMSE), mean absolute error (MAE), Pearson correlation coefficient (R), and standard deviation of regression residuals (SD). SD was calculated by performing a linear regression on the predicted values and the true pK, reflecting prediction accuracy and stability. To compare and validate WLDA's performance, this invention selected four representative protein-ligand binding affinity prediction models as baselines: interaction-based GNN models (SIGN, EHIGN, GraphLambda, CurvAGN), sequence attention models (DEAttentionDTA), multimodal learning models (MFE), and language model-enhanced GNNs (GEMS), to comprehensively evaluate WLDA's prediction accuracy and generalization ability on standard benchmark sets and independent subsets.
[0082] 1.2 Evaluation of the effectiveness of the proposed method This invention fully validates WLDA on CASF-2013, CASF-2016 and their independent subsets, and compares it with multiple baseline models, including interaction-based GNNs (SIGN, CurvAGN, EHIGN, GraphLambda), sequence attention models (DEAttentionDTA), multimodal models (MFE), and GEMS which integrates language models.
[0083] Experimental results show that WLDA achieves superior performance on all benchmark sets and independent subsets (see Tables 1 and 2). On the CASF-2013 and CASF-2016 benchmark sets, WLDA outperforms other methods in metrics such as RMSE, MAE, Pearson correlation coefficient, and SD. Compared with interaction-based GNN models, WLDA achieves higher correlation and lower error across multiple metrics; compared with multimodal and attention models, such as MFE and DEAttentionDTA, WLDA also demonstrates higher prediction accuracy; compared with GEMS, which integrates language models, WLDA further improves prediction performance, fully demonstrating its effectiveness on standard benchmarks.
[0084] Table 1. Performance validation of the WLDA method and multi-class models on the CASF-2013 dataset.
[0085] Table 2. Performance validation of the WLDA method and multi-class models on the CASF-2016 dataset.
[0086] On independent subsets, WLDA demonstrates significantly better generalization ability (see Tables 3 and 4). On CASF-2013-indep, the RMSE was 1.456, MAE was 1.171, Pearson correlation coefficient was 0.787, and SD was 1.434; on CASF-2016-indep, WLDA's RMSE was 1.322, MAE was 1.035, Pearson correlation coefficient was 0.833, and SD was 1.290. This indicates that WLDA exhibits excellent generalization ability for novel, unseen protein-small molecule complexes, effectively mitigating the bias between the training and test distributions.
[0087] Table 3. Generalization performance verification of the WLDA method on the CASF-2013-indep dataset.
[0088] Table 4. Generalization performance verification of the WLDA method on the CASF-2016-indep dataset.
[0089] Ablation experiments further validated the contributions of each component of the WLDA model to its performance. Referring to Table 5, removing the GRBA module increased WLDA's RMSE, MAE, and SD, while decreasing the Pearson correlation coefficient, indicating that the GRBA module significantly contributed to prediction accuracy and stability; removing the effective complexity control term ( This will also lead to a performance degradation, albeit a small one, demonstrating that the effective complexity of controlling the reference distribution helps maintain alignment stability, thereby ensuring generalization ability.
[0090] Table 5. Experimental results of WLDA ablation on the CASF-2016 dataset.
[0091] As shown in Table 6, on the validation set within the training distribution, WLDA's performance did not decrease after introducing GRBA; in fact, it slightly improved in RMSE, MAE, SD, and Pearson correlation coefficient, indicating that the distribution alignment mechanism does not affect the prediction ability within the training distribution.
[0092] Table 6 Performance validation of the WLDA method on the validation set within the training distribution.
[0093] This invention also analyzes the key hyperparameters of WLDA. See [link / reference] Figure 3In (a) and (b), the experimental results show that the performance is best when the temperature coefficient λ is 0.15 and does not change much within a reasonable range, indicating that the model is not sensitive to the temperature coefficient λ. The number of learnable reference distributions K is best when it is 128. Too large or too small a number slightly reduces the performance, indicating that an appropriate number of reference distributions can achieve a balance between model flexibility and regularization.
[0094] In summary, the WLDA method demonstrates high prediction accuracy and strong generalization ability on both the standard benchmark set and strictly independent subsets. The GRBA module and the effective complexity control term play a key role in the model's performance and stability, while having no negative impact on the prediction of samples within the training distribution. It provides a reliable tool for predicting the binding affinity of novel protein-small molecule complexes.
[0095] The following are embodiments of the apparatus of the present invention, which can be used to execute embodiments of the method of the present invention. For details not disclosed in the apparatus embodiments, please refer to the embodiments of the method of the present invention.
[0096] See Figure 4 In one embodiment of the present invention, a protein-ligand binding affinity prediction system based on Wasserstein hidden distribution alignment is provided, including an encoding module, a reference distribution establishment module in Wasserstein space, an alignment module and a prediction module; The encoding module is used to encode the protein-ligand complex as a Gaussian distribution in Wasserstein space to obtain the protein-ligand feature distribution. The Wasserstein space reference distribution establishment module is used to establish a reference distribution in Wasserstein space for protein-ligand feature distributions. The alignment module is used to align the protein-ligand feature distribution to a reference distribution in Wasserstein space based on Wasserstein barycentric projection, so as to obtain the aligned protein-ligand complex characterization. The prediction module is used to input the aligned protein-ligand complex characterization into the optimized binding affinity prediction module for prediction, so as to obtain the protein-ligand binding affinity.
[0097] All relevant content of each step involved in the aforementioned embodiments of the protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment can be referred to in the functional description of the corresponding functional module of the protein-ligand binding affinity prediction system based on Wasserstein hidden distribution alignment in the embodiments of the present invention, and will not be repeated here.
[0098] In one embodiment of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment.
[0099] In one embodiment of the present invention, a computer-readable storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the operating system of the terminal. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory. Volatile memory, such as at least one disk storage medium. One or more instructions stored in the computer-readable storage medium can be loaded and executed by a processor to implement the protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment in the above embodiments.
[0100] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0101] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.
[0102] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0103] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for predicting protein-ligand binding affinity based on Wasserstein hidden distribution alignment, characterized in that, Includes the following steps: The protein-ligand complex is encoded as a Gaussian distribution in Wasserstein space, and the protein-ligand feature distribution is obtained. Establish a reference distribution in Wasserstein space for protein-ligand characteristic distributions; The protein-ligand feature distribution is aligned to the reference distribution in Wasserstein space based on Wasserstein barycentric projection, and the aligned protein-ligand complex characterization is obtained. The aligned protein-ligand complex characterization is input into the optimized binding affinity prediction module for prediction, and the protein-ligand binding affinity is obtained.
2. The protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment according to claim 1, characterized in that, Encoding protein-ligand complexes as a Gaussian distribution in Wasserstein space yields the protein-ligand feature distribution, including the following steps: The protein-ligand complex is represented by an interaction diagram; The latent representation set of each node in the complex is obtained by extracting features from the graph structure of the interaction graph through the encoder. The protein-ligand feature distribution is obtained based on the hidden representation set of each node in the complex.
3. The protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment according to claim 1, characterized in that, Protein-ligand characteristic distribution Calculated using the following formula: , in, The mean vector representing the characteristics of the complex; The variance vector representing the feature; Denotes the diagonal covariance matrix. It follows a Gaussian distribution. For the first i The total number of nodes in a complex i This is the complex number. j This represents the node number within the complex. For the first The embedding vector of each node. represents the parameters of a Gaussian distribution.
4. The protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment according to claim 1, characterized in that, The reference distribution in Wasserstein space is established by the following equation: in, For reference distribution, For the reference distribution One portion, The mean, For variance, Parameters for the reference distribution It follows a Gaussian distribution. It is a diagonal covariance matrix. The component index in the reference distribution. This represents the total number of components in the reference distribution.
5. The protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment according to claim 1, characterized in that, Aligning the protein-ligand feature distribution to a reference distribution in Wasserstein space based on Wasserstein barycentric projection, the aligned protein-ligand complex characterization is obtained, including the following steps: Based on the training sample distribution and the reference distribution of the Wasserstein space, the objective function containing the entropy regularization term is obtained. Solving the objective function containing the entropy regularization term yields the optimal transfer matrix; Based on the optimal transfer matrix, the protein-ligand complex feature distribution is mapped to the reference distribution through Wasserstein centroid projection using the Gaussian reference centroid alignment mechanism to obtain the protein-ligand feature distribution parameters. Based on the protein-ligand feature distribution parameters, the protein-ligand node features are mapped to the aligned distribution space to obtain the aligned protein-ligand complex characterization.
6. The protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment according to claim 5, characterized in that, The objective function containing the entropy regularization term is: satisfy: in, This is the complex number. The component index in the reference distribution. For the training sample distribution, For the reference distribution One portion, For the training sample distribution and the reference distribution, the th Transmission weights between components The squared 2-Wasserstein distance Here is the entropy regularization coefficient. B Batch size; The training sample distribution is as follows: in, It follows a Gaussian distribution. The mean vector representing the characteristics of the complex; The variance vector representing the feature; Let represent the diagonal covariance matrix.
7. The protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment according to claim 5, characterized in that, The optimal transfer matrix is: in, For the optimal transfer matrix, This indicates the index of the reference distribution component used in the denominator normalization term. For reference distribution components; The loss function for the affinity prediction module is: in, For loss function, To incorporate the affinity regression loss function, As a reference for distributed complexity control terms, The features of the complex after being encoded by the encoder. This is an interaction diagram of the complex. These are weight hyperparameters; The reference distributed complexity control terms are: in, It is the numerical stability constant. Indicates the first The average transmission quality received in the current batch by a reference distribution.
8. A protein-ligand binding affinity prediction system based on Wasserstein hidden distribution alignment, characterized in that, include: The encoding module is used to encode the protein-ligand complex as a Gaussian distribution in Wasserstein space to obtain the protein-ligand feature distribution. The Wasserstein space reference distribution establishment module is used to establish a reference distribution in Wasserstein space for protein-ligand feature distributions. The alignment module is used to align the protein-ligand feature distribution to a reference distribution in Wasserstein space based on Wasserstein barycentric projection, so as to obtain the aligned protein-ligand complex characterization. The prediction module is used to input the aligned protein-ligand complex characterization into the optimized binding affinity prediction module for prediction, so as to obtain the protein-ligand binding affinity.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the protein-ligand binding affinity prediction method based on Wasserstein hidden distribution alignment as described in any one of claims 1 to 7.