A drug-target affinity prediction method based on federated decision trees
Drug-target affinity prediction is performed among pharmaceutical institutions through the federated decision tree method. Gradient histogram encryption and Diffie-Hellman key exchange are used to protect data privacy, enabling multi-party collaborative learning and improving the performance and efficiency of drug-target affinity prediction.
Patent Information
- Application Number
- CN202411007995.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-07-25
AI Technical Summary
Existing technologies make it difficult to share drug-target binding affinity data and three-dimensional conformation data among multiple pharmaceutical institutions, and lack methods for multi-party collaborative learning while protecting privacy, resulting in inefficient drug-target affinity prediction.
A federated decision tree-based approach is adopted to learn drug and protein features locally, generate gradient histograms and transmit them with homomorphic encryption. The central server aggregates and updates the global model, and each institution updates the local model. The Diffie-Hellman key exchange algorithm is used to protect data privacy.
It has achieved multi-party collaborative learning while protecting data privacy, and improved the performance of drug target affinity prediction. The performance improvement is significant without affecting the prediction ability, ensuring data privacy and security.
Smart Images

Figure CN118824354B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of bioinformatics and computer application technology, and more specifically, to a drug-target affinity prediction method based on a federated decision tree. Background Art
[0002] Drug-target affinity prediction plays an important role in the drug development process. It can help screen candidate compounds, understand interaction mechanisms, guide drug optimization and design, and evaluate drug side effects and risks. These prediction results are of great significance for improving drug development efficiency, reducing development costs, and promoting the discovery and launch of new drugs. However, achieving high-performance scoring functions requires the support of sufficient drug-target binding affinity data and their binding three-dimensional conformational data. Due to commercial confidentiality, legal and regulatory requirements, and privacy protection, drug-target binding affinity data and their binding three-dimensional conformational data are often difficult to share among multiple pharmaceutical institutions. Therefore, developing a method that can protect private data from being leaked while allowing data from multiple pharmaceutical institutions to be used for drug-target affinity prediction is of great significance for drug discovery.
[0003] Federated learning is a distributed machine learning approach that aims to train a shared global model without centrally storing participants' original datasets. Federated learning offloads model training to local devices, while retaining the data locally. Participants train the model locally and transmit only updates to the model parameters back to a central server or the cloud. The central server then updates the global model based on these updates. This approach ensures that individual data remains private and secure.
[0004] Currently, federated learning is rarely used for drug-target binding affinity prediction tasks. In particular, there is currently no multi-party collaborative learning scheme for drug-target affinity prediction based on 3D data that can ensure data privacy during the learning process. Therefore, it is necessary to develop a drug-target affinity prediction method based on federated decision trees to fill the gap in the field of drug-target affinity prediction based on 3D data while preserving data privacy. Summary of the Invention
[0005] The purpose of the present invention is to provide a drug-target affinity prediction method based on a federated decision tree to overcome the defects of the prior art.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A method for drug-target affinity prediction based on a federated decision tree comprises the following steps:
[0008] S1. Each pharmaceutical institution participating in the federated learning performs feature learning of drugs and proteins locally using the method based on multi-shell extended connectivity fingerprints;
[0009] S2. Each pharmaceutical institution calculates the first-order and second-order gradients of the GBDT model based on local data and the initial splitting point, and generates a gradient histogram;
[0010] S3. Each pharmaceutical institution performs a homomorphic encryption operation on the obtained gradient histogram and transmits the encrypted gradient histogram to the central server;
[0011] S4. The central server aggregates the encrypted gradient histograms collected from each pharmaceutical institution to obtain a global gradient histogram, generates the next tree-building operation according to the decision tree construction process of the GBDT model, and forwards the operation information to each pharmaceutical institution;
[0012] S5. Each pharmaceutical institution uses the decision tree construction information sent by the central server to update its local GBDT model and predicts the drug-target affinity based on the local GBDT model. <00企业名称>Furthermore, in step S1, the extended connectivity fingerprint is used to transform the drug-protein interaction into a feature vector representation, and the shell model is used to divide the protein structure into different hierarchical shells.
[0014] Furthermore, in step S2, local learning of the drug-target affinity using the GBDT algorithm is also included to obtain local data.
[0015] Furthermore, step S2 specifically includes:
[0016] For a given loss function l, the GBDT model processes the feature set D = {x i , y i}(0 < i ≤ n), and the calculation formula of the loss function L is:
[0017]
[0018] In the formula, is the predicted value of the GBDT model, f k is the kth decision tree, η is the regularization value of the decision tree, I represents the set of samples, the length of a single feature is a, and the loss function L of the GBDT model calculates the first derivative g and the second derivative h i and the second derivative h i of the tth decision tree according to the predicted value
[0019] In the process of building a decision tree, the GBDT model first calculates the benefit S of the feature set D at each cutting point, and decides whether to fork or generate a leaf node here based on the benefit S. Given the cutting point and all node weights λ, sample I will be divided into left sample I L and right sample I R , with the goal of minimizing the loss function of the GBDT model, the calculation formula of the benefit S is:
[0020]
[0021] Calculate the benefits S of all cutting points according to the above formula, select the cutting point with the largest S to perform a fork operation on the decision tree. If the benefits S are all less than zero or the set maximum tree depth has been reached, stop the fork and make the node a leaf node;
[0022] With the goal of minimizing the loss function of the GBDT model, the optimal leaf output value V is:
[0023]
[0024] For a given B split points Each split point c i Divide sample I into two parts, the left sample is represented as I i , then the gradient histogram is represented as a vector of length B, each value of the vector is the sum of the first-order gradient g and the second-order gradient h of all left samples at the segmentation point. The calculation formula of the gradient histogram H is:
[0025]
[0026] The gradient histogram of the right sample is obtained by subtracting the gradient histogram of the left sample from the gradient histogram of the parent node.
[0027] Furthermore, in step S3, different pharmaceutical institutions use the Diffie-Hellman key exchange algorithm to exchange public keys with each other, and the key is used as noise to encrypt the histogram. The encrypted histogram cannot be deciphered alone, and the summation operation on the central server can obtain the correct final histogram.
[0028] Furthermore, the specific steps of using the Diffie-Hellman key exchange algorithm to exchange public keys with each other and using the keys as noise to encrypt the histogram are as follows:
[0029] For N pharmaceutical institutions, in the initialization phase, pharmaceutical institution P i For all other pharmaceutical institutions j Send key k ij , and received the key k ji , in the encryption stage, the pharmaceutical organization Pi Add noise to all values of the local plaintext gradient histogram i
[0030] noise i =∑ j∈[N] k ji -∑ j∈[N] k ji .
[0031] Compared with the existing technology, the advantages of the present invention are: the present invention proposes a new privacy protection problem for the first time, namely how to conduct multi-party collaborative learning and protect data privacy for the drug target affinity prediction method based on 3D structure. At present, there is no similar research work that applies federated learning to drug target affinity prediction based on 3D data to achieve multi-party collaborative learning and protect data privacy; the present invention designs and proposes a multi-party collaborative learning solution to solve the data privacy problem in drug target affinity prediction based on 3D data. With the help of federated learning technology, the present invention allows multiple data holders to jointly participate in the training of the model while protecting data privacy, and protects the privacy of the data holders by only sharing the update of model parameters; while ensuring data privacy, the present invention can achieve performance comparable to that of plaintext transmission-based centralized learning, without significantly reducing the predictive ability of the model. At the same time, compared with the results of local learning of each pharmaceutical institution, the federated learning method shows a huge performance improvement. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0033] Figure 1 It is a flow chart of the drug-target affinity prediction method based on the federal decision tree of the present invention.
[0034] Figure 2 This is the result of the present invention's prediction of drug-target affinity.
[0035] Figure 3 Provide local and independent learning results for drug-target affinity prediction for each institution. DETAILED DESCRIPTION
[0036] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more precise definition of the protection scope of the present invention.
[0037] Example 1
[0038] See Figure 1 As shown, this embodiment discloses a drug-target affinity prediction method based on a federated decision tree, comprising the following steps:
[0039] Step S1, local feature representation learning: Each pharmaceutical organization participating in the joint learning uses a method based on multi-shell extended connectivity fingerprints to locally learn the features of drugs and proteins; the multi-shell extended connectivity fingerprint drug-target affinity prediction method is a feature learning method based on 3D structural data, which can effectively extract relevant features of drugs and proteins.
[0040] Step S2: Local calculation of GBDT gradient histogram: After local feature learning is completed, each pharmaceutical organization performs first-order and second-order gradient calculations on the GBDT model based on local data and initial segmentation points to generate a gradient histogram.
[0041] Step S3: Model encryption and transmission: Each pharmaceutical organization performs homomorphic encryption on its gradient histogram and transmits the encrypted gradient histogram to the central server. During this process, both the data and its features are hidden, preventing the organization's private information from being leaked.
[0042] Step S4, model aggregation: The central server aggregates the encrypted gradient histograms collected from each pharmaceutical institution to obtain a global gradient histogram, generates the next tree building operation according to the decision tree construction process of the GBDT model, and forwards the operation information to each pharmaceutical institution.
[0043] Step S5, local model update: Each pharmaceutical organization uses the decision tree construction information sent by the central server to update its local GBDT model and perform drug-target affinity prediction based on the local GBDT model.
[0044] Step S6, evaluation: conduct a comparative experiment on this embodiment with multiple advanced three-dimensional structure-based drug target affinity prediction scoring functions; conduct an ablation experiment on this embodiment to evaluate the performance improvement of joint learning compared to local learning of each institution.
[0045] In this embodiment, in step S1, each pharmaceutical organization uses the three-dimensional conformational data of highly sensitive drug target binding to perform local feature representation, thereby avoiding the risk of information leakage. A method combining extended connectivity fingerprint and shell model is used to characterize the interaction and affinity between drugs and proteins. The extended connectivity fingerprint is based on the atomic environment characteristics and converts the drug-protein interaction into a feature vector representation. The shell model divides the protein structure into shells of different levels, thereby capturing information at more scales. After this stage, each pharmaceutical organization obtains feature vectors of the same dimension.
[0046] In the embodiment, in step S2, local learning of the drug target affinity is further performed using the GBDT algorithm to obtain local data.
[0047] The GBDT model in the embodiment is a model that continuously fits residuals. Its core idea is to gradually improve the performance of the overall model by continuously training weak classifiers to fit residuals. Specifically, GBDT uses regression trees as the base learners. In each iteration, a new tree is trained using the residuals of the previous model, and then the new tree is added to the model to improve the prediction performance. The training process of GBDT is a process of forward step-by-step model construction. Each time, the negative gradient of the original objective function is learned, which is equivalent to using a new classifier to fit the residuals. The training of each classifier is based on the residuals weighted by all the previous classifiers, so as to minimize the sum of the squares of the residuals. After the number of iterations reaches the stop condition, the outputs of each classifier are weighted and summed to obtain the final output, which is the output of the overall model.
[0048] For a given loss function l, the GBDT model for the feature set D = {x i , y i}(0 < i ≤ n), the calculation formula of the loss function L is:
[0049]
[0050] In the formula, is the predicted value of the GBDT model, f k is the k-th decision tree, η is the regularization value of the decision tree, I represents the set of samples, the length of a single feature is a, and the loss function L of the GBDT model calculates the first derivative g and the second derivative h i at the t-th decision tree according to the predicted value i of the (t - 1)-th decision tree,
[0051] In the process of constructing the decision tree, the GBDT model first calculates the gain S of the feature set D at each cut point, and decides whether to branch here or generate a leaf node according to the gain S. Given the cut point and the weights λ of all nodes, the sample I will be divided into the left sample I L and the right sample I R . With the goal of minimizing the loss function of the GBDT model, the calculation formula of the gain S is:
[0052]
[0053] Calculate the benefits S of all cutting points according to the above formula, select the cutting point with the largest S to perform a fork operation on the decision tree. If the benefits S are all less than zero or the set maximum tree depth has been reached, stop the fork and make the node a leaf node;
[0054] With the goal of minimizing the loss function of the GBDT model, the optimal leaf output value V is:
[0055]
[0056] It is very time-consuming to continuously calculate the benefits S of all split points when building a decision tree. Usually, the gradient histogram H is used to save the gradient (g, h). The federated tree of the GBDT algorithm based on federated learning calculates the gradient histogram locally and transmits the gradient histogram between multiple clients to complete collaborative learning. Specifically, for a given B split point Each split point c i Divide sample I into two parts, the left sample is represented as I i , then the gradient histogram is represented as a vector of length B, each value of the vector is the sum of the first-order gradient g and the second-order gradient h of all left samples at the segmentation point. The calculation formula of the gradient histogram H is:
[0057]
[0058] The gradient histogram of the right sample is calculated by subtracting the gradient histogram of the left sample from the gradient histogram of the parent node. The histogram is the only statistical information required to update node values. Its dimensionality is very low, equal to the number of cut points. Furthermore, the histogram does not contain the original eigenvalues, protecting the privacy of the original data. Considering effectiveness, efficiency, and privacy, using the gradient histogram as data for transmission between different clients is a very suitable choice.
[0059] In this embodiment, it is assumed in step S3 that the various organizations strictly abide by the agreement, the communication channels are not monitored, and the organizations do not collude with each other. This embodiment uses a homomorphic encryption algorithm to protect the local gradient histograms of different organizations from being leaked to other organizations. For different pharmaceutical organizations, the Diffie-Hellman key exchange algorithm is used to exchange public keys with each other. The key is used as noise to encrypt the histogram, and the encrypted histogram cannot be deciphered alone. The correct final histogram can be obtained by performing a summation operation on the central server. This ensures that each organization or a man-in-the-middle attack cannot obtain the histogram information of other organizations, and at the same time has no effect on the final gradient histogram result. Specifically, for N pharmaceutical organizations, in the initialization stage, the pharmaceutical organization P i For all other pharmaceutical institutions j Send key k ij , and received the key k ji, in the encryption stage, the pharmaceutical organization P i Add noise to all values of the local plaintext gradient histogram i
[0060] noise i =∑ j∈[N] k ij -∑ j∈[N] k ji .
[0061] The central server cannot decipher the gradient histogram of a single mechanism with noise added. However, by summing the gradient histograms of all mechanisms, the noise of all mechanisms will cancel each other out, and the server can still obtain the correct global gradient histogram.
[0062] In this embodiment, the samples of each pharmaceutical institution in step S4 are inconsistent but the feature space is the same. During the aggregation process, since the initial segmentation points of each institution are exactly the same, the gradient histogram represents the sum of the first-order and second-order gradients of the left samples of the same segmentation point, then it is only necessary to simply add the gradient histograms of each institution to obtain the global gradient histogram. After the server obtains the global gradient histogram, it calculates the segmentation benefit S of each segmentation point according to the formula. If the current depth is the maximum depth or all the segmentation benefits S are negative, the bifurcation is terminated and the value of the leaf node is generated according to the formula; otherwise, the segmentation point with the largest benefit is selected for segmentation. After the aggregation is completed, the server sends the results to each pharmaceutical institution.
[0063] In this embodiment, in step S5, each mechanism selects a split point for forking or updates the predicted value of a leaf node based on the information shared by the server. If a forking operation is performed, the local sample is divided into a left sample and a right sample, and the gradient histogram is repeatedly calculated for the left and right nodes, respectively, and then homomorphically encrypted and transmitted to the server. If the current decision tree is updated and the set number of decision trees has not yet been reached, a new decision tree is added and the above decision tree construction operation is repeated. The details of this algorithm are shown in Algorithm 1.
[0064]
[0065] This paper addresses a new privacy protection issue: how to implement multi-party collaborative learning for 3D structure-based drug-target affinity prediction while protecting data privacy. Currently, no similar research has applied federated learning to 3D data-based drug-target affinity prediction to achieve multi-party collaborative learning while protecting data privacy.
[0066] This paper designs and proposes a multi-party collaborative learning solution to address the data privacy issue in drug target affinity prediction based on 3D data. With the help of federated learning technology, this paper allows multiple data holders to jointly participate in model training while protecting data privacy. By only sharing the updates of model parameters, the privacy of data holders is protected.
[0067] While ensuring data privacy, the present invention can achieve performance comparable to that of plaintext transmission-based centralized learning without significantly reducing the model's predictive ability. Furthermore, compared with the results of local learning by each pharmaceutical institution, the joint learning method shows a significant performance improvement.
[0068] Example 2
[0069] In order to better illustrate the performance of the method of the present invention, this example implemented a rigorous procedure to compare the present invention with the multi-shell extended connectivity fingerprint drug target affinity prediction method based on local learning based on plaintext transmission of aggregated drug target datasets from various institutions, as well as five other state-of-the-art three-dimensional structure-based drug target affinity prediction scoring functions, as shown in Table 2.
[0070] Table 2 Comparison of the present invention with the most advanced existing methods (Pearson: Pearson correlation coefficient RMSE: root mean square error)
[0071]
[0072] Experimental results demonstrate that the present invention not only effectively performs joint learning without leaking the original three-dimensional conformational data and data feature vectors, but also demonstrates excellent scoring performance. According to the results in Table 1, the Pearson correlation coefficient of the present invention is 0.872, while the RMSE is 1.178. These results demonstrate that the present invention can efficiently capture drug-target interactions and accurately predict affinity.
[0073] Example 3
[0074] In order to evaluate the performance improvement of federated learning compared to local learning in each institution, independent experiments were conducted on a single pharmaceutical institution and compared with the experimental results of federated learning. Figure 2 Figure 3 The prediction results of drug-target affinity by joint learning and local independent learning of a single pharmaceutical institution are listed separately.
[0075] The comparison shows that federated learning demonstrates significant advantages across all evaluation metrics. It not only significantly improves the prediction model's performance in terms of Pearson and Spearman correlation coefficients, but also significantly reduces the model's RMSE and MAE errors. These results fully demonstrate the significant performance advantages that federated learning can achieve over independent learning within each institution. By fully leveraging information and feature representations from multiple data sources, federated learning can effectively improve the performance of the scoring function.
[0076] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, the patent owner may make various changes or modifications within the scope of the appended claims. As long as they do not exceed the scope of protection described in the claims of the present invention, they should be within the scope of protection of the present invention.
Claims
1. A drug-target affinity prediction method based on a federated decision tree, characterized in that: The following steps are involved: S1. Each pharmaceutical organization participating in the joint learning will use the multi-shell extended connectivity fingerprint method to learn the characteristics of drugs and proteins locally; S2. Each pharmaceutical organization calculates the first-order and second-order gradients of the GBDT model based on local data and initial segmentation points to generate a gradient histogram; S3. Each pharmaceutical institution performs homomorphic encryption on the gradient histogram it obtains and transmits the encrypted gradient histogram to the central server. S4. The central server aggregates the encrypted gradient histograms collected from each pharmaceutical institution to obtain a global gradient histogram, generates the next tree building operation according to the decision tree construction process of the GBDT model, and forwards the operation information to each pharmaceutical institution; S5. Each pharmaceutical organization uses the decision tree construction information sent by the central server to update its local GBDT model and perform drug-target affinity prediction based on the local GBDT model.
2. The drug-target affinity prediction method based on a federated decision tree according to claim 1, characterized in that: In step S1, the extended connectivity fingerprint is used to convert the drug-protein interaction into a feature vector representation, and the shell model is used to divide the protein structure into shells of different levels.
3. The drug-target affinity prediction method based on a federated decision tree according to claim 1, characterized in that: The step S2 also includes using the GBDT algorithm to perform local learning on drug-target affinity to obtain local data.
4. The drug-target affinity prediction method based on a federated decision tree according to claim 1, characterized in that: The step S2 specifically includes: For a given loss function l, the GBDT model for the feature set D = {x i , y i} (0 < i ≤ n), the calculation formula of the loss function L is: Where, is the predicted value of the GBDT model, f k is the kth decision tree, η is the regularization value of the decision tree, I represents the set of samples, the length of a single feature is a, and the loss function L of the GBDT model is based on the predicted value of the t-1th decision tree. Find its first-order derivative g at the tth decision tree i and the second-order derivative h i , In the process of building a decision tree, the GBDT model first calculates the benefit S of the feature set D at each cutting point, and decides whether to fork or generate a leaf node here based on the benefit S. Given the cutting point and all node weights λ, sample I will be divided into left sample I L and right sample I R , with the goal of minimizing the loss function of the GBDT model, the calculation formula of the benefit S is: Calculate the benefits S of all cutting points according to the above formula, select the cutting point with the largest S to perform a fork operation on the decision tree. If the benefits S are all less than zero or the set maximum tree depth has been reached, stop the fork and make the node a leaf node; With the goal of minimizing the loss function of the GBDT model, the optimal leaf output value V is: For a given B split points Each split point c i Divide sample I into two parts, the left sample is represented as I i , then the gradient histogram is represented as a vector of length B, each value of the vector is the sum of the first-order gradient g and the second-order gradient h of all left samples at the segmentation point. The calculation formula of the gradient histogram H is: The gradient histogram of the right sample is obtained by subtracting the gradient histogram of the left sample from the gradient histogram of the parent node.
5. The drug-target affinity prediction method based on a federated decision tree according to claim 1, characterized in that: In step S3, different pharmaceutical institutions use the Diffie-Hellman key exchange algorithm to exchange public keys with each other. The key is used as noise to encrypt the histogram, and the encrypted histogram cannot be deciphered alone. The correct final histogram can be obtained by performing a summation operation on the central server.
6. The method for drug-target affinity prediction based on a federated decision tree according to claim 5, characterized in that: The specific steps of using the Diffie-Hellman key exchange algorithm to exchange public keys with each other and using the keys as noise to encrypt the histogram are as follows: For N pharmaceutical institutions, in the initialization phase, pharmaceutical institution P i For all other pharmaceutical institutions j Send key k ij , and received the key k ji , in the encryption stage, the pharmaceutical organization P i Add noise to all values of the local plaintext gradient histogram i
Citation Information
Patent Citations
Drug target affinity prediction method based on multi-shell and extended connectivity fingerprints
CN116206678A
Protein-ligand affinity prediction method and device based on federated learning
CN117558339A