Multi-molecule relative combination free energy global correction method and system based on free energy state function attributes
Through a global correction method of multi-molecule relative binding free energy based on the properties of free energy state functions, the absolute binding free energy of large-scale molecular networks is directly solved, which solves the problems of high computational complexity and cumbersome weight integration in existing technologies and realizes efficient and accurate drug screening and optimization processes.
Patent Information
- Application Number
- CN202510699197.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-19
AI Technical Summary
The existing relative binding free energy calculation method based on free energy perturbation has high computational complexity and cumbersome weight integration in large-scale molecular networks, making it difficult to adapt to complex network structures, affecting the accuracy and efficiency of drug screening and optimization processes.
A global correction method for the relative binding free energy of multiple molecules based on the properties of the free energy state function is adopted. By constructing a weight vector for the optimization problem, the absolute binding free energy of all molecules is directly solved, and the weights are used to reflect the credibility of the free energy difference of each pair of molecules. It supports adaptive weight iterative optimization, reduces computational complexity, and adapts to large-scale molecular networks.
It significantly improves the accuracy and high-throughput adaptability of free energy predictions such as FEP-RBFE, facilitates integration into automated drug screening and optimization processes, and improves computational efficiency and accuracy.
Smart Images

Figure CN120673908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer-aided drug design, and more specifically, to a method and system for global correction of multi-molecule relative binding free energy based on free energy state function properties. Background Art
[0002] Predicting the binding affinity between drug molecules and targets remains a core challenge in computer-aided drug design. In recent years, relative binding free energy (RBFE) calculations based on free energy perturbation (FEP) have become the most rigorous computational method for achieving high-precision affinity predictions, gaining widespread application and rapid development in the pharmaceutical industry. Thanks to continued advances in high-performance computing hardware, molecular force fields, and sampling algorithms, demand for the large-scale application of FEP-RBFE in molecular screening, lead optimization, and other processes is growing. In particular, deep integration with machine learning frameworks is further accelerating the drug discovery process.
[0003] In a typical FEP-RBFE workflow, researchers construct a perturbation network connecting target molecules and obtain the relative binding free energy between molecular pairs through simulation ( Theoretically, based on the state function properties of free energy, the sum of the free energies of any closed thermodynamic loop in a perturbed network should be strictly zero (i.e., Hess's law). However, in practice, this physical consistency is often violated due to factors such as insufficient sampling and imperfect force fields, resulting in free energy "hysteresis" or inconsistency in closed loops, affecting predictive accuracy. Therefore, developing effective free energy correction methods to enforce or promote loop consistency has been confirmed by multiple studies to be a key approach to improving the reliability of RBFE calculations.
[0004] Currently, mainstream free energy correction methods are mostly based on cycle closure consistency. For example, Schrödinger proposed a loop closure correction algorithm based on Lagrange multipliers; Li et al. proposed the Weighted Cycle Closure (WCC) method, which incorporates the computational uncertainty of each path as a weight; and methods such as MBARnet and CBayesMBAR treat loop consistency as an optimization constraint and directly integrate it into the free energy estimator. While these methods perform well in small or sparse networks, they inherently rely on explicitly identifying and traversing all closed loops in the network to ensure physical consistency. As the number of molecules and network complexity increase, the number of loops required to be enumerated grows exponentially (e.g., O(N!)), making traditional methods computationally unscalable for large-scale FEP-RBFE networks and severely restricting their practical application in high-throughput drug screening.
[0005] In addition, there are statistical errors and uncertainties in the original calculation results of FEP-RBFE. The ideal correction method should not only have polynomial level (such as O(n k ), k should be as small as possible) to ensure that the correction can be completed efficiently in large-scale molecular networks. It should also be able to flexibly integrate the confidence information of the free energy difference of each molecule (such as explicit weights), adapt to the complex or disconnected structure of the network, and be easy to integrate into existing automated virtual screening and molecular optimization processes to meet the needs of industrial-level applications.
[0006] However, current mainstream graph-based RBFE correction methods still face challenges when dealing with large-scale networks, such as high complexity, cumbersome implementation, and limited adaptability to weight integration and network structure. Therefore, a new free energy correction method with a solid mathematical foundation, low computational complexity, flexible weighting mechanism, and amenable to high-throughput integration is urgently needed to comprehensively improve the reliability and applicability of FEP-RBFE in modern drug discovery. Summary of the Invention
[0007] In order to solve the problem of high computational complexity in the prior art, the present invention provides a global correction method and system for multi-molecule relative binding free energy based on the properties of free energy state functions, which has the characteristics of flexible weighting mechanism and easy high-throughput integration.
[0008] In order to achieve the above-mentioned purpose of the present invention, the technical solutions adopted are as follows: A global correction method for multi-molecule relative binding free energy based on free energy state function properties, comprising the following steps: collecting and preprocessing molecular free energy data of drugs; constructing a weight vector for the optimization problem; ; based on and , construct its optimization problem; Solve the optimization problem to obtain the unique solution for all molecules and obtain the absolute free energy after global correction ; Output global correction and the corrected relative free energy difference of any molecular pair .
[0009] Preferably, the molecular free energy data of the drug is collected and preprocessed, and the specific steps are: Each row includes molecular pairs, free energy differences Molecular free energy data; each row of free energy data can also optionally include the corresponding uncertainty ; Establish molecular ID index mapping for free energy data, unify molecular numbers, normalize input data, set initial weights, and obtain relative binding free energy calculation results for multiple molecular pairs .
[0010] Furthermore, the The specific value can be any function, empirical rule, data driven, model predicted, manually specified, or automatically generated, and its specific form can be any of constant, variable parameter, segmented, dynamically updated, adaptively adjusted, and mixed sources.
[0011] Furthermore, we construct the weight vector of the optimization problem When constructing a linear equation system for characterizing the free energy difference relationship between the molecular pairs, the design moment A and the right-hand side term b are also constructed, specifically: Let N be the total number of molecules and P be the number of pairs of molecules for which the relative binding free energy has been calculated; Construct a P×N molecular pair design matrix A, where each row corresponds to a pair (i, j): the i-th column is -1, the j-th column is +1, and the rest are 0; Construct a P-dimensional vector b, .
[0012] Furthermore, based on the relative binding free energy calculation results , construct its optimization problem, specifically: in, is the total number of molecules, and are the absolute binding free energies of molecules i and j, is the calculated result of the relative binding free energy between molecules i and j.
[0013] Furthermore, based on the calculated result of relative binding free energy ΔΔG, the optimization problem is constructed as follows:
[0014] Where A is the molecule-molecule sparse design matrix, for The diagonal matrix, for The measurement vector, For all Column vector of .
[0015] Furthermore, after constructing the optimization problem, if it is determined that there is In any case where the data is unknown or the noise exceeds the preset threshold, adaptive weighted iterative optimization is used to determine , the specific steps are: initialization ; Iteratively repeat steps a to d until convergence: a. Press Current Solution ; b. Calculate the residuals for each pair of molecules ; c. Adoption Or any update of the robust weighting factor , where ε is a value that makes the denominator non-zero; d. Determine the updated Whether the weight convergence is less than the set threshold; Output after convergence .
[0016] Furthermore, the optimization problem is solved to obtain a unique solution for all molecules. The specific steps are: The absolute free energy ΔG of the reference molecule is fixed, and the optimization problem is solved to obtain the unique solution for all molecules. The specific steps are as follows: For this reference molecule Assign a known or assumed absolute binding free energy value, denoted by , build ; Constrain Substitute into the original optimization problem by directly substituting or using a penalty function; Solve the optimization problem and get the globally corrected .
[0017] Furthermore, the optimization problem is solved to obtain a unique solution for all molecules. The specific steps are: For this reference molecule Assign a known or assumed absolute binding free energy value, denoted by , build ; By adding corresponding constraints to the optimization problem, the original underdetermined optimization problem is transformed into a system of equations that is exactly or overdetermined but has a unique least squares solution. The sparse linear algebra method is used to solve the global correction containing all molecules. of vector A global correction system for multi-molecule relative binding free energy based on free energy state function properties, including a data processing module, an initialization module, an optimization solution module, and a result output module; The data processing module is used to collect and pre-process the molecular free energy data of the drug to obtain the relative binding free energy calculation results of multiple molecular pairs. ; The initialization module is used to construct the weight vector of the optimization problem ; The optimization solution module is used based on and , construct its optimization problem; solve the optimization problem to obtain the unique solution for all molecules and obtain the absolute free energy after global correction ; The result output module is used to output the globally corrected and the corrected relative free energy difference of any molecular pair .
[0018] The beneficial effects of the present invention are as follows: This paper proposes a global correction method for multi-molecule relative binding free energies based on the properties of free energy state functions. This method does not rely on loop traversal consisting of molecular perturbation pairs, but directly solves for the absolute binding free energies of all molecules. It uses weights to reflect the credibility of the free energy difference between each pair of molecules, supporting both explicit weighting based on uncertainty and adaptive weighted iterative optimization. This method has low computational complexity and is suitable for large-scale molecular networks. It significantly improves the accuracy and high-throughput adaptability of alchemical free energy predictions such as FEP-RBFE, facilitating integration into automated affinity-guided molecular screening and optimization processes. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a flow chart of a method for global correction of multi-molecule relative binding free energy based on free energy state function properties of the present invention.
[0020] Figure 2 2 is a linear correlation analysis result diagram of the correction value obtained by the present invention and the experimental reference value in Example 2.
[0021] Figure 3This is a schematic diagram of the computational time consumption of the algorithm when the number of molecule pairs is increased and the network complexity is perturbed under the condition that the number of molecules is fixed at 20 in Example 3. DETAILED DESCRIPTION
[0022] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0023] Example 1 like Figure 1 As shown, a global correction method for multi-molecule relative binding free energy based on free energy state function properties includes the following steps: collecting and preprocessing the molecular free energy data of the drug, and obtaining the relative binding free energy calculation results containing multiple molecular pairs. ;based on , construct the weight vector of the optimization problem ; based on and , construct its optimization problem; Solve the optimization problem to obtain the unique solution for all molecules and obtain the absolute free energy after global correction ; Output global correction and the corrected relative free energy difference of any molecular pair .
[0024] In a specific embodiment, the molecular free energy data of the drug is collected and preprocessed, and the specific steps are as follows: Each row includes molecular pairs, free energy differences In this example, each row contains molecular free energy data of ligand i , ligand j ), the relative free energy difference measured by FEP-RBFE , and optional uncertainty .
[0025] Establish molecular ID index mapping for free energy data, unify molecular numbers, normalize input data, set initial weights, and obtain relative binding free energy calculation results for multiple molecular pairs .
[0026] In a specific embodiment, the weight vector of the optimization problem is constructed , specifically: If each row of free energy data includes the corresponding uncertainty , then based on After normalization, set If not included, take .
[0027] In one embodiment, based on the relative binding free energy calculation results , construct its optimization problem, specifically: in, is the total number of molecules, and are the absolute binding free energies of molecules i and j, is the calculated result of the relative binding free energy between molecules i and j.
[0028] In a specific embodiment, after constructing the optimization problem, if it is determined that there is In any case where the data is unknown or the noise exceeds the preset threshold, adaptive weighted iterative optimization is used to determine , the specific steps are: initialization ; Iteratively repeat steps a to d until convergence: a. Press Current Solution ; b. Calculate the residuals for each pair of molecules ; c. Update ,in To make the denominator non-zero, take 10 -8 ; d. Determine the updated Whether the weight convergence is less than the set threshold; Output after convergence .
[0029] In a specific embodiment, solving the optimization problem to obtain a unique solution for all molecules involves the following steps: The absolute free energy of the reference molecule Fix, solve the optimization problem to obtain the unique solution for all molecules, the specific steps are: For this reference molecule Assign a known or assumed absolute binding free energy value, denoted by , build ; Constrain Directly substitute into the original optimization problem and change the optimization variable from indivual( ) is reduced to In each term of the sum middle: if , then the item becomes ; if , then the item becomes
[0030] Then to this The variables are minimized to obtain the globally corrected .
[0031] Example 2 In this embodiment, the The specific acquisition rules are based on any of mathematics, statistics, physics, chemistry, data-driven or combination rules, and the acquisition methods include but are not limited to residuals, confidence intervals, distribution estimates, robust loss functions, experience additions, external model predictions, prior knowledge, Bayesian probability inference, machine learning inference, manual settings, and multi-source fusion. The specific forms are any of constants, variable parameters, segmentation, dynamic updates, adaptive adjustments, and mixed sources.
[0032] In this embodiment, in a specific embodiment, based on the relative binding free energy calculation results , construct its optimization problem, specifically: in, is the total number of molecules, and are the absolute binding free energies of molecules i and j, is the calculated result of the relative binding free energy between molecules i and j.
[0033] In this embodiment, the optimization problem is solved to obtain a unique solution for all molecules. The specific steps are: The absolute free energy ΔG of the reference molecule is fixed, and the optimization problem is solved to obtain the unique solution for all molecules. The specific steps are as follows: For this reference molecule Assign a known or assumed absolute binding free energy value, denoted by , build ; Add the constraint as a penalty term to the objective function: in is a very large positive constant (penalty weight). As it approaches infinity, the optimization process forces tends to 0, that is , this method keeps the number of optimization variables to , and obtain the globally corrected .
[0034] Figure 2The results show that after the SFC and WSFC methods proposed in this paper are used to calibrate a molecular perturbation network containing 30 molecules and 69 molecular pairs, the correction values ( and ) and experimental reference value ( ) is the linear correlation analysis result diagram. By comparing the or original calculated value) and after correction ( and ) The fit between the data points and the experimental values can be used to evaluate the accuracy improvement of the correction method. Figure 2 The results show that the SFC and especially the WSFC methods can improve the consistency between the calculated and experimental values, as shown in Figure (C). and The R² value reaches 0.9520, which is better than the 0.7148 of the original sample in (A) and the 0.8854 of the SFC in (B). This shows that the method of this invention can effectively improve the accuracy of free energy predictions such as FEP-RBFE.
[0035] Example 3 In this embodiment, based on the relative binding free energy calculation results , construct its optimization problem, specifically:
[0036] Where A is the molecule-molecule sparse design matrix, for The diagonal matrix, for The measurement vector, For all Column vector of .
[0037] In a specific embodiment, after constructing the optimization problem, if it is determined that there is In any case where the data is unknown or the noise exceeds the preset threshold, adaptive weighted iterative optimization is used to determine , the specific steps are: initialization ; Iteratively repeat steps a to d until convergence: a. Press Current Solution ; b. Calculate the residuals for each pair of molecules ; c. Using Huber robust weighting factor, update through robust loss function ; d. Determine the updated Weight convergence of Is it less than the set threshold? ; Output after convergence .
[0038] In a specific embodiment, solving the optimization problem to obtain a unique solution for all molecules involves the following steps: For this reference molecule Assign a known or assumed absolute binding free energy value, denoted by , build ; By adding corresponding constraints to the optimization problem, the original underdetermined optimization problem is transformed into a system of equations that is exactly or overdetermined but has a unique least squares solution: In the original Design Matrix Add a line below to become , in this row, only Column (corresponding to reference molecule ) is 1, and the rest are 0; In the original The target vector Add an element below to become , the value of this element is ; Weight Matrix Expand accordingly and give it greater weight to enforce its exact satisfaction.
[0039] The sparse linear algebra method is used to solve the global correction containing all molecules. of vector; In this embodiment, the error of the molecule pair is also output; for each molecule , traverse all the original input molecular pairs and the corresponding relative binding free energy difference calculated by FEP-RBFE . Collect all the conditions that meet or These are molecules that are directly For each "edge" connected to the molecule Related molecular pairs (Right now or ): Get the relative binding free energy of the original input ; The absolute binding free energy obtained after correction was used and , calculate the corrected relative binding free energy of the molecule pair: ; Calculate the residual for this numerator pair, which is the difference between the original input value and the corrected back-calculated value: . Let the pairwise error of the molecule pair be the absolute value of the above residual: ; Collect all molecules Correlated pairwise error; output numerator Path dependence error , defined as the maximum of these related pairwise errors, .
[0040] like Figure 3 As shown, with a fixed number of molecules at 20, the computational time of the SFC method (blue curve) is much lower than that of the WCC method (orange curve). Furthermore, as the number of molecular pairs or network complexity increases, the computational time of the WCC method increases sharply, while the computational time of the SFC method increases much more gradually. This demonstrates that the SFC method of the present invention has low computational complexity and is suitable for processing large-scale molecular networks, demonstrating its advantages in efficiency and scalability, which are crucial for high-throughput drug screening applications.
[0041] Example 4 A global correction system for multi-molecule relative binding free energy based on free energy state function properties, including a data processing module, an initialization module, an optimization solution module, and a result output module; The data processing module is used to collect and pre-process the molecular free energy data of the drug to obtain the relative binding free energy calculation results of multiple molecular pairs. ; The initialization module is used to construct the weight vector of the optimization problem ; The optimization solution module is used to calculate the results based on the relative binding free energy , construct its optimization problem; solve the optimization problem to obtain the unique solution for all molecules and obtain the absolute free energy after global correction ; The result output module is used to output the globally corrected and the corrected relative free energy difference of any molecular pair .
[0042] Obviously, the above embodiments of the present invention are merely examples for the purpose of illustrating the present invention, and are not intended to limit the embodiments of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.
Claims
1. A global correction method for multi-molecule relative binding free energy based on free energy state function properties, characterized in that: The following steps are involved: Collect and preprocess the molecular free energy data of drugs to obtain the relative binding free energy calculation results of multiple molecular pairs ; Construct the weight vector of the optimization problem ; based on and , construct its optimization problem; Solve the optimization problem to obtain the unique solution for all molecules and obtain the absolute free energy after global correction ; Output global correction and the corrected relative free energy difference of any molecular pair .
2. The method for global correction of multi-molecule relative binding free energy based on free energy state function properties according to claim 1, characterized in that: Collect and preprocess the molecular free energy data of the drug. The specific steps are as follows: Each row includes molecular pairs, free energy differences Molecular free energy data; each row of free energy data can also optionally include the corresponding uncertainty ; Establish molecular ID index mapping for free energy data, unify molecular numbers, normalize input data, set initial weights, and obtain relative binding free energy calculation results for multiple molecular pairs .
3. The method for global correction of multi-molecule relative binding free energy based on free energy state function properties according to claim 2, characterized in that: The The specific value can be any function, empirical rule, data driven, model predicted, manually specified, or automatically generated, and its specific form can be any of constant, variable parameter, segmented, dynamically updated, adaptively adjusted, and mixed sources.
4. The method for global correction of multi-molecule relative binding free energy based on free energy state function properties according to claim 2, characterized in that: Construct the weight vector for the optimization problem When constructing a linear equation system for characterizing the free energy difference relationship between the molecular pairs, a design matrix A and a right-hand term b are also constructed, specifically: Let N be the total number of molecules and P be the number of pairs of molecules for which the relative binding free energy has been calculated; Construct a P×N molecular pair design matrix A, where each row corresponds to a pair (i, j): the i-th column is -1, the j-th column is +1, and the rest are 0; Construct a P-dimensional vector b, .
5. The method for global correction of multi-molecule relative binding free energy based on free energy state function properties according to claim 3, characterized in that: Based on the results of relative binding free energy calculation , construct its optimization problem, specifically: in, is the total number of molecules, and are the absolute binding free energies of molecules i and j, is the calculated result of the relative binding free energy between molecules i and j.
6. The method for global correction of multi-molecule relative binding free energy based on free energy state function properties according to claim 4, characterized in that: Based on the results of relative binding free energy calculation , construct its optimization problem, specifically: Where A is the molecule-molecule sparse design matrix, for The diagonal matrix, for The measurement vector, For all Column vector of .
7. The method for global correction of relative binding free energy of multiple molecules based on free energy state function properties according to any one of claims 5 and 6, characterized in that: After constructing the optimization problem, if it is judged that there is In any case where the data is unknown or the noise exceeds the preset threshold, adaptive weighted iterative optimization is used to determine , the specific steps are: initialization ; Iteratively repeat steps a to d until convergence: a. Press Current Solution ; b. Calculate the residuals for each pair of molecules ; c. Adoption Or any update of the robust weighting factor ,in is a value that makes the denominator non-zero; d. Determine the updated Whether the weight convergence is less than the set threshold; Output after convergence .
8. The method for global correction of multi-molecule relative binding free energy based on free energy state function properties according to any one of claim 5, characterized in that: Solve the optimization problem to obtain the unique solution for all molecules. The specific steps are: The absolute free energy of the reference molecule Fix, solve the optimization problem to obtain the unique solution for all molecules, the specific steps are: For this reference molecule Assign a known or assumed absolute binding free energy value, denoted by , build ; Constrain Substitute into the original optimization problem by directly substituting or using a penalty function; After global correction .
9. The method for global correction of multi-molecule relative binding free energy based on free energy state function properties according to any one of claim 6, characterized in that: For this reference molecule Assign a known or assumed absolute binding free energy value, denoted by , build ; By adding corresponding constraints to the optimization problem, the original underdetermined optimization problem is transformed into a system of equations that is exactly or overdetermined but has a unique least squares solution. The sparse linear algebra method is used to solve the global correction containing all molecules. of vector.
10. A global correction system for relative binding free energy of multiple molecules based on free energy state function properties, characterized in that: Including data processing module, initialization module, optimization solution module, and result output module; The data processing module is used to collect and pre-process the molecular free energy data of the drug to obtain the relative binding free energy calculation results of multiple molecular pairs. ; The initialization module is used to construct the weight vector of the optimization problem ; The optimization solution module is used based on and , construct its optimization problem; Solve the optimization problem to obtain the unique solution for all molecules and obtain the absolute free energy after global correction ; The result output module is used to output the globally corrected and the corrected relative free energy difference of any molecular pair .