A method, apparatus, medium, and procedure for evaluating the affinity of protein-protein complexes and screening protein binders.

CN122575462APending Publication Date: 2026-08-14SHANGHAI MOLECULAR HEART INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-09
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,界面预测模板建模分数(ipTM,Interface Predicted Template Modeling score)、预测局部距离差异测试(pLDDT,Predicted Local Distance Difference Test)等结构置信度指标并不能直接反映结合亲和力,无法区分真实结合物与模型生成的结构性陷阱(decoys),其预测的高置信度分数的结合结构分仅表明模型对预测复合物结构具有信心,但并不能保证预测界面具备溶液中稳定结合所需的理化特性

Benefits of technology

[0022]与现有技术相比,本申请基于待评估的蛋白质-蛋白质复合物,利用结构预测模型生成对应的多个复合物预测结构构象;对于每个复合物预测结构构象,提取所述复合物预测结构构象对应的结合界面信息;基于每个复合物预测结构构象对应的结合界面信息,通过亲和力评估模型,确定所述待评估的蛋白质-蛋白质复合物对应的亲和力评分,其中,所述亲和力评估模型包括统计物理力场、图神经网络与混合密度网络,所述亲和力评估模型在蛋白质-小分子复合物数据集上训练获得。通过对复合物对应的多个结构构象的亲和力进行评分,综合考虑蛋白质界面天然的灵活性与构象多样性,避免依赖单一静态结构,提升评分准确性。并且,本申请所采用的亲和力评估模型完全在蛋白质-小分子复合物数据集上训练,避免数据泄露。该亲和力评估模型引入统计物理力场来直接评估界面互补性而非模型自信度。并通过图神经网络与混合密度网络来预测用于统计物理力场计算的相关参数,使计算评分对AI生成结构的几何噪声具有鲁棒性,提升评分准确性。进一步地,还可以利用该亲和力评分为目标蛋白质筛选合适的蛋白质结合子。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575462A_ABST
    Figure CN122575462A_ABST
Patent Text Reader

Abstract

The purpose of this application is to provide a method, apparatus, medium, and program product for evaluating the affinity of protein-protein complexes and screening protein binders. The method includes: generating multiple corresponding structural conformations based on the protein-protein complex to be evaluated using a structure prediction model; extracting corresponding binding interface information for each structural conformation; and determining the affinity score of the protein-protein complex to be evaluated based on the binding interface information using an affinity evaluation model. The model includes a statistical physical force field, a graph neural network, and a hybrid density network, and is trained on a protein-small molecule complex dataset. This avoids reliance on a single static structure, and the use of a model trained on a protein-small molecule complex dataset avoids data leakage. The graph neural network and hybrid density network are used to predict relevant parameters for force field calculations, improving the accuracy of the scoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics technology, and more particularly to a technique for evaluating the affinity of protein-protein complexes. Background Technology

[0002] In the fields of protein engineering and drug discovery, generative deep learning models can efficiently generate a large number of candidate protein binders. Existing standard computational workflows typically use structure prediction models to predict the binding structure between the designed protein and the target protein. These models then generate structure confidence scores to assess the accuracy of the predicted binding structures. However, structure confidence metrics such as the Interface Predicted Template Modeling score (ipTM) and the Predicted Local Distance Difference Test (pLDDT) do not directly reflect binding affinity and cannot distinguish between real binders and model-generated structural decoys. High confidence scores for these models only indicate that the model is confident in predicting the complex structure, but they do not guarantee that the predicted interface possesses the physicochemical properties required for stable binding in solution.

[0003] Existing affinity scoring methods, such as scoring functions based on classical physics (e.g., FoldX, HADDOCK) and interface evaluation algorithms (e.g., Rosetta interface energy (ΔG), PRODIGGGY), are designed for natural or near-natural protein complexes. They are extremely sensitive to minute atomic overlaps (clashes), geometric distortions, and other "distribution shifts" in AI-generated structures, leading to divergent energy assessment values ​​and making accurate predictions difficult. Deep learning-based affinity prediction models, on the other hand, are trained directly on protein-protein complex datasets, making them prone to data leakage. These models learn specific sequence or structural patterns rather than true physical interactions. Summary of the Invention

[0004] One object of this application is to provide a method, apparatus, medium, and procedure for evaluating the affinity of protein-protein complexes and screening protein binders.

[0005] According to one aspect of this application, a method for assessing the affinity of protein-protein complexes is provided, the method comprising:

[0006] Based on the protein-protein complex to be evaluated, multiple corresponding complex predicted structural conformations are generated using a structure prediction model.

[0007] For each predicted structural conformation of a complex, the binding interface information corresponding to the predicted structural conformation of the complex is extracted.

[0008] Based on the binding interface information corresponding to the predicted structural conformation of each complex, the affinity score of the protein-protein complex to be evaluated is determined by an affinity assessment model. The affinity assessment model includes a statistical physical force field, a graph neural network, and a hybrid density network, and is trained on a protein-small molecule complex dataset.

[0009] According to another aspect of this application, a method for screening protein binders is provided, the method comprising:

[0010] For candidate protein-protein complexes formed by the target protein and each candidate protein binder, the affinity score corresponding to each candidate protein-protein complex is determined, wherein the affinity score corresponding to each candidate protein-protein complex is determined by the method described above for protein-protein complex affinity assessment.

[0011] Based on the affinity scores of each candidate protein-protein complex, one or more corresponding protein binders are selected from each candidate protein binder.

[0012] According to one aspect of this application, a computer device is provided, including a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of any of the methods described above.

[0013] According to one aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described above.

[0014] According to one aspect of this application, a computer program product is provided, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of any of the methods described above.

[0015] According to one aspect of this application, an apparatus for evaluating the affinity of protein-protein complexes is provided, the apparatus comprising:

[0016] The module is used to generate multiple predicted structural conformations of the corresponding complexes based on the protein-protein complex to be evaluated using a structural prediction model.

[0017] The first and second modules are used to extract the binding interface information corresponding to the predicted structural conformation of each complex for each complex.

[0018] The first and third modules are used to determine the affinity score of the protein-protein complex to be evaluated based on the binding interface information corresponding to the predicted structural conformation of each complex and through an affinity assessment model. The affinity assessment model includes a statistical physical force field, a graph neural network, and a hybrid density network, and is trained on a protein-small molecule complex dataset.

[0019] According to another aspect of this application, an apparatus for screening protein binders is provided, the apparatus comprising:

[0020] Module 21 is used to determine the affinity score of each candidate protein-protein complex formed by the target protein and each candidate protein binder, wherein the affinity score of each candidate protein-protein complex is determined by the method described above for protein-protein complex affinity assessment.

[0021] The second module is used to screen one or more protein binders from each candidate protein-protein complex based on the affinity score corresponding to each candidate protein-protein complex.

[0022] Compared with existing technologies, this application generates multiple predicted structural conformations of the protein-protein complex to be evaluated using a structure prediction model. For each predicted structural conformation, the binding interface information corresponding to that conformation is extracted. Based on the binding interface information corresponding to each predicted structural conformation, an affinity score is determined for the protein-protein complex to be evaluated using an affinity assessment model. This affinity assessment model includes a statistical physical force field, a graph neural network, and a hybrid density network, and is trained on a protein-small molecule complex dataset. By scoring the affinity of multiple structural conformations corresponding to the complex, the inherent flexibility and conformational diversity of the protein interface are comprehensively considered, avoiding reliance on a single static structure and improving scoring accuracy. Furthermore, the affinity assessment model used in this application is trained entirely on a protein-small molecule complex dataset, avoiding data leakage. This affinity assessment model introduces a statistical physical force field to directly evaluate interface complementarity rather than model confidence. A graph neural network and a hybrid density network are used to predict relevant parameters used for calculating the statistical physical force field, making the calculated score robust to the geometric noise of the AI-generated structure and improving scoring accuracy. Furthermore, this affinity score can be used to screen suitable protein binders for target proteins. Attached Figure Description

[0023] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0024] Figure 1 This diagram illustrates a method for assessing the affinity of protein-protein complexes according to one embodiment of this application.

[0025] Figure 2 A flowchart of a method for screening protein binders according to an embodiment of this application is shown;

[0026] Figure 3 This diagram illustrates a structural diagram of an apparatus for evaluating the affinity of protein-protein complexes according to one embodiment of this application.

[0027] Figure 4 This diagram illustrates a structural diagram of an apparatus for screening protein binders according to an embodiment of this application.

[0028] Figure 5 Exemplary systems that can be used to implement the various embodiments described in this application are shown.

[0029] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation

[0030] The present application will now be described in further detail with reference to the accompanying drawings.

[0031] In a typical configuration of this application, the terminal, the device of the service network, and the trusted party all include one or more processors (e.g., a central processing unit (CPU)), input / output interfaces, network interfaces, and memory.

[0032] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. Memory is an example of computer-readable media.

[0033] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0034] The devices referred to in this application include, but are not limited to, user equipment, network equipment, or devices composed of user equipment and network equipment integrated through a network. The user equipment includes, but is not limited to, any mobile electronic product capable of human-computer interaction (e.g., via a touchpad), such as smartphones and tablets. These mobile electronic products can use any operating system, such as Android or iOS. The network equipment includes an electronic device capable of automatically performing numerical calculations and information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and embedded devices. The network equipment includes, but is not limited to, computers, network hosts, single network servers, multiple network server clusters, or clouds composed of multiple servers. Here, a cloud consists of a large number of computers or network servers based on cloud computing, where cloud computing is a type of distributed computing, consisting of a virtual supercomputer composed of a group of loosely coupled computer clusters. The network includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, VPN network, wireless ad hoc network, etc. Preferably, the device can also be a program running on the user equipment, network device, or a device formed by integrating user equipment and network device, network device, touch terminal, or network device and touch terminal through a network.

[0035] Of course, those skilled in the art should understand that the above-described devices are merely examples, and other existing or future devices that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0036] In the description of this application, "multiple" means two or more, unless otherwise expressly and specifically defined.

[0037] Figure 1A flowchart illustrating a method for evaluating the affinity of protein-protein complexes according to an embodiment of this application is shown. The method includes steps S11, S12, and S13. In step S11, device 1 generates multiple predicted structural conformations of the protein-protein complex to be evaluated using a structural prediction model. In step S12, device 1 extracts the binding interface information corresponding to each predicted structural conformation. In step S13, device 1 determines the affinity score of the protein-protein complex to be evaluated based on the binding interface information corresponding to each predicted structural conformation using an affinity evaluation model. The affinity evaluation model includes a statistical physical force field, a graph neural network, and a hybrid density network, and is trained on a protein-small molecule complex dataset.

[0038] In some embodiments, the device 1 includes, but is not limited to, user equipment or network equipment with information processing or computing capabilities, such as tablet computers, computers, and servers.

[0039] In step S11, device 1 generates multiple predicted structural conformations of the corresponding complex based on the protein-protein complex to be evaluated using a structural prediction model.

[0040] In some embodiments, the protein-protein complex can be any known protein-protein complex; it can also be a protein-protein complex designed based on the target protein using a corresponding generative model (e.g., ProtGPT2, BindCraft, Boltzgen, etc.). The sequence information of the protein-protein complex can be input into the structure prediction model to obtain multiple conformations of the protein-protein complex structure predicted by the structure prediction model. The structure prediction model can be a model that supports multi-conformation sampling, which can provide multiple structural states for the same complex design. For example, models such as AlphaFold3, AlphaFold2-Multimer, Chai-1, or RoseTTAFold All-Atom can be selected. For the aforementioned structure prediction model, one or more random seeds can be set according to actual needs to generate the required number of predicted complex structural conformations. Here, generating diverse structural conformations of the protein-protein complex is used for subsequent affinity assessment model scoring, aiming to fully consider the uncertainty of the protein-protein complex structural state, capture the epistemic uncertainty of model prediction, and improve the accuracy of complex affinity assessment.

[0041] Those skilled in the art should understand that the above-mentioned generative models and structural prediction models are merely examples, and other existing or future models that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0042] In step S12, device 1 extracts the binding interface information corresponding to each predicted structural conformation of the complex. For example, for each predicted structural conformation of the complex, a region in the protein-protein complex where multiple chains that need to have their affinity calculated interact can be identified and extracted. Based on the information of residues in the region, the corresponding binding interface information is determined. Specifically, step S12 includes: for each predicted structural conformation of the complex, determining multiple interface residues from the predicted structural conformation of the complex, wherein the interface residues contain at least one heavy atom whose distance from the opposite side chain of the chain to which the interface residue belongs satisfies a preset distance condition; and determining the binding interface information corresponding to the predicted structural conformation of the complex based on the multiple interface residues. For example, for each pair of interacting chains in the predicted structural conformation of the complex, its corresponding binding interface can be defined by the interface residues in which the heavy atom (e.g., carbon, oxygen, nitrogen, etc., non-hydrogen atoms) is distanced from the opposite side chain to satisfy a preset distance condition. The corresponding binding interface information can then be determined based on the determined interface residues. The combination of interacting chains can be determined based on the actual affinity calculation requirements. For example, if a protein-protein complex consists of only two chains, A and B, then the combination of the interacting chains is chain A and chain B. If a protein-protein complex consists of two or more chains, the user specifies which chains' affinities are calculated, resulting in various combinations. For example, if a protein-protein complex consists of chains A, B, C, and D, chain A and chain B can be specified as the combination of the interacting chains, or chain A and chain B+D can be specified as the combination of the interacting chains, or chain A+C and chain D can be specified as the combination of the interacting chains; the possible combinations of the interacting chains are not limited here. Correspondingly, the opposing side chain is one or more chains corresponding to the combination of the interacting chains and the chain to which the interface residues belong. For example, for the combination of chain A and chain B, chains A and B are opposing side chains. For the combination of chain A and chain B+D, chains B and D are considered as a whole, the opposing side chain of chain A is the whole of chains B and D, and the opposing side chain of chain B / D is chain A. In some embodiments, the preset distance condition includes the distance between the heavy atom and the opposing side chain being less than or equal to a preset distance threshold.

[0043] In step S13, device 1 determines the affinity score of the protein-protein complex to be evaluated based on the binding interface information corresponding to the predicted structural conformation of each complex using an affinity assessment model. The affinity assessment model includes a statistical physical force field, a graph neural network, and a hybrid density network, and is trained on a protein-small molecule complex dataset.

[0044] In some embodiments, the affinity assessment model employs a zero-shot transfer strategy, meaning it is trained solely on a protein-small molecule complex dataset and directly applied to protein-protein complex affinity assessment tasks. This fundamentally avoids data leakage. The statistical physical force field is a molecular force field constructed based on statistical physics principles, used to describe intermolecular interactions and calculate system potential energy. Key physical parameters are learned through graph neural networks and hybrid density networks. For each predicted complex conformation, parameters for potential energy calculation can be predicted using graph neural networks and hybrid density networks based on the corresponding binding interface information. These parameters can then be used to calculate the affinity score of the predicted complex conformation using the statistical physical force field. The arithmetic mean of the affinity scores for all predicted complex conformations is used to determine the final affinity score for the protein-protein complex to be evaluated. Here, the physical parameters used for force field calculations are predicted and obtained through graph neural networks and hybrid density networks. They can be dynamically adjusted according to the local environment, ensuring that even in AI-predicted structures containing tiny atomic collisions or geometric deviations, the force field can gracefully filter noise and capture accurate physical signals, thus improving the accuracy of affinity assessment.

[0045] In some embodiments, step S13 includes: step S131 (not shown), for each complex predicted structural conformation corresponding to the binding interface information, device 1 determines multiple corresponding molecular maps by taking one protein in the protein-protein complex to be evaluated as the ligand and pocket respectively, and correspondingly taking another protein as the pocket and ligand respectively; step S132 (not shown), for each molecular map, device 1 determines the affinity score corresponding to the molecular map through an affinity assessment model, wherein the affinity assessment model includes a statistical physical force field, a graph neural network, and a mixed density network, and the affinity assessment model is trained on a protein-small molecule complex dataset; step S133 (not shown), device 1 determines the affinity score corresponding to the complex predicted structural conformation based on the affinity score corresponding to each molecular map; step S134 (not shown), device 1 determines the affinity score corresponding to the protein-protein complex to be evaluated based on the affinity score corresponding to the complex predicted structural conformation.

[0046] In some embodiments, since the binding interface of a protein-protein complex does not have a clear distinction between the binding pocket and the ligand compared to the binding interface of a protein-small molecule complex, calculating a score in only one direction would introduce human bias. To eliminate input perspective bias and compensate for the geometric asymmetry of the protein binding interface, one protein (denoted as A) in the protein-protein complex to be evaluated can be treated as the ligand, and the other protein (denoted as B) as the pocket, thus determining the corresponding A→B molecular map; then, the roles are reversed, with protein B as the ligand and protein A as the pocket, to determine the corresponding B→A molecular map. Based on these two molecular maps, the corresponding affinity scores are determined using an affinity assessment model. The average of the affinity scores from these two molecular maps is then used as the affinity score corresponding to the predicted structural conformation of the complex, making the final affinity score closer to the true affinity score, which is independent of direction.

[0047] In some embodiments, an atom of an interface residue in the interface information is used as a node, and atomic pairs whose distance to the corresponding atom of a ligand atom and the atom of a pocket atom is less than or equal to a preset threshold are used as edges to construct the corresponding molecular graph. The node information in the molecular graph includes, but is not limited to, atom type, chemical bond information, hydrophobicity, charge, accessible surface area, and ligand / pocket identifier; the edge information in the molecular graph includes, but is not limited to, the distance between the corresponding atom pairs. In some embodiments, the node information can be calculated and obtained using cheminformatics tools such as RDKit.

[0048] In some embodiments, the molecular graph is input into an affinity assessment model, and parameters for calculating the statistical physical force field are determined sequentially through a graph neural network and a hybrid density network. Based on these parameters and the statistical physical force field, a corresponding affinity score is calculated. Specifically, step S132 includes: step S1321, for each molecular graph, device 1 determines the corresponding atom pair feature information through the graph neural network; step S1322, based on the atom pair feature information, device 1 predicts the corresponding physical parameter information through the hybrid density network; step S1323, based on the physical parameter information, device 1 determines the affinity score corresponding to the molecular graph using the statistical physical force field.

[0049] In some embodiments, the graph neural network includes a gated graph attention layer (GatedGAT) and an interaction network layer. In some embodiments, the molecular graph is input into the graph neural network, and the gated graph attention layer integrates the information of each node with the information of its neighboring nodes. For example, the attention score of a neighboring node can be determined based on the node information, the neighboring node information, and the atomic pair distance between the node and its neighboring nodes. Based on the attention scores of each neighboring node, the node information of each neighboring node is weighted and summed to determine aggregated neighbor information. Combined with this aggregated neighbor information, the node information is updated through a gating mechanism. In some embodiments, multiple layers of the gated graph attention layer are typically stacked to iterate the node information multiple times, thereby capturing more information from neighboring nodes. Then, based on the node information embedded with neighboring node information, the interaction network layer performs ligand-side and pocket-side information exchange. For example, the ligand node information set and the pocket node information set can be determined based on the ligand or pocket identifier in the node information. Based on the ligand node information set and the pocket node information set, the corresponding ligand global information and pocket global information are determined respectively. For example, the sum or mean of all node information in the ligand / pocket node information set can be calculated as the corresponding ligand / pocket global information. Then, node information belonging to the ligand node information set is concatenated with the global pocket information to update and obtain new node information; similarly, node information belonging to the pocket node information set is concatenated with the global ligand information to update and obtain new node information. This allows node information to capture the overall features of the opposite side. Furthermore, for each atom pair corresponding to an edge, the atom pair feature information can be determined based on the finally updated node information corresponding to the two atoms in the atom pair and the corresponding edge information.

[0050] In some embodiments, the Mixture Density Network (MDN) is a neural network architecture specifically designed to handle regression problems with inherent uncertainty and multi-valued mappings. Instead of directly predicting a fixed parameter value, it predicts the probability distribution the parameters follow, thus capturing the uncertainty of atom-pair interactions. In some embodiments, inputting the atom-pair feature information into the MDN yields the parameters μ and σ of the Gaussian components corresponding to the atom-pair feature information, output by the MDN. In some embodiments, the physical parameter information includes van der Waals force energy parameters and / or hydrogen bond energy parameters. These parameters are used for subsequent calculations of the corresponding energies in the statistical physical force field. For various types of physical parameter information, a matching MDN can be used for prediction. For example, for van der Waals force energy parameters, since van der Waals forces involve multiple interaction modes, a multi-head MDN is needed to learn multiple Gaussian components to more accurately fit the true statistical distribution. A three-head MDN is typically used. Each hybrid density network head can output corresponding Gaussian component parameters (μ and σ) and corresponding weights α, and the outputs of each hybrid density network head are used as van der Waals force parameters. For hydrogen bond energy parameters, since hydrogen bond interactions are subject to strict geometric constraints, using multiple components can easily lead to overfitting. A single-head hybrid density network can be used to obtain simpler and more stable prediction results, and the Gaussian component parameters (μ and σ) corresponding to the feature information of each atom pair output by the single-head hybrid density network are used as hydrogen bond energy parameters.

[0051] After obtaining the physical parameter information, this information can be substituted into a statistical physical force field for calculation to obtain the corresponding affinity score. Specifically, step S1323 includes determining the corresponding physical energy information based on the physical parameter information and using the statistical physical force field; and determining the affinity score corresponding to the molecular diagram based on the physical energy information. In some embodiments, the affinity score corresponding to the molecular diagram in the statistical physical force field is the corresponding total binding affinity energy. The physical energy information can be determined based on the aforementioned calculated physical parameter information and the corresponding potential energy function in the statistical physical force field. Then, the total binding affinity energy is calculated based on the physical energy information to obtain the corresponding affinity score. In some embodiments, the physical energy information includes van der Waals energy, hydrogen bond energy, hydrophobic potential energy, and / or metallic potential energy. For van der Waals energy, the van der Waals interaction corresponding to each atomic pair (r, s) is calculated using the 12-6 Lennard-Jones potential function:

[0052]

[0053] in, The weights determined for the feedforward network. Where n is the number of hybrid density network heads used to predict van der Waals force parameters. and These are the weights and Gaussian component parameters of the output of the k-th head mixed density network, respectively. The sum of the radii of the van der Waals forces. This represents the atom pair distance corresponding to the atom pair (r, s). In some embodiments, the atom pair distance... To accommodate conformational flexibility, the correction amount of the atomic pair distance can also be obtained through a small MLP based on the atomic pair characteristics. Correcting the distance between atomic pairs Therefore, the sum of the van der Waals interactions between each pair of atoms can be taken as the van der Waals force energy. For hydrogen bond energies, the statistical Gaussian potential function can be used to calculate the hydrogen bond energies corresponding to each atom pair (r, s):

[0054]

[0055] in, The weights determined for the feedforward network and The Gaussian component parameters of this atom pair are predicted by a mixed density network used to predict hydrogen bond energy parameters. The sum of the radii of the van der Waals forces. The distance between atomic pairs (r, s) is the same as that used in the aforementioned van der Waals force energy calculation, and therefore will not be repeated here, but is included by reference. Furthermore, the sum of the hydrogen bond energies corresponding to each atomic pair can be taken as the hydrogen bond energy. For hydrophobic and metallic potentials, since these interactions are less affected by the surrounding chemical environment, they can be described using relatively fixed energy functions with relatively fixed parameters, and prediction can be performed without using a mixed density network. Piecewise linear functions can be used to calculate the hydrophobic / metallic interactions corresponding to each atom pair (r, s):

[0056]

[0057] Here, w is the global interaction weight optimized during the training of the affinity assessment model. (This is relevant to the calculation of hydrophobic and metallic interactions.) , The values ​​differ. This is typically true for metal-metal interactions. =-0.7, =0.0, for hydrophobic interactions =0.5, =1.5. Therefore, the sum of the hydrophobic / metallic interactions between each atom pair can be taken as the hydrophobic potential energy / metallic potential energy. This allows for the calculation of corresponding affinity scores by combining various physical energies. .in, ,in, For learnable weights, This represents the number of rotatable keys.

[0058] After determining the affinity scores corresponding to each molecular diagram, the average affinity score for the predicted structural conformation of the complex, independent of orientation, can be calculated. For example, the affinity score E for the molecular diagram A→B is... A→B Affinity score E corresponding to the molecular diagram of B→A B→A The affinity score S corresponding to the predicted structural conformation of the complex PPA =(E A→B +E B→A ) / 2. Determine the affinity score corresponding to the predicted structural conformation of each complex. Subsequently, the affinity score of the corresponding protein-protein complex is:

[0059] .

[0060] In some embodiments, the method further includes: step S14 (not shown), determining a predicted affinity score for the protein-small molecule complex structure based on the protein-small molecule complex structure in the protein-small molecule complex dataset using an affinity assessment model to be trained; step S15 (not shown), determining a corresponding recombination loss based on the predicted affinity score and the binding affinity corresponding to the protein-small molecule complex; and step S16 (not shown), updating the affinity assessment model to be trained based on the recombination loss to obtain a trained affinity assessment model.

[0061] In some embodiments, the experimentally resolved crystal structure of the protein-small molecule complex is used as the dataset for training the affinity assessment model. Here, the process of determining the predicted affinity score of the protein-small molecule complex structure using the affinity assessment model to be trained is similar to steps S12 and S13 described above. Based on the protein-small molecule complex structure, corresponding interface information is extracted, and the corresponding predicted affinity score is determined based on this interface information using the affinity assessment model to be trained. In model training, the accurate crystal structure resolved experimentally is used for affinity score prediction, without needing to employ multiple conformations. Furthermore, considering the clear distinction between pockets and ligands in the protein-small molecule complex structure, the affinity score prediction can be performed using only a one-way score. That is, the affinity score corresponding to the molecular diagram constructed with the protein as the pocket and the small molecule as the ligand is directly used as the predicted affinity score for the protein-small molecule complex structure. Specifically, step S14 includes: extracting sample binding interface information corresponding to the protein-small molecule complex structure based on the protein-small molecule complex structure in the protein-small molecule complex dataset; for the sample binding interface information, determining the corresponding sample molecular map by using the protein in the protein-small molecule complex structure as a pocket and the small molecule in the protein-small molecule complex structure as a ligand; and determining the predicted affinity score of the protein-small molecule complex structure based on the sample molecular map using the affinity assessment model to be trained. Here, the process of determining the sample binding interface information and the process of determining the sample molecular map and determining the predicted affinity score based on the sample molecular map are similar to the processes of determining the binding interface information and the process of determining the molecular map and determining the affinity score based on the molecular map in steps S12 and S13, and therefore will not be repeated here, but are included by reference. In some embodiments, the physical scale of the aforementioned model inference process needs to be consistent with the training process, that is, the aforementioned inference process uses the same preset distance conditions as the training process to screen interface residues that meet the conditions, thereby ensuring that the model can be transferred to the protein-protein complex assessment task with zero samples to accurately assess atomic interactions.

[0062] In some embodiments, after determining the predicted affinity score, the corresponding composite loss can be determined based on the experimentally resolved binding affinity of the protein-small molecule complex, combined with the composite loss function used. In some embodiments, the composite loss includes mean squared error loss, derivative constraint loss, and negative log-likelihood loss for mixed density networks. The mean squared error loss L... MSE The mean squared error between the predicted affinity score and the binding affinity is calculated using the mean squared error loss function. The derivative constraint loss L... DERThis is determined by taking the first and second derivatives of the predicted affinity with respect to the atomic coordinates of all atoms in the protein-small molecule complex and applying physical constraints. The negative log-likelihood loss of the mixed density network corresponds to the mixed density network used, and typically includes the van der Waals mixed density network negative log-likelihood loss L. MDN-vdW Negative log-likelihood loss L in hydrogen bond mixed density network MDN-hbond Among them, L MDN-vdW It is the negative log-likelihood loss calculated for the van der Waals force energy parameter and the distances to all atomic pairs. L MDN-hbond This is the negative log-likelihood loss calculated by relating the hydrogen bond energy parameter to the distance between the atom pairs forming the hydrogen bond. The atom pairs are defined as those formed in the protein-small molecule complex crystal structure where the nearest distance to the small molecule is less than or equal to a preset distance threshold, consisting of a protein atom and a corresponding small molecule atom. The composite loss calculated above is L = c1·L. MSE +c2·L DER +c3·L MDN-vdW +c4·L MDN-hbond (Where c1, c2, c3, and c4 are preset coefficients), the corresponding parameter gradients can be determined through backpropagation, and the model parameters of the affinity assessment model can be updated using an optimizer.

[0063] Figure 2 A flowchart of a method for screening protein conjugates according to an embodiment of this application is shown. The method includes steps S21 and S22. In step S21, the device 2 determines the affinity score corresponding to each candidate protein-protein complex formed by the target protein and each candidate protein conjugate, wherein the affinity score corresponding to each candidate protein-protein complex is determined by the aforementioned method for protein-protein complex affinity assessment. In step S22, the device 2 screens out one or more corresponding protein conjugates from each candidate protein conjugate based on the affinity score corresponding to each candidate protein-protein complex.

[0064] In some embodiments, the device 2 includes, but is not limited to, user devices or network devices with information processing or computing capabilities, such as tablet computers, computers, and servers. In some embodiments, the candidate protein binders can be generated based on the target protein using generative deep learning models. These generative deep learning models include, but are not limited to, ProtGPT2, BindCraft, and Boltzgen. For each candidate protein-protein complex formed by the candidate protein binder and the target protein, the aforementioned affinity assessment method can be used to determine the affinity score corresponding to each candidate protein-protein complex. Furthermore, the affinity scores can be used to rank the complexes, selecting one or more protein binders with the highest scores, thereby improving protein binder screening efficiency, providing reliable protein binders for downstream experiments, and improving the efficiency of drug / diagnostic reagent development.

[0065] Figure 3 This diagram illustrates a device structure for evaluating the affinity of protein-protein complexes according to an embodiment of this application. The device 1 includes a first module 11, a second module 12, and a third module 13. The first module 11 generates multiple predicted structural conformations of the protein-protein complex to be evaluated using a structure prediction model. The second module 12 extracts the binding interface information corresponding to each predicted structural conformation. The third module 13 determines the affinity score of the protein-protein complex to be evaluated based on the binding interface information corresponding to each predicted structural conformation using an affinity evaluation model. The affinity evaluation model includes a statistical physical force field, a graph neural network, and a hybrid density network, and is trained on a protein-small molecule complex dataset. Figure 3 The specific implementation methods corresponding to the first module 11, the second module 12 and the third module 13 shown are the same as or similar to the specific embodiments of the aforementioned steps S11, S12 and S13, and therefore will not be repeated here. They are included here by reference.

[0066] In some embodiments, the first-third module 13 includes a first-third-first unit 131 (not shown), a first-third-second unit 132 (not shown), a first-third-third unit 133 (not shown), and a first-third-fourth unit 134 (not shown). The first-third-first unit 131, for each complex predicted structural conformation corresponding to the binding interface information, identifies one protein in the protein-protein complex to be evaluated as the ligand and pocket, and correspondingly, another protein as the pocket and ligand, to determine multiple corresponding molecular maps. The first-third-second unit 132, for each molecular map, determines the affinity score corresponding to the molecular map using an affinity assessment model, wherein the affinity assessment model includes a statistical physical force field, a graph neural network, and a mixed density network, and the affinity assessment model is trained on a protein-small molecule complex dataset. The first-third-third unit 133, based on the affinity scores corresponding to each molecular map, determines the affinity score corresponding to the complex predicted structural conformation. The first-third-fourth unit 134, based on the affinity scores corresponding to each complex predicted structural conformation, determines the affinity score corresponding to the protein-protein complex to be evaluated. Here, the specific implementation of unit 131, unit 132, unit 133, and unit 134 is the same as or similar to the specific embodiments of steps S131, S132, S133, and S134 mentioned above, so they will not be repeated here, but are included by reference.

[0067] In some embodiments, the 132 unit 132 includes a 1321 subunit 1321 (not shown), a 1322 subunit 1322 (not shown), and a 1323 subunit 1323 (not shown). For each molecular graph, the 1321 subunit 1321 determines the corresponding atom pair feature information through the graph neural network; the 1322 subunit 1322 predicts the corresponding physical parameter information based on the atom pair feature information through the hybrid density network; and the 1323 subunit 1323 determines the affinity score corresponding to the molecular graph based on the physical parameter information using the statistical physical force field. Here, the specific implementations of the 1321 subunit 1321, 1322 subunit 1322, and 1323 subunits are the same as or similar to the specific embodiments of steps S1321, S1322, and S1323 described above, and therefore will not be repeated here, but are incorporated herein by reference.

[0068] In some embodiments, the device 1 further includes a fourth module 14 (not shown), a fifth module 15 (not shown), and a sixth module 16 (not shown). The fourth module 14, based on the protein-small molecule complex structures in the protein-small molecule complex dataset, determines a predicted affinity score for the protein-small molecule complex structure using an affinity assessment model to be trained. The fifth module 15 determines a corresponding recombination loss based on the predicted affinity score and the binding affinity corresponding to the protein-small molecule complex. The sixth module 16 updates the affinity assessment model to be trained based on the recombination loss to obtain a trained affinity assessment model. Here, the specific implementations of the fourth module 14, the fifth module 15, and the sixth module 16 are the same as or similar to the specific embodiments of steps S14, S15, and S16 described above, and therefore will not be repeated here, but are incorporated herein by reference.

[0069] Figure 4 This diagram illustrates a device structure for screening protein binders according to an embodiment of this application. The device 2 includes a single module 21 and a double module 22. The single module 21 determines the affinity score corresponding to each candidate protein-protein complex formed by the target protein and each candidate protein binder, wherein the affinity score corresponding to each candidate protein-protein complex is determined by the method according to any one of claims 1 to 10. The double module 22, based on the affinity score corresponding to each candidate protein-protein complex, screens out one or more corresponding protein binders from each candidate protein binder. Here, the... Figure 4 The specific implementation methods corresponding to the two modules 21 and 22 shown are the same as or similar to the specific embodiments of the aforementioned steps S21 and S22, and therefore will not be repeated here. They are included here by reference.

[0070] Figure 5 Exemplary systems that can be used to implement the various embodiments described in this application are shown; such as Figure 5 As shown in some embodiments, system 300 can function as any of the devices described in each of the embodiments. In some embodiments, system 300 may include one or more computer-readable media having instructions (e.g., system memory or NVM / storage device 320) and one or more processors (e.g., one or more processors 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement the module and thus perform the actions described in this application.

[0071] In one embodiment, the system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of the processors 305 and / or any suitable device or component communicating with the system control module 310.

[0072] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.

[0073] System memory 315 can be used, for example, to load and store data and / or instructions for system 300. In one embodiment, system memory 315 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, system memory 315 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).

[0074] In one embodiment, the system control module 310 may include one or more input / output (I / O) controllers to provide interfaces to the NVM / storage device 320 and (one or more) communication interfaces 325.

[0075] For example, NVM / storage device 320 may be used to store data and / or instructions. NVM / storage device 320 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drive (HDD), one or more optical disc (CD) drives, and / or one or more digital universal optical disc (DVD) drives).

[0076] NVM / storage device 320 may include storage resources that are physically part of a device on which system 300 is mounted, or that can be accessed by the device without necessarily being part of it. For example, NVM / storage device 320 may be accessed via a network through one or more communication interfaces 325.

[0077] One or more communication interfaces 325 may provide the system 300 with an interface to communicate over one or more networks and / or with any other suitable device. The system 300 may wirelessly communicate with one or more components of a wireless network in accordance with any of one or more wireless network standards and / or protocols.

[0078] In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 (e.g., memory controller module 330). In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 to form a system-in-package (SiP). In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die. In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die to form a system-on-a-chip (SoC).

[0079] In various embodiments, system 300 may be, but is not limited to, a server, workstation, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, system 300 may have more or fewer components and / or different architectures. For example, in some embodiments, system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0080] In addition to the methods and devices described in the above embodiments, this application also provides a computer-readable storage medium storing computer code that, when executed, performs the method described in any of the preceding embodiments.

[0081] This application also provides a computer program product that, when executed by a computer device, performs the method described in any of the preceding claims.

[0082] This application also provides a computer device, the computer device comprising:

[0083] One or more processors;

[0084] Memory, used to store one or more computer programs;

[0085] When the one or more computer programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described in any of the preceding methods.

[0086] It should be noted that this application can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.

[0087] Furthermore, a portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0088] Communication media include media through which communication signals containing, for example, computer-readable instructions, data structures, program modules, or other data are transmitted from one system to another. Communication media can include guided transmission media (such as cables and wires (e.g., optical fibers, coaxial cables, etc.)) and wireless (unguided transmission) media capable of propagating energy waves, such as sound, electromagnetic, RF, microwave, and infrared. Computer-readable instructions, data structures, program modules, or other data can be embodied as modulated data signals in, for example, wireless media (such as carrier waves or similar mechanisms embodied as part of spread spectrum technology). The term "modulated data signal" refers to a signal whose one or more characteristics are altered or set in a manner that encodes information in the signal. Modulation can be analog, digital, or a hybrid modulation technique.

[0089] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memories such as random access memory (RAM, DRAM, SRAM); and non-volatile memories such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other media now known or hereafter developed capable of storing computer-readable information / data for use by a computer system.

[0090] Herein, one embodiment of this application includes an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the apparatus is triggered to run a method and / or technical solution based on the foregoing embodiments of this application.

[0091] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.

Claims

1. A method for assessing the affinity of protein-protein complexes, characterized in that, The method includes: Based on the protein-protein complex to be evaluated, multiple corresponding complex predicted structural conformations are generated using a structural prediction model. For each predicted structural conformation of a complex, the binding interface information corresponding to the predicted structural conformation of the complex is extracted. Based on the binding interface information corresponding to the predicted structural conformation of each complex, the affinity score of the protein-protein complex to be evaluated is determined by an affinity assessment model. The affinity assessment model includes a statistical physical force field, a graph neural network, and a hybrid density network, and is trained on a protein-small molecule complex dataset.

2. The method according to claim 1, characterized in that, The step of extracting the binding interface information corresponding to each predicted structural conformation of the complex includes: For each predicted structural conformation of the complex, a plurality of interface residues are determined from the predicted structural conformation of the complex, wherein the interface residues contain at least one heavy atom whose distance from the opposite side chain of the chain to which the interface residue belongs satisfies a preset distance condition. Based on the multiple interface residues, the binding interface information corresponding to the predicted structural conformation of the complex is determined.

3. The method according to claim 1 or 2, characterized in that, Based on the binding interface information corresponding to the predicted structural conformation of each complex, an affinity assessment model is used to determine the affinity score of the protein-protein complex to be evaluated. The affinity assessment model includes a statistical physical force field, a graph neural network, and a hybrid density network. The affinity assessment model is trained on a protein-small molecule complex dataset and includes the following: For the binding interface information corresponding to the predicted structural conformation of each complex, one protein in the protein-protein complex to be evaluated is respectively regarded as the ligand and the pocket, and the other protein is respectively regarded as the pocket and the ligand, to determine the corresponding multiple molecular maps. For each molecular graph, an affinity score is determined using an affinity assessment model, wherein the affinity assessment model includes a statistical physical force field, a graph neural network, and a hybrid density network, and is trained on a protein-small molecule complex dataset. Based on the affinity scores corresponding to each molecular diagram, the affinity scores corresponding to the predicted structural conformations of the complex are determined. The affinity score of the protein-protein complex to be evaluated is determined based on the affinity score corresponding to the predicted structural conformation of each complex.

4. The method according to claim 3, characterized in that, For each molecular graph, an affinity score is determined using an affinity assessment model. This affinity assessment model includes a statistical physical force field, a graph neural network, and a hybrid density network. The model is trained on a protein-small molecule complex dataset and includes the following parameters: For each molecular graph, the corresponding atom pair feature information is determined through the graph neural network. Based on the atom pair feature information, the corresponding physical parameter information is predicted through the hybrid density network; Based on the physical parameter information, the affinity score corresponding to the molecular diagram is determined using the statistical physical force field.

5. The method according to claim 4, characterized in that, The physical parameter information includes van der Waals force parameters and / or hydrogen bond energy parameters.

6. The method according to claim 4 or 5, characterized in that, The determination of the affinity score corresponding to the molecular diagram based on the physical parameter information and using the statistical physical force field includes: Based on the physical parameter information, the corresponding physical energy information is determined using the statistical physical force field; Based on the physical energy information, the affinity score corresponding to the molecular diagram is determined.

7. The method according to claim 6, characterized in that, The physical energy information includes van der Waals force energy, hydrogen bond energy, hydrophobic potential energy, and / or metallic potential energy.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Based on the protein-small molecule complex structures in the protein-small molecule complex dataset, the predicted affinity score of the protein-small molecule complex structure is determined by the affinity assessment model to be trained. Based on the predicted affinity score and the binding affinity corresponding to the protein-small molecule complex, the corresponding complexation loss is determined; Based on the composite loss, the affinity assessment model to be trained is updated to obtain a trained affinity assessment model.

9. The method according to claim 8, characterized in that, The step of determining the predicted affinity score of the protein-small molecule complex structure based on the protein-small molecule complex structure in the protein-small molecule complex dataset, using the affinity assessment model to be trained, includes: Based on the protein-small molecule complex structures in the protein-small molecule complex dataset, extract the sample binding interface information corresponding to the protein-small molecule complex structures; For the sample binding interface information, the protein in the protein-small molecule complex structure is used as a pocket and the small molecule in the protein-small molecule complex structure is used as a ligand to determine the corresponding sample molecular map; Based on the sample molecular map, the predicted affinity score of the protein-small molecule complex structure is determined using the affinity assessment model to be trained.

10. The method according to claim 8 or 9, characterized in that, The composite loss includes mean square error loss, derivative constraint loss, and negative log-likelihood loss for hybrid density networks.

11. A method for screening protein binders, characterized in that, The method includes: For candidate protein-protein complexes formed by the target protein and each candidate protein binder, the affinity score corresponding to each candidate protein-protein complex is determined, wherein the affinity score corresponding to each candidate protein-protein complex is determined by the method of any one of claims 1 to 10. Based on the affinity scores of each candidate protein-protein complex, one or more corresponding protein binders are selected from each candidate protein binder.

12. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 11.

13. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 11.

14. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method as described in any one of claims 1 to 11.