A method and system for drug performance evaluation based on chemical space clustering

By using chemical spatial clustering and hierarchical calibration frameworks, combined with soft allocation strategies and ensemble learning models, drug performance evaluation results are generated, solving the reliability problem of drug performance prediction in existing technologies and enabling differentiated evaluation and credibility judgment of molecules with different structural types.

CN122091275BActive Publication Date: 2026-08-25XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610552061.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-24
Publication Date
2026-08-25
Estimated Expiration
2046-04-24

AI Technical Summary

Technical Problem

Existing drug performance prediction methods can only output point prediction values ​​and cannot provide differentiated reliability assessments for different structural types of molecules and different safety sensitivities in R&D scenarios, leading to potential biases in R&D decisions.

Method used

A chemical spatial clustering-based approach is adopted. By acquiring molecular structure information, calculating chemical spatial distance, performing density-aware clustering, constructing a hierarchical calibration framework, and combining a soft assignment strategy and an ensemble learning model, drug performance evaluation results carrying confidence intervals associated with structure types are generated.

Benefits of technology

It provides statistical reliability information for drug performance evaluation results, adapts to the decision-making needs of different R&D stages, reduces the probability of errors in homogeneous molecular clustering, and improves the applicability and credibility of evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122091275B_ABST
    Figure CN122091275B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of drug performance evaluation, and particularly relates to a drug performance evaluation method and system based on chemical space clustering, comprising: extracting structural features of known drug molecules in a calibration data set and calculating chemical space distance; adopting a density-aware clustering strategy to cluster the calibration data set, respectively performing statistical calibration in each chemical similarity cluster, and constructing a chemical space hierarchical calibration framework; adopting a graph neural network integrated model to predict the drug performance of a drug molecule to be tested, obtaining a prediction result and uncertainty estimation, determining the assignment weight of each chemical similarity cluster based on the structural features and the chemical space hierarchical calibration framework, and generating a confidence interval; according to the safety level corresponding to the drug performance to be evaluated, adjusting the sensitivity of the confidence interval, and outputting the prediction result and the confidence interval after reliability evaluation. The technical scheme of the present application provides reliable quantitative information with statistical basis for differentiated decision-making in early drug development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of drug performance evaluation technology, and in particular to a drug performance evaluation method and system based on chemical spatial clustering. Background Technology

[0002] Drug development is a lengthy, costly, and risky process. Statistics show that bringing a new drug to market typically takes over a decade and involves investments of hundreds of millions of dollars, with approximately 90% of candidate compounds failing during clinical trials. Poor ADMET properties (absorption, distribution, metabolism, excretion, and toxicity) are a major cause of drug development failure. Accurately predicting the ADMET properties of candidate molecules early in drug development and promptly eliminating potentially risky compounds is crucial for reducing development costs and increasing the success rate.

[0003] In recent years, with the rapid development of machine learning and deep learning technologies, data-driven ADMET prediction methods have made significant progress. Existing technologies mainly fall into two categories: First, traditional machine learning methods, which use algorithms such as random forests and support vector machines to build prediction models based on molecular descriptors (e.g., LogP, molecular weight, number of hydrogen bond donors and acceptors) and fingerprint features. However, these methods rely on manual feature engineering and have limited ability to represent complex molecules. Second, graph neural network methods, which represent molecules as graph structures with atoms as nodes and chemical bonds as edges. Message-passing neural networks or graph convolutional networks are used to learn molecular representations. Representative tools such as ADMETlab and pkCSM have been widely applied in drug development practice, showing a significant improvement in prediction accuracy compared to traditional methods.

[0004] Chinese patent application CN114842920A discloses a method for predicting molecular properties. This method involves inputting a molecular structure diagram into a graph transformation coding network to extract atomic features and bond features of chemical bonds. These features are then processed by a feature fusion network to obtain a molecular structure feature vector, ultimately outputting a quantitative value of the molecule's target properties. This achieves automatic prediction of molecular chemical properties, improving prediction efficiency without human intervention. However, this method only outputs point prediction values ​​and cannot reflect the reliability differences in prediction conclusions across different structural types of molecules. In real-world drug development scenarios, candidate drugs with different indications and safety sensitivities have significantly different confidence requirements for prediction results. A single numerical value is insufficient to support researchers in making reasonable judgments at high-risk decision points (such as candidate compound screening and toxicity risk assessment), posing a potential risk of decision-making bias due to ignoring prediction uncertainties. Summary of the Invention

[0005] In view of this, the present invention provides a drug performance evaluation method and system based on chemical spatial clustering to solve the technical problem that existing drug performance prediction methods can only output point prediction values ​​and cannot provide differentiated reliability assessments for different structural types of molecules and different safety sensitivities in research and development scenarios.

[0006] The technical solution of this invention is implemented as follows: On one hand, this invention provides a drug performance evaluation method based on chemical spatial clustering, comprising: S1. Obtain molecular structure information of multiple known drug molecules in the calibration dataset, extract structural features, and determine the chemical spatial distance between each molecule based on the structural features; S2. Based on chemical spatial distance, a density-aware clustering strategy is used to cluster the calibration dataset to obtain multiple chemical similarity clusters. Statistical calibration is then performed within each chemical similarity cluster to obtain the calibration results corresponding to each chemical similarity cluster, thereby constructing a chemical spatial hierarchical calibration framework. S3. Obtain the molecular structure information of the drug molecule to be tested, use a prediction model to predict the drug performance of the drug molecule to be tested, obtain the prediction results and uncertainty estimates, and extract the structural features of the drug molecule to be tested. S4. The structural features and chemical spatial hierarchical calibration framework of the drug molecule to be tested calculates the correlation degree based on structural distance and normalizes it into allocation weights through a soft allocation strategy. S5. Based on the weighted distribution, the calibration results of each chemical similarity cluster are fused to obtain the fused calibration result of the drug molecule to be tested. Combined with the prediction results and uncertainty estimation, the confidence interval of the drug molecule to be tested is generated. S6. Based on the safety level corresponding to the performance of the drug to be evaluated, the confidence interval is sensitively adjusted, and the prediction results and confidence intervals are evaluated for reliability. The evaluation results are then output as the drug performance evaluation results of the drug molecule to be tested.

[0007] In the above technical solution, preferably, the structural features are composed of skeletal features and topological features. The skeletal features are obtained by extracting the Bemis-Murcko skeleton of the molecule and performing Morgan fingerprint encoding, and the topological features are obtained by performing Morgan circular fingerprint encoding on the overall structure of the molecule.

[0008] In the above technical solution, preferably, in step S1, the chemical spatial distance is a skeleton-topological joint weighted distance, and the calculation method of the skeleton-topological joint weighted distance includes: ; ; in, These are the weighting coefficients for the global topological distance components in the joint distance metric; For calibration of concentrated molecules With molecules The skeleton-topology joint weighted distance between them, with a value range of [0,1]; For molecules With molecules The Tanimoto coefficient between Morgan's circular fingerprints; For molecules With molecules The Tanimoto coefficients between Bemis-Murcko skeleton features; The Spearman rank correlation coefficient between random variables X and Y; To calibrate the non-uniform fractional differences of all molecular pairs in the set.

[0009] In the above technical solution, preferably, step S2 specifically includes: S21. Based on the chemical spatial distance, calculate the local density of each molecule in the calibration dataset in the chemical space to obtain the local density value corresponding to each molecule; S22. Calculate the local density of each molecule in the calibration dataset based on the skeleton-topology joint weighted distance. Divide the calibration dataset into high-density regions, medium-density transition regions and low-density regions according to the local density. Use a differentiated clustering strategy for different density regions to obtain multiple chemical similarity clusters. At the same time, calculate the chemical spatial coverage radius of the calibration dataset. S23. Within each chemical similarity cluster, based on the significance level determined according to the safety level of the task to be evaluated, the quantiles of the non-consistency scores of the calibration molecules within the cluster are calculated independently, and the calibration quantiles of each chemical similarity cluster are determined. Each chemical similarity cluster and its corresponding calibration quantile together constitute a chemical spatial hierarchical calibration framework.

[0010] In the above technical solution, preferably, in step S22, adopting a differentiated clustering strategy for different density regions to obtain multiple chemical similarity clusters specifically includes: determining the optimal values ​​of the number of candidate clusters in high-density regions and the number of candidate clusters in medium-density transition regions respectively through a silhouette coefficient maximization strategy. The silhouette coefficient maximization strategy is as follows: in each density region, the number of candidate clusters within a preset range is clustered sequentially, the average silhouette coefficient corresponding to each number of candidate clusters is calculated, and the number of clusters corresponding to the maximum average silhouette coefficient is determined as the optimal number of clusters in that density region.

[0011] In the above technical solution, preferably, in step S3, the prediction model is a multi-architecture graph neural network ensemble model trained using an ensemble learning strategy. The ensemble model contains multiple graph neural network sub-models with different architectures, and each sub-model is trained independently on the training set. When predicting the drug molecule to be tested, each sub-model takes the molecular graph representation obtained by converting the molecular structure information as input and independently outputs the corresponding drug performance prediction value. The mean of the prediction values ​​of each sub-model is used as the prediction result, and the standard deviation of the prediction values ​​of each sub-model is used as the uncertainty estimate.

[0012] In the above technical solution, preferably, step S4 specifically includes: S41. Based on the structural characteristics of the drug molecule to be tested, calculate the skeleton-topological joint weighted distance from the drug molecule to the center of each chemical similarity cluster in the chemical spatial hierarchical calibration framework. S42. Determine the local adaptive temperature parameters corresponding to the drug molecule to be tested based on the skeleton-topology joint weighted distance from the drug molecule to the nearest cluster center and the chemical space coverage radius of the calibration dataset. S43. The skeleton-topology joint weighted distance of the drug molecule to be tested is scaled using the local adaptive temperature parameter, and the scaled distance is converted into a weight value that satisfies the normalization constraint through the normalization exponential function to obtain the assigned weight of the drug molecule to be tested corresponding to each chemical similarity cluster.

[0013] In the above technical solution, preferably, in step S43, the weighting of each chemical similarity cluster corresponding to the drug molecule to be tested is calculated as follows: ; ; in, The soft assignment weights of the molecule to be tested to the k-th chemical similarity cluster satisfy the following conditions: and ; The denoted distance is the framework-topological joint weighted distance from the molecule to be tested to the center of the k-th cluster; K is the total number of chemically similar clusters. The local adaptive temperature parameter corresponding to the molecule to be tested; Reference temperature parameter; This is the distance sensitivity coefficient; The distance is the skeleton-topological joint weighted distance from the molecule to be tested to the nearest cluster center; To calibrate the chemical space coverage radius of the dataset.

[0014] In the above technical solution, preferably, step S6 specifically includes: S61. Classify the drug performance to be evaluated into safety-critical, important, and screening types according to clinical safety level. Verify whether the nominal coverage of the confidence interval generated in step S5 meets the coverage target of the safety level corresponding to the task to be evaluated. If it meets the target, directly label the confidence interval as the corresponding safety level and coverage level. If it does not meet the target, re-perform the confidence interval fusion calculation in step S5 with the significance level of the actual corresponding safety level to obtain the confidence interval after safety level correction. S62. Calculate the comprehensive confidence score from three dimensions: relative width of the corrected confidence interval, prediction uncertainty, and cluster allocation concentration. Divide the evaluation results into multiple confidence levels based on the comprehensive confidence score. When the comprehensive confidence score is lower than a preset threshold, add an extrapolation warning and a quantitative description of the chemical spatial location of the drug molecule to be tested to the output results. Output the confidence interval, prediction results, and confidence level after sensitivity adjustment as the drug performance evaluation results of the drug molecule to be tested.

[0015] In the above technical solution, preferably, the drug performance evaluation method further includes: S7. When new known drug molecule data is obtained and the amount of new data reaches a preset update threshold, the new data is included in the calibration dataset. Based on the expanded calibration dataset, the skeleton-topology joint weighted distance between each molecule is recalculated and the weighting coefficient is updated. The local density of each molecule is recalculated and the cluster structure and chemical space coverage radius are updated by maximizing the contour coefficient. The quantiles of the expanded non-consistency scores in each chemical similarity cluster are recalculated to update the calibration quantiles of each cluster.

[0016] In addition, the present invention also provides a drug performance evaluation system based on chemical spatial clustering to achieve the above-described method, comprising: The chemical space modeling module is used to construct a chemical space clustering model based on the structural features of calibration set molecules and form a chemical space hierarchical calibration framework associated with the safety level of the task to be evaluated. The performance prediction module is used to predict drug performance based on the molecular structure information of the drug molecule to be tested, and to obtain the predicted value and uncertainty estimate. The calibration and evaluation module is used to perform statistical calibration and reliability evaluation on the predicted values ​​based on the distribution relationship of the drug molecules to be tested in the chemical space clustering model, and generate a drug performance evaluation report. The storage module is used to store the chemical spatial clustering model, the chemical spatial hierarchical calibration framework, and the drug performance evaluation report; The dynamic update module is used to trigger the update of the chemical spatial clustering model and the chemical spatial hierarchical calibration framework when new experimental data meet preset conditions, and write the update results into the storage module.

[0017] The present invention has the following advantages over the prior art: (1) This invention provides a drug performance evaluation method and system based on chemical spatial clustering. By combining the chemical spatial hierarchical calibration framework with the statistical calibration confidence interval, the drug performance evaluation results are expanded from a single predicted value to a complete evaluation output carrying the confidence interval associated with the structure type. Researchers can obtain the statistical reliability information of the predicted conclusion under the current molecular structure type while obtaining the predicted value of the drug properties. At the same time, the confidence interval width and credibility judgment criteria are differentiated according to the safety level of the evaluation task, so that the evaluation conclusion can adapt to different decision-making scenarios from early high-throughput screening to evaluation of safety key indicators, and assist researchers in making a decision on candidate compounds that matches the risk sensitivity at different research and development stages.

[0018] (2) In terms of chemical spatial characterization, the present invention adopts a joint weighted distance measure that combines skeletal features and global topological features to measure the chemical spatial distance between molecules. Compared with a single distance measure that only relies on global topological features, it can capture the structural differences at both the skeletal level and the local topological level during the clustering stage. This helps to reduce the probability that molecules with similar global topology but different skeletal structures are incorrectly classified into the same similarity cluster, improve the homogeneity of structure-activity relationship behavior within the cluster, and enhance the statistical representativeness of the calibration quantiles of each cluster.

[0019] (3) In terms of soft assignment weight smoothness control, the present invention introduces a local adaptive temperature parameter, which dynamically adjusts the smoothness of the soft assignment weight according to the position of the molecule to be tested in the chemical space. Different weight distributions are applied to molecules in dense and sparse regions of the chemical space, which helps to alleviate the problem of inconsistent actual and nominal coverage between molecules in different chemical space positions caused by heteroscedasticity of drug property data under global fixed temperature parameters, and improves the conditional coverage effectiveness of confidence intervals to a certain extent.

[0020] (4) Regarding the continued applicability of the calibration framework, the present invention supports incremental triggering of calibration framework updates. The update process only involves recalculation of the calibration phase and does not require retraining of the underlying graph neural network model. When new experimental data accumulates to a preset scale, parameter updates can be completed quickly, which helps to maintain the consistency between the evaluation results and the current data distribution and adapts to the actual scenario of continuous data accumulation in drug development. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of the drug performance evaluation method based on chemical spatial clustering of the present invention; Figure 2 This is a flowchart of the chemical spatial distance calculation method of the present invention; Figure 3 This is a flowchart illustrating the construction process of the density-sensing clustering and hierarchical calibration framework of the present invention. Figure 4 This is a flowchart of the soft-assignment weight calculation and confidence interval generation process of the present invention; Figure 5 This is a flowchart of the credibility assessment and output process for the present invention. Detailed Implementation

[0023] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0024] like Figure 1 As shown, this invention provides a drug performance evaluation method based on chemical spatial clustering, comprising: S1. Obtain molecular structure information of multiple known drug molecules in the calibration dataset, extract structural features, and determine the chemical spatial distance between each molecule based on the structural features; S2. Based on chemical spatial distance, a density-aware clustering strategy is used to cluster the calibration dataset to obtain multiple chemical similarity clusters. Statistical calibration is then performed within each chemical similarity cluster to obtain the calibration results corresponding to each chemical similarity cluster, thereby constructing a chemical spatial hierarchical calibration framework. S3. Obtain the molecular structure information of the drug molecule to be tested, use a prediction model to predict the drug performance of the drug molecule to be tested, obtain the prediction results and uncertainty estimates, and extract the structural features of the drug molecule to be tested. S4. Based on the structural characteristics and chemical spatial hierarchical calibration framework of the drug molecule to be tested, the degree of correlation between the drug molecule to be tested and each chemical similarity cluster is determined, and a soft allocation strategy is used to obtain the allocation weight of the drug molecule to be tested corresponding to each chemical similarity cluster. S5. Based on the weighted distribution, the calibration results of each chemical similarity cluster are fused to obtain the fused calibration result of the drug molecule to be tested. Combined with the prediction results and uncertainty estimation, the confidence interval of the drug molecule to be tested is generated. S6. Based on the safety level corresponding to the performance of the drug to be evaluated, the confidence interval is sensitively adjusted, and the prediction results and confidence intervals are evaluated for reliability. The evaluation results are then output as the drug performance evaluation results of the drug molecule to be tested.

[0025] In one embodiment of the present invention, such as Figure 2 As shown, in step S1, the structural features are composed of skeletal features and topological features. The skeletal features are obtained by extracting the Bemis-Murcko skeletal structure of the molecule and performing Morgan fingerprint encoding, while the topological features are obtained by performing Morgan circular fingerprint encoding on the overall structure of the molecule. The chemical spatial distance is the skeletal-topological joint weighted distance.

[0026] For each known drug molecule in the calibration dataset, the SMILES (Simplified Molecular Linear Input Specification) is used. The Molecular Input Line Entry System (MILS) string was used as the input format for molecular structure information. Two types of molecular fingerprints were extracted using the RDKit cheminformatics toolkit. The first type is global topological features, generated using the Morgan circular fingerprint algorithm (radius 2, bit count 2048). This fingerprint captures the local topological chemical environment within a radius of 2 hops, centered on each atom, and is sensitive to the global structural similarity of the molecule. This is denoted as... The corresponding fingerprint vector. The second category is skeletal features, i.e., skeletal features. First, the core skeletal structure of the molecule is extracted using the Bemis-Murcko skeletal decomposition method—that is, the ring system and connecting chain structure retained after removing side chains. Then, the extracted skeletal structure is encoded using the Morgan circular fingerprint (radius of 2, bit count of 2048) with the same parameters. This fingerprint specifically characterizes the structural features at the skeletal level of the molecule, denoted as . The corresponding fingerprint vectors. The two types of fingerprints capture different levels of information about the molecular structure, and are both used for subsequent calculations of chemical spatial distances.

[0027] Based on the two types of fingerprints mentioned above, the Tanimoto coefficient is used to measure the similarity between molecules, and then the backbone-topology joint weighted distance between any two molecules i and j in the calibration dataset is calculated. : ; in, is the skeleton-topological joint weighted distance between molecule i and molecule j, with a value range of [0,1]. is the Tanimoto coefficient between the Morgan circular fingerprints of molecule i and molecule j, which reflects the global topological similarity between the two molecules and takes a value in the range [0,1]. The Tanimoto coefficient is the skeletal feature ratio between molecules i and j, reflecting the skeletal structural similarity between the two molecules, and its value ranges from [0,1]. These are the weighting coefficients for the global topological distance components, with values ​​ranging from [0,1]. The weighting coefficients for the skeleton distance components are 1, and their sum is 1.

[0028] in, The data-driven approach was used to determine: firstly, an integrated prediction model was employed to calculate the inconsistency score for each molecule in the calibration dataset. (That is, the scalar value of the model prediction error after normalization, reflecting the degree of prediction bias of each molecule under the current prediction task), and then for any molecule pair Calculate the non-consistent fraction difference This value reflects the degree of difference in the structure-activity relationship behavior between the two molecules under the current prediction task; This is the basic inconsistent score without task weighting factors. Let be the model's predicted value for molecule i. The true label value of molecule i. Let be the standard deviation of the model's prediction for molecule i; Then calculate the global topological distance respectively. and Between, skeleton distance and Spearman rank correlation coefficient between The weighting coefficients are determined based on the relative magnitudes of the correlation coefficients: ; in, Represents the global topological distance vector of all molecular pairs in the calibration set. With the corresponding Spearman rank correlation coefficient between vectors Similarly, when skeletal differences are more strongly correlated with structure-activity relationship differences (such as metabolic pathway-related endpoints), the correlation is stronger. When the distance approaches 0, the skeleton distance dominates the clustering structure; when the correlation of global topological differences is stronger (such as the endpoint dominated by substituent effects). Approaching 1. Calculated only once during the calibration phase, and separately for different ADMET prediction tasks, the distance metric naturally corresponds to the task characteristics. This jointly weighted distance metric incorporates structural information at both the skeleton and topological levels. Compared to a single global fingerprint distance, it exhibits stronger structure-behavior relationship discrimination capabilities on datasets with high structural diversity, reducing clustering errors caused by different skeletons but similar global topologies.

[0029] In one embodiment of the present invention, such as Figure 3 As shown, step S2 includes: S21. First, determine the salience level based on the safety level of the task to be evaluated. and task weight factor .

[0030] Before proceeding with the clustering and calibration process, the significance level required for subsequent calculations is first determined based on the clinical safety level of the task to be evaluated. and task weight factor The tasks to be evaluated are divided into three categories according to safety level: safety-critical endpoints (including cardiotoxicity hERG inhibition and hepatotoxicity DILI), which correspond to higher confidence requirements. This corresponds to a nominal coverage rate of 95%–99%, and a task weighting factor. Important endpoints (including oral bioavailability, blood-brain barrier permeability, and plasma protein binding rate). This corresponds to a nominal coverage rate of 90%–95%, and a task weighting factor. Screening endpoints (including water solubility and membrane permeability) are used for early-stage large-scale virtual screening, where the accuracy requirements for prediction are relatively relaxed. This corresponds to a nominal coverage rate of 80%–90%, and a task weighting factor. The determined and This remains unchanged throughout the entire calculation process from S2 to S5.

[0031] S22. Calculate the local density of each molecule in the calibration dataset based on the skeleton-topology joint weighted distance. Divide the calibration dataset into high-density regions, medium-density transition regions and low-density regions according to the local density. Then, adopt a differentiated clustering strategy for different density regions to obtain multiple chemical similarity clusters.

[0032] Based on skeleton-topology joint weighted distance Calculate the local chemical spatial density of each molecule i in the calibration set. : ; in, denoted as the local chemical space density of molecule i, a larger value indicates a denser chemical space around the molecule; For calibrating the dataset; The distance cutoff parameter controls the spatial decay rate of the local density, and is taken as the median of the intermolecular distance distribution in the calibration set.

[0033] based on ,by (75th percentile of local density) and (The 25th percentile of local density) is a double threshold, setting local densities higher than... Molecules are grouped into high-density regions, with local densities lower than [a certain value]. Molecules with high density are assigned to low-density regions, while the remaining molecules are assigned to medium-density transition regions. Within high-density regions, candidates are grouped according to their number of clusters. Perform K-means clustering, and in the medium-density transition region, according to... K-means clustering is performed, treating low-density regions as a single sparse cluster. For each density region, the number of candidate clusters is traversed within a preset range, and the average silhouette coefficient (a comprehensive indicator of intra-cluster compactness and inter-cluster separation; a higher value indicates better clustering quality) is calculated for each candidate cluster. The cluster that maximizes the average silhouette coefficient is selected. and Combining these elements, along with a single sparse cluster in a low-density region, yields multiple chemically similar clusters. The total number of clusters is denoted as […]. .

[0034] Simultaneously calculate the chemical space coverage radius of the calibration set. , is defined as the 95th percentile of the skeleton-topological joint weighted distance from all molecules in the calibration set to their respective nearest cluster centers.

[0035] S23. Within each chemical similarity cluster, a layered coverage error minimization method is used, based on the determination of the corresponding significance level according to the safety level. Calculate the calibration quantiles for each cluster separately. This allows for the construction of a chemical spatial hierarchical calibration framework.

[0036] In each chemical similarity cluster k ( Within the cluster, for each calibrated molecule i, the ensemble model's prediction value for that molecule is used. and the true value Based on this, calculate the normalized non-consistent score. The score is defined as the absolute prediction error divided by the uncertainty estimate of the ensemble model, and then weighted by the task weight factor. The scalar obtained by weighting, i.e. ,in This represents the prediction standard deviation of the ensemble model. Collect the standard deviations of all calibration molecules within cluster k. A layered coverage error minimization method is adopted to minimize the nominal coverage target. Calculate the calibration quantiles of this cluster. Even if the empirical coverage of the calibration molecules within a cluster is closest to the non-consistent fractional quantile corresponding to the target coverage, then each of the K clusters obtains its corresponding calibration quantile. This is the core parameter of the chemical spatial stratification calibration framework.

[0037] The density-aware clustering strategy described above combines density information of chemical spatial distribution with a skeleton-topology bilayer distance metric, enabling high-density structurally dense regions and low-density sparse regions to obtain clustering granularities adapted to their distribution characteristics. The consistency of structure-property relationships within each cluster is improved compared to single global clustering, which helps to improve the statistical representativeness of the calibration quantiles of each cluster in the hierarchical calibration framework.

[0038] In one embodiment of the present invention, such as Figure 4 As shown, step S3 includes: the prediction model is a multi-architecture graph neural network ensemble model trained using an ensemble learning strategy. The ensemble model contains multiple graph neural network sub-models with different architectures, and each sub-model is trained independently on the training set; when predicting the drug molecule to be tested, each sub-model takes the molecular graph representation obtained by converting the molecular structure information as input and independently outputs the corresponding drug performance prediction value; the mean of the prediction values ​​of each sub-model is used as the prediction result, and the standard deviation of the prediction values ​​of each sub-model is used as the uncertainty estimate.

[0039] Specifically, the SMILES string of the drug molecule to be tested is converted into a molecular graph representation, where atoms in the molecule are graph nodes and chemical bonds are graph edges. Node features include atomic number, hybridization type, formal charge, and whether it is in a loop, while edge features include bond type and whether it is an aromatic bond. The ensemble model includes five graph neural network sub-models with different architectures: Graph Convolutional Network (GCN), Graph Attention Network (GAT), Message Passing Neural Network (MPNN), Graph Isomorphic Network (GIN), and a hybrid architecture of GNN and traditional molecular descriptors. Each sub-model is independently trained on the training set using mean squared error (MSE) as the loss function and Adam optimizer.

[0040] For the drug molecule to be tested, the five sub-models independently output predicted values, and the average of the five predicted values ​​is used as the ensemble prediction result. The standard deviation of the five predicted values ​​is used as an uncertainty estimate. , It reflects the consistency of the predictions made by each sub-model for the current molecule to be tested.

[0041] Meanwhile, global topological features and backbone features of the drug molecules to be tested are extracted in the same way as in step S1, and used as structural feature inputs for calculating chemical spatial distance.

[0042] In one embodiment of the present invention, step S4 includes: S41. Based on the structural characteristics of the drug molecule to be tested, calculate the skeleton-topological joint weighted distance from the drug molecule to the center of each chemical similarity cluster in the chemical spatial hierarchical calibration framework.

[0043] Using weighting coefficients Calculate the molecular structure of the drug to be tested To the center of the kth chemical similarity cluster Skeleton-Topological Joint Weighted Distance ( ): ; in, Let be the center of the k-th chemical similarity cluster, and take the mean of the molecular fingerprint vectors of all members within the cluster; The Tanimoto coefficient is the relationship between the global topological features of the molecule under test and the global topological features of the cluster center. is the Tanimoto coefficient between the skeletal features of the molecule to be tested and the skeletal features of the cluster center. Wherein The distance is the skeleton-topological joint weighted distance from the molecule to be tested to the nearest cluster center.

[0044] S42. Based on the joint weighted distance of each skeleton-topology, the softmax function with local adaptive temperature parameter is used to calculate the soft assignment weight of the drug molecule to each chemical similarity cluster.

[0045] ; ; in, The soft assignment weights of the molecule to be tested to the k-th chemical similarity cluster satisfy the following conditions: and ; The local adaptive temperature parameter corresponding to the molecule to be tested; The reference temperature parameter controls the baseline smoothness of the soft distribution; For distance sensitivity coefficient, control Follow Increased growth rate; The distance is the skeleton-topological joint weighted distance from the molecule to be tested to the nearest cluster center; To calibrate the chemical space coverage radius. and By applying a layered coverage error minimization method to the calibration set, the difference between the nominal coverage and the actual coverage on the calibration set is minimized.

[0046] When the molecule to be tested is very close to the center of a cluster ( When it is smaller, near The soft-assignment weights are more concentrated, allowing for greater utilization of calibration information from the dominant clusters; when the analyte is far from all cluster centers ( When the size is large, tending towards the chemical spatial extrapolation region, As the weights of each cluster increase, the weights of each cluster tend to be more evenly distributed. The system automatically integrates the calibration information of multiple clusters in a more conservative manner, resulting in a reasonable amplification effect of uncertainty.

[0047] In one embodiment of the present invention, step S5 includes: Using weight allocation Calibration quantiles for each chemical similarity cluster Weighted fusion is performed to obtain the fusion calibration quantile for the analyte molecule. : ; in, The local calibration information is integrated by weighting the distance between the molecule to be tested and each chemically similar cluster.

[0048] Combined predictions from the ensemble model With fusion calibration quantile Generate confidence intervals for statistical calibration: ; in, The mean of the predictions from the ensemble model; The standard deviation of the predictions of the integrated model (uncertainty estimate); To be at the significance level The generated confidence interval has a nominal coverage rate of The width of the confidence interval is determined by... and Together, it is determined that when the overlap between the test molecule and the training chemical space is low, The larger the range, the wider it naturally becomes, reflecting higher forecast uncertainty.

[0049] In one embodiment of the present invention, such as Figure 5 As shown, step S6 includes: S61. The drug performance to be evaluated is classified into safety-critical, important and screening types according to the clinical safety level. Based on the safety level of the task to be evaluated, the confidence interval is adjusted for sensitivity according to the corresponding significance level to obtain the confidence interval after safety level correction.

[0050] The confidence interval has been determined in advance. Therefore, the sensitivity adjustment in this step is reflected in labeling the confidence interval of the final output with a confidence level linked to the corresponding safety level, ensuring that the nominal coverage of the confidence interval in the assessment report is consistent with the safety sensitivity of the task. For safety-critical endpoints (such as hERG inhibition and DILI), the confidence interval corresponds to 95% to 99% coverage to ensure coverage reliability at high-risk decision points (such as toxicity screening); for important endpoints, it corresponds to 90% to 95% coverage; and for screening endpoints, it corresponds to 80% to 90% coverage, providing a reasonable range of uncertainty while meeting the efficiency requirements of large-scale screening.

[0051] S62. Calculate the comprehensive confidence score from three dimensions: relative width of the corrected confidence interval, prediction uncertainty, and cluster allocation concentration. Divide the evaluation results into multiple confidence levels based on the comprehensive confidence score. When the comprehensive confidence score is lower than a preset threshold, add an extrapolation warning and a quantitative description of the chemical spatial location of the drug molecule to be tested to the output results. Output the confidence interval, prediction results, and confidence level after sensitivity adjustment as the drug performance evaluation results of the drug molecule to be tested.

[0052] The overall credibility score is calculated by weighting indicators from three dimensions: ; in, The overall credibility score ranges from [0,1], with values ​​closer to 1 indicating a more credible prediction. The relative width of the confidence interval reflects the degree to which the confidence interval widens relative to the predicted value. ; The relative width of the empirical maximum confidence interval is used to normalize W to [0,1]. The coefficient of variation of the predictions of the ensemble model for the target molecule reflects the degree of consistency in the predictions of the various sub-models within the ensemble model. ; Assign cluster concentration, which is the maximum soft assignment weight of the molecule to be tested across all chemically similar clusters. , The larger the value, the more clearly the chemical spatial assignment of the molecule being tested is defined, and the better it matches the calibration information of a certain local chemical structure type; , , All are weighting coefficients, satisfying + + =1.

[0053] in accordance with Five levels of credibility: extremely high credibility ( ), outputting prediction results and suggesting their direct use in decision-making; high reliability ( The output predicts the results, and it is recommended to combine this with expert judgment; medium confidence level ( Output the prediction results and mark warnings, suggesting supplementary experimental verification; low confidence ( The output predicts the results and marks them with a high warning, and it is not recommended to use them directly for critical decisions; extremely low reliability ( An extrapolation warning is added, indicating that the molecule may be outside the effective coverage of the calibrated chemical space. Simultaneously, a quantitative description of the chemical spatial location of the analyte is output—including the framework-topologically weighted distance to the nearest cluster center. With calibration set coverage radius The ratio is used as a quantitative reference for the degree of extrapolation.

[0054] The aforementioned multidimensional credibility assessment framework integrates three independent information dimensions—confidence interval width, model internal uncertainty, and chemical space allocation concentration—in a weighted manner. Compared to credibility judgment methods that rely solely on prediction error or interval width, this framework provides a more comprehensive characterization of the prediction reliability of molecules with different structural types. It helps researchers adopt differentiated decision-making approaches for candidate molecules with varying degrees of safety sensitivity.

[0055] In one embodiment of the present invention, the drug performance evaluation method further includes: S7. When new known drug molecule data is obtained and the amount of new data reaches the preset update threshold, the new data is added to the calibration dataset. The cluster structure and chemical spatial coverage radius are updated to be maximized using the silhouette coefficient. Within each chemical similarity cluster, the quantiles of the expanded inconsistency scores are recalculated to update the calibration quantiles of each cluster. The update adopts an incremental update method, that is, only the quantile recalculation is performed locally on the chemical similarity clusters affected by the new data, while the original calibration quantiles are retained for unaffected clusters, in order to control the overall update computation. The preset update threshold can be flexibly configured according to the actual data accumulation rate. A typical value is that the update is triggered when the number of new molecules is not less than 10% of the total amount of the current calibration set.

[0056] The present invention also provides a drug performance evaluation system based on chemical spatial clustering. The system includes a chemical spatial modeling module, a performance prediction module, a calibration evaluation module, a storage module, and a dynamic update module. The modules work together through an internal data interface.

[0057] The chemical space modeling module is responsible for receiving molecular structure data of known drugs in the calibration set, constructing a chemical space representation based on skeleton features and global topological feature extraction methods, and completing density-aware clustering and hierarchical statistical calibration based on skeleton-topology joint weighted distance. Finally, a clustering calibration framework containing calibration quantiles and safety level parameters of each similarity cluster is formed for subsequent modules to call.

[0058] The performance prediction module receives the molecular structure information of the drug molecule to be tested, uses a multi-architecture graph neural network ensemble model to output the predicted mean and uncertainty estimate of the drug molecule to be tested, and transmits the results to the calibration evaluation module.

[0059] The calibration evaluation module receives the output of the performance prediction module, calculates the soft allocation weight of the molecule to be tested relative to the center of each cluster, weights and fuses the calibration quantiles of each cluster to generate a confidence interval, and combines the safety level sensitivity adjustment and the three-dimensional comprehensive confidence score to finally output a drug performance evaluation report.

[0060] The storage module is used to persistently store the clustering model and calibration framework generated by the chemical spatial modeling module, the data of each calibration set, and the evaluation report output by the calibration evaluation module, providing a unified data read and write interface for each functional module.

[0061] The dynamic update module continuously monitors the accumulation of new experimental data. When the scale of new data meets the specified threshold conditions, it triggers an incremental update of the clustering model and calibration framework in the chemical space modeling module and writes the updated results back to the storage module to maintain the consistency between the system's evaluation capability and the actual data distribution.

[0062] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A drug performance evaluation method based on chemical spatial clustering, characterized in that, include: S1. Obtain molecular structure information of multiple known drug molecules in the calibration dataset, extract structural features, and determine the chemical spatial distance between each molecule based on the structural features; S2. Based on chemical spatial distance, a density-aware clustering strategy is used to cluster the calibration dataset to obtain multiple chemical similarity clusters. Statistical calibration is then performed within each chemical similarity cluster to obtain the calibration results corresponding to each chemical similarity cluster, thereby constructing a chemical spatial hierarchical calibration framework. S3. Obtain the molecular structure information of the drug molecule to be tested, use a prediction model to predict the drug performance of the drug molecule to be tested, obtain the prediction results and uncertainty estimates, and extract the structural features of the drug molecule to be tested. S4. Based on the structural features and chemical spatial hierarchical calibration framework of the drug molecule to be tested, the correlation degree is calculated based on the structural distance, and the correlation is normalized into allocation weights through a soft allocation strategy. S5. Based on the weighted distribution, the calibration results of each chemical similarity cluster are fused to obtain the fused calibration result of the drug molecule to be tested. Combined with the prediction results and uncertainty estimation, the confidence interval of the drug molecule to be tested is generated. S6. Based on the safety level corresponding to the performance of the drug to be evaluated, the confidence interval is sensitively adjusted, and the prediction results and confidence intervals are evaluated for reliability. The evaluation results are then output as the drug performance evaluation results of the drug molecule to be tested. Step S4 specifically includes: S41. Based on the structural characteristics of the drug molecule to be tested, calculate the skeleton-topological joint weighted distance from the drug molecule to the center of each chemical similarity cluster in the chemical spatial hierarchical calibration framework. S42. Determine the local adaptive temperature parameters corresponding to the drug molecule to be tested based on the skeleton-topology joint weighted distance from the drug molecule to the nearest cluster center and the chemical spatial coverage radius of the calibration dataset. S43. The skeleton-topology joint weighted distance of the drug molecule to be tested is scaled using the local adaptive temperature parameter, and the scaled distance is converted into a weight value that satisfies the normalization constraint through the normalization exponential function to obtain the assigned weight of the drug molecule to be tested corresponding to each chemical similarity cluster.

2. The drug performance evaluation method based on chemical spatial clustering as described in claim 1, characterized in that, The structural features are composed of skeletal features and topological features. The skeletal features are obtained by extracting the Bemis-Murcko skeletal structure of the molecule and encoding it with Morgan fingerprints. The topological features are obtained by encoding the overall structure of the molecule with Morgan circular fingerprints.

3. The drug performance evaluation method based on chemical spatial clustering as described in claim 2, characterized in that, In step S1, the chemical spatial distance is a skeleton-topological joint weighted distance, and the calculation method of the skeleton-topological joint weighted distance includes: ; ; in, These are the weighting coefficients for the global topological distance components in the joint distance metric; For calibration of concentrated molecules i With molecules j The skeleton-topology joint weighted distance between them, with a value range of [0,1]; For molecules i With molecules j The Tanimoto coefficient between Morgan's circular fingerprints; For molecules i With molecules j The Tanimoto coefficients between Bemis-Murcko skeleton features; The Spearman rank correlation coefficient between random variables X and Y; To calibrate the non-uniform fractional differences of all molecular pairs in the set.

4. The drug performance evaluation method based on chemical spatial clustering as described in claim 1, characterized in that, Step S2 specifically includes: S21. Based on the chemical spatial distance, calculate the local density of each molecule in the calibration dataset in the chemical space to obtain the local density value corresponding to each molecule; S22. Calculate the local density of each molecule in the calibration dataset based on the skeleton-topology joint weighted distance. Divide the calibration dataset into high-density regions, medium-density transition regions and low-density regions according to the local density. Use a differentiated clustering strategy for different density regions to obtain multiple chemical similarity clusters. At the same time, calculate the chemical spatial coverage radius of the calibration dataset. S23. Within each chemical similarity cluster, based on the significance level determined according to the safety level of the task to be evaluated, the quantiles of the non-consistency scores of the calibration molecules within the cluster are calculated independently, and the calibration quantiles of each chemical similarity cluster are determined. Each chemical similarity cluster and its corresponding calibration quantile together constitute a chemical spatial hierarchical calibration framework.

5. The drug performance evaluation method based on chemical spatial clustering as described in claim 4, characterized in that, In step S22, the differential clustering strategy for different density regions to obtain multiple chemical similarity clusters specifically includes: determining the optimal values ​​of the number of candidate clusters in high-density regions and the number of candidate clusters in medium-density transition regions by using the silhouette coefficient maximization strategy. The silhouette coefficient maximization strategy is as follows: in each density region, the number of candidate clusters within a preset range is clustered sequentially, the average silhouette coefficient corresponding to each number of candidate clusters is calculated, and the number of clusters corresponding to the maximum average silhouette coefficient is determined as the optimal number of clusters in that density region.

6. The drug performance evaluation method based on chemical spatial clustering as described in claim 1, characterized in that, In step S3, the prediction model is a multi-architecture graph neural network ensemble model trained using an ensemble learning strategy. The ensemble model contains multiple graph neural network sub-models with different architectures, and each sub-model is trained independently on the training set. When predicting the drug molecule to be tested, each sub-model takes the molecular graph representation obtained by converting the molecular structure information as input and independently outputs the corresponding drug performance prediction value. The mean of the prediction values ​​of each sub-model is used as the prediction result, and the standard deviation of the prediction values ​​of each sub-model is used as the uncertainty estimate.

7. The drug performance evaluation method based on chemical spatial clustering as described in claim 1, characterized in that, In step S43, the weights of each chemical similarity cluster corresponding to the drug molecule to be tested are calculated as follows: ; ; in, The soft assignment weights of the molecule to be tested to the k-th chemical similarity cluster satisfy the following conditions: and ; The denoted distance is the framework-topological joint weighted distance from the molecule to be tested to the center of the k-th cluster; K is the total number of chemically similar clusters. The local adaptive temperature parameter corresponding to the molecule to be tested; Reference temperature parameter; This is the distance sensitivity coefficient; The distance is the skeleton-topological joint weighted distance from the molecule to be tested to the nearest cluster center; To calibrate the chemical space coverage radius of the dataset.

8. The drug performance evaluation method based on chemical spatial clustering as described in claim 1, characterized in that, Step S6 specifically includes: S61. Classify the drug performance to be evaluated into safety-critical, important, and screening types according to clinical safety level. Verify whether the nominal coverage of the confidence interval generated in step S5 meets the coverage target of the safety level corresponding to the task to be evaluated. If it meets the target, directly label the confidence interval as the corresponding safety level and coverage level. If it does not meet the target, re-perform the confidence interval fusion calculation in step S5 with the significance level of the actual corresponding safety level to obtain the confidence interval after safety level correction. S62. Calculate the comprehensive confidence score from three dimensions: relative width of the corrected confidence interval, prediction uncertainty, and cluster allocation concentration. Divide the evaluation results into multiple confidence levels based on the comprehensive confidence score. When the comprehensive confidence score is lower than a preset threshold, add an extrapolation warning and a quantitative description of the chemical spatial location of the drug molecule to be tested to the output results. Output the confidence interval, prediction results, and confidence level after sensitivity adjustment as the drug performance evaluation results of the drug molecule to be tested.

9. The drug performance evaluation method based on chemical spatial clustering as described in claim 1, characterized in that, The drug performance evaluation method also includes: S7. When new known drug molecule data is obtained and the amount of new data reaches a preset update threshold, the new data is included in the calibration dataset. Based on the expanded calibration dataset, the skeleton-topology joint weighted distance between each molecule is recalculated and the weighting coefficient is updated. The local density of each molecule is recalculated and the cluster structure and chemical space coverage radius are updated by maximizing the contour coefficient. The quantiles of the expanded non-consistency scores in each chemical similarity cluster are recalculated to update the calibration quantiles of each cluster.

10. A drug performance evaluation system based on chemical spatial clustering, characterized in that, The system is used to implement the method as described in any one of claims 1 to 9, comprising: The chemical space modeling module is used to construct a chemical space clustering model based on the structural features of calibration set molecules and form a chemical space hierarchical calibration framework associated with the safety level of the task to be evaluated. The performance prediction module is used to predict drug performance based on the molecular structure information of the drug molecule to be tested, and to obtain the predicted value and uncertainty estimate. The calibration and evaluation module is used to perform statistical calibration and reliability evaluation on the predicted values ​​based on the distribution relationship of the drug molecules to be tested in the chemical space clustering model, and generate a drug performance evaluation report. The storage module is used to store the chemical spatial clustering model, the chemical spatial hierarchical calibration framework, and the drug performance evaluation report; The dynamic update module is used to trigger the update of the chemical spatial clustering model and the chemical spatial hierarchical calibration framework when new experimental data meet preset conditions, and write the update results into the storage module.

Citation Information

Patent Citations

  • Molecular property prediction method and device, storage medium and electronic equipment

    CN114842920A

  • Method for quantitatively screening new drug primer based on chemical space

    CN118506921A

  • Affinity prediction method and apparatus, method and apparatus for training affinity prediction model, device and medium

    US20220215899A1