Molecular druggability potential scoring method based on comparative learning variational auto-encoder

By constructing a molecular drugability scoring method based on contrastive learning variational autoencoders, the problem of accurate evaluation of molecular drugability in existing technologies is solved, low-dimensional modeling and continuous scoring of molecular drugability are achieved, and the efficiency and interpretability of drug research and development are improved.

CN120783901APending Publication Date: 2025-10-14EAST CHINA UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510870966.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing drug developability assessment methods are unable to accurately express the comprehensive drug properties of molecules, lack the ability to understand and model the potential structure and category information of molecules in high-dimensional space, and have problems such as strong sample dependence and weak generalization ability.

Method used

A molecular drugability potential scoring method based on contrastive learning variational autoencoder is constructed. By building a structurally separable, interpretable, and discriminative latent space, adopting a triplet contrastive learning mechanism, and combining ADMET characteristics and physicochemical properties, low-dimensional modeling and continuous scoring of molecular drugability are achieved.

Benefits of technology

It achieves low-dimensional modeling of the druggability distribution of molecules, provides a more universal, interpretable and practical scoring mechanism, is suitable for real compound screening scenarios, and reduces the risk of failure in drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783901A_ABST
    Figure CN120783901A_ABST
Patent Text Reader

Abstract

The invention relates to the field of drug development, and discloses a molecular druggability potential scoring method based on a comparative learning variational auto-encoder, which comprises the following steps: constructing a sample set comprising drug molecules and non-drug molecules, preprocessing the sample set and predicting to obtain ADMET characteristic spectrums, evaluating and sequencing the importance of the ADMET characteristic spectrums, and determining the druggability potential of the drug molecules according to the druggability potential of the drug molecules and the non-drug molecules. Screening out a druggability related characteristic set; a UniMol-based multi-task learning model is adopted to predict ADMET properties, an RDKit chemoinformatics tool is adopted to calculate physicochemical properties and synthesis feasibility scores, and a one-dimensional molecular feature vector is formed through fusion; taking the one-dimensional molecular feature vector as the input of a variational auto-encoder model adopting a fused triple contrast learning mechanism for training, and constructing a potential space; and mapping the approved drug molecules and the to-be-evaluated molecules into a potential space, and calculating druggability scores based on the Euclidean distance between the to-be-evaluated molecules and the distribution center and the local drug molecule density. The method is suitable for druggability comprehensive evaluation of drug screening, pilot optimization and molecular design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of drug development, and in particular to a method for scoring the druggability potential of molecules based on contrastive learning variational autoencoders. Background Art

[0002] During the new drug development process, experimental development often requires significant human, material, and financial resources. Therefore, identifying candidate compounds with favorable safety, efficacy, and pharmacokinetic (ADMET) properties at an early stage becomes crucial for improving drug development efficiency and reducing the risk of failure. Drug-likeness studies assess the potential of compounds to become potential drugs, providing crucial guidance for subsequent experimental development.

[0003] Existing drug developability (druggability) assessment methods mainly rely on the following two methods:

[0004] 1. Empirical rule-based scoring methods (such as quantitative estimate of drug-likeness (QED) and Lipinski rule): usually based on a weighted combination of several physicochemical properties of a molecule;

[0005] 2. Supervised learning models: These models are trained using known drug and non-drug datasets to classify or score unknown molecules. Representative work includes DBPP-Predictor (2024). Although this work integrates ADMET signatures with physicochemical properties and employs a binary classification (TCC) framework to construct a scoring mechanism (DBPP_Score), it is still essentially a supervised classification model and suffers from typical issues of "overconfidence" and "overdispersion of scores," making it difficult to reflect subtle differences between drug-like molecules.

[0006] Traditional drugability scoring methods (such as QED) rely on limited empirical rules and fixed physicochemical properties, making it difficult to accurately express a molecule's comprehensive drug properties. In particular, they lack the ability to understand and model the underlying structure and classification information of molecules in high-dimensional space. Furthermore, existing supervised learning-based scoring methods suffer from strong sample dependence and weak generalization capabilities, making them difficult to apply to real-world compound screening scenarios.

[0007] Therefore, there is still a lack of a druggability scoring method that can integrate multi-source druggability information, has discriminative ability and continuous scoring mechanism, and has good generalization and interpretability. Summary of the Invention

[0008] The purpose of this application is to provide a molecular drugability potential scoring method based on contrastive learning variational autoencoders, which realizes low-dimensional modeling of the molecular drugability distribution by constructing a structurally separable, interpretable and discriminative latent space, and calculates continuous scores through latent vectors, providing a molecular drugability evaluation mechanism with more universality, generalizability and practical guidance.

[0009] In a first aspect, the present application discloses a method for scoring the druggability potential of molecules based on contrastive learning variational autoencoders, comprising:

[0010] Construct a sample set including drug molecules and non-drug molecules, perform standardized preprocessing on each molecule, predict the preprocessed molecules and obtain ADMET signature profiles, use random forest and mutual information algorithms to evaluate and rank the importance of the ADMET signature profiles and other pseudo-labels related to drug similarity, and screen for a set of drugability-related features;

[0011] A multi-task learning model based on Uni-Mol was used to predict the ADMET properties of the pretreated molecules. The RDKit cheminformatics tool was used to calculate the physicochemical properties and synthetic feasibility scores of the pretreated molecules and fuse them to form a one-dimensional molecular feature vector for each molecule.

[0012] Training the one-dimensional molecular feature vector as input to a variational autoencoder model that uses a fused triplet contrastive learning mechanism to construct a latent space that can distinguish between drugs and non-drugs, wherein the triplet contrastive learning loss of the variational autoencoder model includes reconstruction loss, KL divergence, and triplet contrastive loss; and

[0013] Approved drug molecules and molecules to be evaluated are mapped to the latent space, and the drugability score of the molecule to be evaluated is calculated based on the Euclidean distance between the molecule to be evaluated and the distribution center of the drug molecules in the latent space and the local drug molecule density, and the drugability score result is output.

[0014] In a preferred example, the drugability-related feature set includes at least 16 ADMET properties, wherein the at least 16 ADMET properties are selected from the following group: F50, PgP inhibitor, BSEP inhibitor, plasma protein binding ratio (Plasma Protein Binding Ratio, PPB), steady state volume of distribution (Steady state volume of distribution, VDss), CYP3A4 inhibitor, CYP3A4 substrate, CYP2D6 substrate, CYP2C9 inhibitor, plasma clearance (Clplasma), acute drug liver injury (Drug-induced hepatotoxicity, DILI), FDA recommended daily dose (FDAMDD), Ames, micronucleus (Micronucleus), reproductive toxicity (Reproductive toxicity), neurotoxicity (Nephrotoxicity).

[0015] In a preferred example, the set of drugability-related features also includes at least 5 physicochemical properties and a synthetic feasibility score, and the at least 5 physicochemical properties are selected from the following group: molecular weight (MW), lipid-water partition coefficient (LogP), number of hydrogen bond donors (HBD), number of hydrogen bond acceptors (HBA) and number of rotatable bonds (nROT).

[0016] In a preferred embodiment, the standardization pre-processing includes: converting salts into corresponding acid or base forms, removing mixtures and inorganic substances, standardizing molecular representation strings, and removing duplicate molecules.

[0017] In a preferred embodiment, the Uni-Mol based multi-task learning model adopts a pre-trained Uni-Mol transformer structure to predict the ADMET properties of the pre-processed molecules through a multi-head self-attention mechanism and a task-specific classification head.

[0018] In a preferred example, the latent space is a three-dimensional space, wherein the maximum correlation features of the first dimension are: synthetic feasibility score, molecular weight, lipid-water partition coefficient, PPB, and PgP inhibitor; the maximum correlation features of the second dimension are: synthetic feasibility score, lipid-water partition coefficient, PgP inhibitor, CYP3A4 inhibitor, and neurotoxicity; and the maximum correlation features of the third dimension are: CYP3A4 inhibitor, CYP2C9 inhibitor, PgP inhibitor, molecular weight, and neurotoxicity.

[0019] In a preferred embodiment, admetSAR 3.0 and ADMETlab 3.0 are used to predict the pretreated molecules and obtain ADMET characteristic spectra.

[0020] In one preferred embodiment, the drug-likeness score is calculated by the formula norm + (1 - a) (1 - D norm ), where a is the score weight, 0 < a < 1, p norm is the normalized density score, and D norm is the normalized distance score.

[0021] In one preferred embodiment, the drug molecules include the dataset DrugBank of FDA-approved drugs, the dataset WorldDrug of other region drugs, and the non-drug molecules include the datasets ZINC, ChEMBL, and GDB17.

[0022] In one preferred embodiment, the triple contrast loss is composed of an anchor sample, a positive sample, and a negative sample, wherein the anchor and the positive sample are both drug molecules, and the negative sample is a non-drug molecule.

[0023] In a second aspect, the present application also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and characterized in that the processor implements the method when executing the computer program.

[0024] In a second aspect, the present application also discloses a non-transitory computer readable storage medium, including a computer program, and characterized in that the processor implements the steps of the method when executing the computer program.

[0025] It should be understood that, within the scope of the present application, the above-mentioned technical features of the present application and the technical features described in detail below (such as the embodiments) can be combined with each other to form new or preferred technical solutions. Due to the limited space, they will not be listed one by one here. BRIEF DESCRIPTION OF DRAWINGS

[0026] Figure 1 is a flowchart of a method for scoring the drug-likeness potential of molecules based on a contrast learning variational autoencoder according to one embodiment of the present application.

[0027] Figure 2 shows the overall workflow diagram of CLaSP in one embodiment of the present application.

[0028] Figure 3 shows the overall architecture diagram of the comprehensive prediction platform of the multi-task learning framework in one embodiment of the present application.

[0029] Figure 4 shows the three-dimensional visualization diagram of the latent space generated by the five dimension reduction methods in one embodiment of the present application.

[0030] Figure 5The performance of DBPP_Score, CLaSP_Score and QED on five datasets in one embodiment of the application is shown.

[0031] Figure 6 The correlation of different dimensions and different features of the latent space in one embodiment of the application is shown.

[0032] Figure 7 Three clinical candidate compounds and representative optimization results in one embodiment of the application are shown. DETAILED DESCRIPTION

[0033] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without these specific details and that numerous changes and modifications can be made to the embodiments described herein.

[0034] The inventors have proposed a molecular drugability potential scoring method based on contrastive learning variational autoencoder through extensive and in-depth research. The variational autoencoder is combined with triple contrastive learning to construct a structured latent space composed of physicochemical properties, synthetic feasibility scores and related ADME properties. The latent space can distinguish between drugs and non-drugs, allowing drugs and non-drugs to form a natural transition structure in a low-dimensional space, and realizing a continuous and interpretable scoring system. The present application provides a computationally efficient, interpretable and generalizable framework for drugability modeling and scoring, which has important research value and industrial application potential.

[0035] The present application has at least the following advantages:

[0036] 1. Fusion of structural features and supervised information, introduction of triple contrastive learning mechanism to guide latent space learning, so that compounds with drug properties are clustered in the space and non-drug molecules are far away, thereby improving the discriminability of the drug-like space.

[0037] 2. Define a continuous drugability scoring function on the latent space to realize seamless connection from latent vector to numerical score, overcoming the defects of traditional classification models that cannot be sorted and are prone to overfitting.

[0038] 3. Combine ADMET multi-task prediction and physicochemical property calculation to construct a unified, structured and interpretable drugability evaluation system, so that the score not only reflects the molecular structure itself, but also reflects its biological efficacy, safety and synthetic feasibility.

[0039] 4. Realize a scoring process that is stable in training, has strong generalization ability and is less dependent on labels, suitable for large-scale virtual screening, lead optimization and molecular redesign in real drug development environments.

[0040] In order to make the purposes, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0041] One aspect of the present application relates to a method for scoring the drugability potential of molecules based on a contrastive learning variational autoencoder, the flow of which is shown in Figure 1 The method comprises the following steps:

[0042] Step 101, a sample set comprising drug molecules and non-drug molecules is constructed, each molecule is subjected to standardization pretreatment, the pretreated molecules are predicted to obtain an ADMET feature map, the importance of the ADMET feature map (pseudo-label related to drug similarity) is evaluated and sorted using a random forest (RF) and mutual information (MI) algorithm, and a drugability-related feature set is screened. The pseudo-label related to drug similarity is the ADMET prediction endpoint related to drug similarity in the ADMETlab3.0 platform and the admetSAR3.0 platform, specifically, the remaining endpoints more related to the ADMET of the drug after excluding endpoints such as environmental endpoints aquatic toxicity, biological enrichment; cosmetic endpoints such as skin irritation and corrosion, etc. The prediction results of these endpoints are the pseudo-labels related to drug similarity. After screening the ADMET property features related to drugability in this step, they are combined with physicochemical property features and synthetic feasibility scores to form a drugability-related feature set.

[0043] Step 102, a multi-task learning model based on Uni-Mol is used to predict the ADMET properties of the pretreated molecules, and the RDKit chemical information tool is used to calculate the physicochemical properties and synthetic feasibility scores of the pretreated molecules, and the one-dimensional molecular feature vector of each molecule is formed by fusion.

[0044] Step 103, the one-dimensional molecular feature vector is used as the input of the variational autoencoder model adopting the fusion triple contrastive learning mechanism to train, and a latent space that can distinguish drugs from non-drugs is constructed. The triple contrastive learning loss of the variational autoencoder model includes reconstruction loss, KL divergence and triple contrastive loss.

[0045] Step 104, the approved drug molecules and the molecules to be evaluated are mapped to the latent space, the drugability score of the molecules to be evaluated is calculated based on the Euclidean distance between the molecules to be evaluated and the distribution center of the drug molecules in the latent space and the local drug molecule density, and the drugability score result is output.

[0046] In order to better understand the technical solutions of the present application, the following specific examples will be described, and the details listed in the examples are mainly for easy understanding, and do not limit the protection scope of the present application.

[0047] The application provides a triple contrast learning variational autoencoder (CLVAE) framework for drug-likeness assessment. First, features related to drug-likeness are screened from an established molecular property prediction platform, with particular emphasis on relevant ADMET property features, and these features are combined with classic physicochemical properties and expert-derived descriptors to form an enhanced prediction platform. Through semi-supervised training of the platform, a well-distributed drug-likeness latent space is constructed. Using the coordinate information of new chemical entities in the latent space, a novel and interpretable drug-likeness score is derived. Finally, we combine the drug-likeness feature prediction platform with the scoring space to establish a contrast learning guided latent space scoring platform (CLaSP), which can quickly generate a comprehensive drug-likeness assessment report for compounds. The overall workflow of CLaSP consists of four modules, as shown in Figure 2 Figure 1, wherein: (A) ADME pseudolabels are generated, and feature selection is based on RF feature importance and MI; (B) a multi-task prediction platform is constructed for estimating ADME properties, physicochemical properties, and synthetic accessibility (SA) scores from SMILES expressions; (C) a drug-likeness latent space is extracted by a contrast learning-based variational autoencoder (VAE); (D) a drug-likeness score is constructed based on the extracted latent space. The drug-likeness scoring system of the application outperforms methods such as QED on multiple benchmark datasets. In a specific case study, it successfully captures the property optimization process of Wee1 inhibitors, highlighting its potential in guiding lead compound optimization.

[0048] To support the selection and evaluation of drugability features, we collected and organized multiple datasets. FDA-approved drugs (FDA_Drug) were used as positive samples. For negative samples, three representative non-drug sources were selected to enhance the generalizability of feature extraction: the ZINC dataset; the ChEMBL database, which contains molecules with known biological activities, although not all molecules meet typical drugability criteria; and the GDB17 database, a computationally enumerated chemical space containing small organic molecules with high structural diversity but low synthetic accessibility. To further support model evaluation and generalization analysis, several external datasets were also collected: WorldDrug, a set of drugs approved by regulatory agencies outside the United States; WITHDRAWN, a set of drugs that have been withdrawn from the market; Investigation, an investigational compound from the DrugBank database; and TCMSP, a set of herbal molecules from the Traditional Chinese Medicine Systems Pharmacology database. All datasets underwent standardized preprocessing, including: (1) conversion of salts to their corresponding acids or bases; (2) removal of mixtures and inorganic compounds; and (3) standardization of SMILES strings and elimination of duplicate molecules. To ensure the robustness of the evaluation, the present invention uses different random seeds, stratifies the data in a ratio of 8:1:1, and divides it into training set, validation set, and test set three times.

[0049] The ADMET properties of molecules are widely considered to be key indicators for evaluating drugability. However, selecting the most relevant ADMET endpoints remains a major challenge in current drugability assessment. In order to determine meaningful endpoints for downstream evaluation, the present invention utilizes two state-of-the-art prediction platforms, admetSAR 3.0 and ADMETlab 3.0, to systematically evaluate the importance of ADMET-related features. The prediction platform performed comprehensive ADMET property predictions on the ZINC, ChEMBL, GDB17, and FDA_Drug datasets and obtained ADMET feature maps. The resulting ADMET feature maps, together with other pseudo-labels related to drugability, were used to evaluate and rank the importance of features through two complementary methods, random forest (RF) and mutual information (MI). The overall workflow of the feature selection process is as follows: Figure 2 Figure 1 shows a model with a multitude of features. RF is a decision tree-based ensemble learning method that estimates feature importance using metrics such as average impurity reduction and permutation importance. MI is a classic feature selection technique derived from information theory that measures dependencies between variables and can capture nonlinear relationships. RF and MI together provide a more comprehensive and robust assessment of feature relevance.

[0050] To identify drugability features, the present invention uses FDA-approved drugs as positive samples and randomly extracts 3,000 molecules from each of the ZINC, ChEMBL, and GDB17 datasets using different random seeds as negative samples. To reduce the bias of specific datasets and better reflect the real-world distribution, an additional 1,000 molecules were extracted from each dataset to construct four comparison groups: FDA vs. ZINC, FDA vs. ChEMBL, FDA vs. GDB17, and FDA vs. ALL. Each group underwent 10 rounds of random sampling, and the RF and MI methods combined with five-fold cross-validation were used to evaluate feature importance. The top 30 features in each method were ranked by average importance, and their intersection was taken as the final drugability feature set.

[0051] After identifying the most informative ADMET properties for druggability assessment, the present invention aims to develop a fast, accurate, and unified method to predict these properties of any small molecule. This will facilitate downstream tasks such as feature-based dimensionality reduction and support efficient druggability scoring. To this end, a comprehensive prediction platform based on a multi-task learning framework was constructed. Its overall architecture is as follows: Figure 3 As shown. The platform extracts drug-related features from SMILES input through two main branches. On the one hand, a trained multi-task Uni-Mol model is used to predict 16 selected relevant ADMET properties. On the other hand, the RDKit toolkit is used to calculate the physicochemical property characteristics and derive the synthetic accessibility (SA) score SA_Score through an expert model. The most commonly used SA score estimates the difficulty of synthesis by analyzing molecular complexity and fragment contribution. These combined features form the basis of downstream drugability scoring. Table 1 shows the combined features selected for drugability scoring.

[0052] Table 1 ADMET properties, physicochemical properties and SA scoring features selected for druggability scoring

[0053]

[0054] like Figure 3 As shown, to develop an AI model capable of deeply learning the relationship between molecular SMILES and druggability, the present invention uses the pre-trained molecular representation model Uni-Mol. To build a robust ADMET property prediction model, a training strategy based on Uni-Mol was followed: molecular features were processed through a feature conversion layer and a multi-head self-attention mechanism, and then connected to a task-specific classification head (task-specific classification head). This formed a multi-task prediction framework based on the pre-trained Uni-Mol model.

[0055] To benchmark the multi-task Uni-Mol model, the present invention implements three mature graph neural network (GNN) architectures as benchmark methods, in which molecules are represented as graphs with atoms as nodes and chemical bonds as edges. Graph convolutional network (GCN), graph isomorphism network (GIN), and graph attention network (GAT) are chosen as representative models, each of which utilizes different message passing mechanisms for molecular representation learning. These models include an atom and bond embedding layer, multiple message passing layers, and a graph-level pooling operation. For multi-task learning, each model simultaneously predicts 15 binary classification tasks—covering ADMET properties such as oral bioavailability (F, binarized at 50%, denoted as F50); plasma protein binding rate (PPB); various CYP interactions; and toxicity endpoints (FDAMDD, DILI, Neurotoxicity, Micronucleus, Reproductive toxicity, Ames) and one regression task, i.e., steady-state distribution volume (VDss). A shared graph embedding backbone network and task-specific prediction heads are used to capture shared and task-specific molecular features. All models are trained and evaluated using the same three random data splits as the multi-task Uni-Mol model to ensure consistency in training, validation, and testing.

[0056] To achieve effective dimensionality reduction while preserving the structural features of molecular data, the present invention proposes a contrastive learning-based variational autoencoder (CLVAE) model. This model incorporates triple contrastive learning and can learn to generate discriminative latent space representations. VAE is a probabilistic generative model that learns to map input data to latent space representations through an encoder and reconstructs the original data through a decoder. This allows high-dimensional molecular data to be compressed while retaining essential features, providing a theoretical basis for downstream analysis. The training objective is to maximize the Evidence Lower Bound (ELBO), which is defined as:

[0057]

[0058] where the first term measures reconstruction fidelity, and the second term KL divergence regularizes the latent distribution by penalizing deviation from the prior. x represents the input of the CLVAE model, and θ and φ represent the parameters of the decoder and encoder, respectively. z is the latent variable representing the hidden generative privacy, and p(z) is the prior distribution, which is generally a normal distribution; p θ (x|z) represents a generative model that generates data x from latent variable z; q φ (z|x) represents the variational distribution (similar to the posterior distribution) learned using a neural network. represents the expectation, represents the expectation of the log-likelihood under the variational distribution, is the reconstruction loss, and can also be understood as how likely it is to generate x from z under the current variational distribution.

[0059] To enhance the discriminative ability of the latent representation, contrastive learning is incorporated into the VAE framework. Contrastive learning is widely adopted in representation learning due to its ability to construct a structured latent space by modeling the similarity and dissimilarity between samples. In the present application, a triplet contrastive loss is used to construct the contrastive objective, which encodes the relative similarity by comparing one anchor sample with one positive sample and one negative sample. Specifically, the triplet contrastive loss is defined as:

[0060]

[0061] where za, zp, and zn represent the anchor (an FDA-approved drug), the positive sample (another FDA-approved drug), and the negative sample (a non-FDA-approved molecule), respectively. d(,) represents the distance metric function between two vectors, representing the distance between two points, such as the Euclidean distance described in the present embodiment. The boundary value margin enforces the minimum separation between the positive sample and the negative sample. The overall CLVAE loss function combines the reconstruction loss, the KL divergence, and the triplet contrastive loss, as follows:

[0062]

[0063] where are the reconstruction loss, the KL divergence, and the triplet contrastive loss, respectively, and λ1, λ2, λ3 are the weight coefficients that balance the contributions of each term. Through this design, the CLVAE constructs a latent space with discriminative and informative properties, which can capture molecular features and improve class separability, thereby laying a solid foundation for establishing a reliable drugability scoring system.

[0064] In this space, the center of the drug distribution can be regarded as representing the prototype features of approved drugs. Therefore, the distance of a candidate molecule to be evaluated from this center reflects its similarity to the ideal drug. In addition, the local density of surrounding drugs reflects the extent to which the candidate molecule is covered by the known drug regions, and the higher the density, the greater the similarity to existing approved compounds. Therefore, the present application calculates two complementary scores: a distance-based score representing the Euclidean distance of the molecule to the center point of the latent space drug, and a density-based score quantifying the concentration of nearby drug points. Then, by combining these two scores through an adjustable weighted sum, the final drugability score CLaSP_Score is obtained, which is calculated as follows:

[0065] CLaSP_Score=αρ norm +(1-α)(1-D norm )

[0066] Where α is the score weight, 0≤α≤1, ρ norm is the normalized density score, which is calculated as:

[0067]

[0068] The density ρ comes from a kernel density estimation (KDE) model trained on a sample of FDA-approved drugs: ρ log =KDE FDA (Z).ρ max and ρ min Indicates the maximum and minimum values ​​corresponding to ρ in all data.

[0069] Among them D norm is the normalized distance score, which is calculated as:

[0070]

[0071] Where Z represents the representation of the compound in the VAE latent space, C FDA is the center point of the FDA-approved drug in the potential space. max and D min It is the maximum and minimum distance from the center of the drug class.

[0072] To further evaluate the performance of CLaSP_Score on real-world compounds, we analyzed its behavior on a set of 1,751 investigational compounds and 266 withdrawn drugs collected from the DrugBank research group. The score distribution of each dataset is shown in Figure 2. Figure 4 As shown in the figure, (A) and (B) show the latent spaces generated by CLVAE and standard VAE, respectively, on the left side of the figure; (C), (D) and (E) show the dimensionality reduction results of PCA, t-SNE and UMAP, on the right side. Each point represents a compound, and the color is distinguished according to the data source (FDA, ChEMBL, ZINC or GDB17). CLVAE shows a more structured and drug-aware spatial distribution, forming a smooth transition from non-drug molecules to drug-like molecules. In contrast, PCA failed to achieve effective distinction, and although t-SNE and UMAP showed a certain clustering effect, they failed to clearly align drug properties. Standard VAE is significantly weaker than CLVAE in category separation ability.

[0073] Comparing the latent space representations of CLVAE with PCA, t-SNE, UMAP, and standard VAE in Table 2, CLVAE constructed clear transition structures between drugs and non-drugs, and was optimal in the following indicators:

[0074] Table 2 Comparison of latent space quality of different model methods

[0075] method ARI↑ Silhouette Score↑ Davies Bouldin Index↓ PCA 0.1718 0.0366 1.1201 TSNE 0.1425 0.0456 1.1394 UMAP 0.2390 0.0700 1.0728 VAE 0.2640 0.0219 1.3355 CLVAE 0.6105 0.3907 0.5835

[0076] To study the sample dependency of different scoring methods, the present application analyzes the performance of DBPP_Score, CLaSP_Score and QED on five data sets, as shown in Figure 5 The score distribution of DBPP_Score, CLaSP_Score and QED on five data sets (FDA, ZINC, Worlddrug, ChEMBL and GDB17) is shown in A-C; the distribution of the three scores in real-world compound classification, including Drugs (containing FDA and Worlddrug), Investigation, Withdrawn and ZINC, is shown in D-F. The drugability score CLaSP_Score can still be reasonably scored on data sets not involved in training (such as WorldDrug, Withdrawn, TCMSP), and has strong generalization ability. Compared with traditional QED and DBPP_Score, CLaSP significantly improves the accuracy and credibility of real drug evaluation.

[0077] Specifically, DBPP_Score shows strong discrimination between FDA and ZINC data sets, and the scores tend to be polarized. This behavior can be attributed to its training settings: DBPP_Score is a supervised machine learning model that uses FDA compounds as positive samples and ZINC compounds as negative samples for training. Although this design enhances classification performance, it inevitably introduces sample-specific bias and overfitting. In contrast, QED scores are more evenly distributed across all data sets, with most molecules having similar score ranges. QED seems to favor molecules from the ZINC database - this is contrary to the results of DBPP_Score, which may be due to the "molecule inflation" effect caused by long-term reliance on QED-based screening, leading to data set management practices that tend to enrich molecules that meet QED optimization features.

[0078] Regarding the CLaSP_Score, it's noteworthy that only FDA compounds were used during the comparative training phase of CLVAE. Despite this, the score achieves high scores on the WorldDrug dataset, which was completely unseen during training. This suggests that the CLaSP_Score has good generalization capabilities. One possible explanation is that many WorldDrug molecules are "me-too" drugs developed with reference to FDA-approved compounds, often with similar or improved ADMET and overall drugability characteristics, which are effectively captured by the CLaSP scoring framework.

[0079] To investigate how druggability features are encoded in the latent space learned by CLVAE, we performed Pearson correlation analysis between key molecular descriptors and each of the three latent dimensions. Figure 6 As shown, the results indicate that each latent dimension captures different aspects of molecular properties: (1) Latent dimension z1 is mainly related to molecular complexity and lipophilicity, such as high molecular weight and synthetic accessibility (SA_Score); (2) Latent dimension z2 reflects hydrophilicity and protein binding properties, especially the number of hydrogen bond donors (HBD); (3) Latent dimension z3 is mainly related to drug safety and metabolism-related characteristics, including CYP enzyme substrate status and neurotoxicity indicators. Among them, SA_Score showed significant correlation with all three latent dimensions. Although affinity to biological targets and related ADMET properties are generally considered to be key to drugability, this finding highlights that synthetic accessibility also plays a crucial role in determining the developability of compounds. This observation supports the inclusion of synthetic accessibility as a meaningful component in the comprehensive evaluation of drugability.

[0080] The ClaSP framework of the present invention is evaluated below by applying the CLaSP score to candidate tumor drugs.

[0081] The cell cycle undergoes stringent checkpoint regulation at key stages, such as S phase, G1 / S, and G2 / M, to ensure that cells complete DNA damage repair before mitosis. WEE1 kinase is a key regulator of the G2 / M checkpoint, and the drug candidate AZD1775 exerts its anti-tumor effects by inhibiting WEE1. Based on this structure, Joel L. Syphers et al. developed a series of more active WEE1 inhibitors, representative compounds of which outperformed AZD1775 in both selectivity and activity, and demonstrated promising efficacy in colorectal cancer patient-derived organoids (PDOs).

[0082] The method of the present invention performed CLaSP scoring on the 31 synthetic compounds and 3 clinical candidate drugs in this study. Figure 7Three clinical candidate compounds and representative optimization results were presented. ZN-c3 (compound 2) was reported to exhibit hematologic toxicity, such as myelosuppression, in early clinical trials, corresponding to a relatively low CLaSP score. SC0191 (compound 3) demonstrated excellent activity and no significant toxicity in preclinical studies, also resulting in a high CLaSP score.

[0083] In addition, all compounds in this series were identified as positive by the CLaSP property predictor in the micronucleus assay, which is consistent with their mechanism of action of inducing DNA damage; the CLaSP score can also well reflect the differences in WEE1 inhibitory activity. As the inhibitory activity increases (such as compounds 32 to 34), the score gradually increases.

[0084] In summary, the scoring method of the present invention can conduct a comprehensive evaluation of the drugability of candidate drugs at an early stage, and provide effective guidance for the screening and optimization of lead compounds.

[0085] On the other hand, one embodiment of the present application further provides a molecular druggability potential scoring system, comprising:

[0086] An input module is used to receive and normalize the structure of preprocessed input molecules, such as normalizing them into SMILES format;

[0087] A feature selection module is used to call admetSAR 3.0 and ADMETlab 3.0 to predict ADMET signature profiles. It uses random forest and mutual information algorithms to evaluate and rank the importance of ADMET signature profiles and other pseudo-labels related to drug similarity, thereby screening out a set of drug-related features.

[0088] Feature prediction module, used to generate one-dimensional molecular feature vectors of input molecules through the multi-task Uni-Mol model and RDKit;

[0089] A variational autoencoder module that integrates triplet contrastive learning to construct a latent space that can distinguish between drugs and non-drugs;

[0090] The scoring calculation module is used to calculate the drugability score of the input molecule based on the Euclidean distance between the input molecule and the distribution center of the drug molecules in the latent space and the local drug molecule density and output the drugability score result.

[0091] The scoring method and system proposed in the present invention have good industrial applicability and can be used in commercial drug discovery processes, reducing development costs and shortening the R&D cycle.

[0092] In another aspect, the present invention further provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus. The processor may invoke logic instructions in the memory to execute a method for scoring druggability potential using a variational autoencoder based on contrastive learning.

[0093] In addition, the logical instructions in the above-mentioned memory can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disk.

[0094] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the molecular drugability potential scoring method based on contrastive learning variational autoencoder provided by the above methods.

[0095] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the above-mentioned molecular drugability potential scoring method based on contrastive learning variational autoencoder.

[0096] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0097] The specification of this application records a large number of technical features, which are distributed in various technical solutions. If all possible combinations of technical features of this application (i.e., technical solutions) are to be listed, the specification will be too lengthy. In order to avoid this problem, the various technical features disclosed in the above-mentioned invention content of this application, the various technical features disclosed in the various embodiments and examples below, and the various technical features disclosed in the accompanying drawings can be freely combined with each other to form various new technical solutions (these technical solutions are all deemed to have been recorded in this specification), unless such a combination of technical features is technically infeasible. For example, in one example, feature A+B+C is disclosed, and in another example, feature A+B+D+E is disclosed. Features C and D are equivalent technical means that play the same role. Technically, only one of them can be used, and it is impossible to use them at the same time. Feature E can be technically combined with feature C. Then, the solution of A+B+C+D should not be considered as having been recorded because it is technically infeasible, while the solution of A+B+C+E should be considered as having been recorded.

[0098] All documents mentioned in this application are considered to be included in their entirety in the disclosure of this application so that they can be used as a basis for modification when necessary. In addition, it should be understood that after reading the above disclosure of this application, those skilled in the art may make various changes or modifications to this application, and these equivalent forms also fall within the scope of protection claimed in this application.

Claims

1. A molecular drugability potential scoring method based on contrastive learning variational autoencoder, characterized in that: include: Construct a sample set including drug molecules and non-drug molecules, perform standardized preprocessing on each molecule, predict the preprocessed molecules and obtain ADMET characteristic profiles, use random forest and mutual information algorithms to evaluate and rank the importance of the ADMET characteristic profiles, and screen out a set of drugability-related features; A multi-task learning model based on Uni-Mol was used to predict the ADMET properties of the pretreated molecules. The RDKit cheminformatics tool was used to calculate the physicochemical properties and synthetic feasibility scores of the pretreated molecules and fuse them to form a one-dimensional molecular feature vector for each molecule. The one-dimensional molecular feature vector is used as input for training a variational autoencoder model that adopts a fusion triplet contrastive learning mechanism, and a latent space that can distinguish between drugs and non-drugs is constructed. The triplet contrastive learning loss of the variational autoencoder model includes reconstruction loss, KL divergence, and triplet contrastive loss; and Approved drug molecules and molecules to be evaluated are mapped to the latent space, and the drugability score of the molecule to be evaluated is calculated based on the Euclidean distance between the molecule to be evaluated and the distribution center of the drug molecules in the latent space and the local drug molecule density, and the drugability score result is output.

2. The method according to claim 1, wherein The set of drugability-related features includes at least 16 ADMET properties, wherein the at least 16 ADMET properties are selected from the following groups: F50, PgP inhibitor, BSEP inhibitor, PPB, VDss, CYP3A4 inhibitor, CYP3A4 substrate, CYP2D6 substrate, CYP2C9 inhibitor, Clplasma, DILI, FDAMDD, Ames, Micronucleus, reproductive toxicity, neurotoxicity; the set of drugability-related features includes: at least 5 physicochemical properties and synthetic feasibility scores, wherein the at least 5 physicochemical properties are selected from the following group: molecular weight, lipid-water partition coefficient, number of hydrogen bond donors, number of hydrogen bond acceptors and number of rotatable bonds.

3. The method according to claim 1, wherein The normalization pre-processing includes converting salts into their corresponding acid or base forms, removing mixtures and inorganic substances, normalizing molecular representation strings, and removing duplicate molecules.

4. The method according to claim 1, wherein The Uni-Mol-based multi-task learning model adopts a pre-trained Uni-Mol transformer structure to predict the ADMET properties of the pre-processed molecules through a multi-head self-attention mechanism and a task-specific classification head.

5. The method according to claim 1, wherein The latent space is a three-dimensional space, in which the maximum correlation features of the first dimension are: synthetic feasibility score, molecular weight, lipid-water partition coefficient, PPB, and PgP inhibitor; the maximum correlation features of the second dimension are: synthetic feasibility score, lipid-water partition coefficient, PgP inhibitor, CYP3A4 inhibitor, and neurotoxicity; the maximum correlation features of the third dimension are: CYP3A4 inhibitor, CYP2C9 inhibitor, PgP inhibitor, molecular weight, and neurotoxicity.

6. The method according to claim 1, wherein admetSAR 3.0 and ADMETlab 3.0 were used to predict the pretreated molecules and obtain the ADMET characteristic spectrum.

7. The method according to claim 1, wherein The drugability score is calculated by the formula αρ norm +(1-α)(1-D norm ) calculation, where α is the score weight, 0≤α≤1, ρ norm is the normalized density score, D norm is the normalized distance score.

8. The method according to claim 1, wherein The triplet samples in the triplet contrast loss consist of anchor samples, positive samples and negative samples, where the anchor samples and positive samples are both drug molecules, and the negative samples are non-drug molecules.

9. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the computer program.

10. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium comprises a computer program, which implements the steps of the method according to any one of claims 1 to 8 when executed by a processor.