Small molecule virtual screening method and device based on information bottleneck training
The small molecule virtual screening method trained by information bottlenecks solves the accuracy and robustness problems of existing small molecule virtual screening technologies by using comparative learning and information compression of drug small molecule pre-trained models and transcriptome change encoders, and achieves effective screening for complex diseases.
Patent Information
- Application Number
- CN202511043974.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies have limited effectiveness in virtual screening of small molecule drugs for complex diseases such as Alzheimer's disease and type 2 diabetes, and perform poorly in cross-cell line drug screening, making it difficult to achieve accurate and robust drug screening.
A small molecule virtual screening method based on information bottleneck training is adopted. By acquiring a training dataset, a fusion vector representation is generated using a drug small molecule pre-trained model and an initial unperturbed encoder. The transcriptome change encoder is then combined with contrastive loss training. An information bottleneck loss function is introduced to optimize the drug small molecule encoder and the transcriptome change encoder, thereby achieving contrastive learning and information compression.
It enables precise and robust screening of small drug molecules, accurately predicts molecular effects in new cell lines, implicitly learns drug-target interaction patterns, and avoids overfitting to inherent cell line properties and experimental noise.
Smart Images

Figure CN120998347A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and apparatus for virtual screening of small molecules based on information bottleneck training. Background Technology
[0002] In recent years, target-based virtual screening technology has achieved significant results in drug development by identifying compounds that bind to disease-related targets. However, this method is only applicable to diseases with clearly defined targets, and its screening effect is very limited for complex diseases such as neurodegenerative diseases (e.g., Alzheimer's disease) or multifactorial metabolic syndromes (e.g., type 2 diabetes).
[0003] With the rapid development of systems biology and omics technologies, phenotype-based virtual screening has gradually become a research hotspot. This strategy does not rely on pre-defined target information but instead screens candidate compounds by assessing the ability of molecules to reverse disease-related phenotypes at the cellular or tissue level. Statistical data shows that between 1999 and 2008, drugs using phenotype-based strategies received more "first-in-class" drug approvals than those using target-based strategies, highlighting their potential for clinical translation. However, existing phenotype-based virtual screening technologies perform poorly in cross-cell line drug screening, reflecting their limitations in modeling molecular perturbation effects.
[0004] Therefore, how to conduct precise and robust virtual screening of small drug molecules is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] This invention provides a method and apparatus for virtual screening of small molecules based on information bottleneck training, which addresses the above-mentioned deficiencies in the prior art and enables accurate and robust virtual screening of small drug molecules.
[0006] This invention provides a small molecule virtual screening method based on information bottleneck training, comprising the following steps.
[0007] A training dataset is obtained, comprising: identifiers of drug small molecule samples; perturbed transcriptome data samples obtained by perturbing the cell population with the drug small molecule samples; and unperturbed transcriptome data samples of the cell population. Using a pre-trained drug small molecule model and an initial unperturbed encoder, a fused vector representation of the drug small molecule samples and the unperturbed transcriptome data samples is obtained based on their identifiers and the unperturbed transcriptome data samples. Using an initial transcriptome change encoder, a predicted vector of transcriptome change samples is obtained based on the transcriptome change samples between the perturbed transcriptome data samples and their corresponding unperturbed transcriptome data samples. The fusion vector representation is paired with the predicted vector representation of the corresponding transcriptome change sample to form a positive sample pair, and the fusion vector representation is paired with the predicted vector representation of the non-corresponding transcriptome change sample to form a negative sample pair. Based on the contrastive loss function and the information bottleneck loss function, the drug small molecule pre-trained model, the initial transcriptome change encoder, and the initial unperturbed encoder are trained using contrastive loss, so as to use the trained drug small molecule encoder, transcriptome change encoder, and unperturbed encoder to perform the task of screening drug small molecules; wherein, the information bottleneck loss function is used to minimize the KL divergence between the distribution of the fusion vector representation and the standard Gaussian distribution.
[0008] According to the present invention, a small molecule virtual screening method based on information bottleneck training is provided. The method utilizes a drug small molecule pre-trained model and an initial unperturbed encoder to obtain a fused vector representation of the drug small molecule sample and the unperturbed transcriptome data sample based on the drug small molecule sample identifier and the unperturbed transcriptome data sample. The method includes: using the drug small molecule pre-trained model to obtain a predicted vector representation of the drug small molecule sample based on the drug small molecule sample identifier; using the initial unperturbed encoder to obtain a predicted vector representation of the unperturbed transcriptome data sample based on the unperturbed transcriptome data sample; and using a variational encoder to fuse the predicted vector representations of the drug small molecule sample and the unperturbed transcriptome data sample to obtain the fused vector representation of the drug small molecule sample and the unperturbed transcriptome data sample.
[0009] According to the present invention, a small molecule virtual screening method based on information bottleneck training is provided. The method involves fusing the predicted vector representation of the drug small molecule sample and the predicted vector representation of the unperturbed transcriptome data sample using a variational encoder to obtain a fused vector representation of the drug small molecule sample and the unperturbed transcriptome data sample. This includes: concatenating the predicted vector representation of the drug small molecule sample and the predicted vector representation of the unperturbed transcriptome data sample to obtain a predicted concatenated vector; and using the variational encoder to perform latent space sampling on the predicted concatenated vector to obtain the fused vector representation.
[0010] According to the present invention, a small molecule virtual screening method based on information bottleneck training is provided, wherein in the contrast loss training process, each batch of sample data comes from the same cell line.
[0011] According to the present invention, a small molecule virtual screening method based on information bottleneck training is provided, wherein the contrastive loss function is as follows: ; in, N This represents the total number of positive and negative sample pairs. k It is a positive integer; The first contrastive loss function is used to train and maximize the fused vector representation of sample pairs k. Its positive samples The probability of a match is calculated using the following formula: ; in, For the first k The fusion vector representation of a pair of positive samples. For the first k The predicted vector representation of transcriptome change samples for each positive sample pair; Temperature coefficient; For the first i The predicted vector representation of transcriptome change samples for each sample pair; in, The second contrastive loss function maximizes the sample pairs through contrastive training. k Vector representation of transcriptome change samples Its positive samples The probability of a match is expressed by the following formula: in, For the first i The fusion vector representation of a sample pair.
[0012] According to the present invention, a small molecule virtual screening method based on information bottleneck training is provided, wherein the information bottleneck loss function is as follows: in, N The number of sample pairs participating in training; The dimension of the final fused vector representation. and The first The fusion vector representation of the nth sample pair Mean and standard deviation of each dimension; It follows a standard Gaussian distribution; The KL divergence between the final fusion vector representation and the standard Gaussian distribution.
[0013] This invention also provides a small molecule virtual screening device based on information bottleneck training, comprising the following modules: The first acquisition module is used to acquire a training dataset, wherein the training dataset includes the identifiers of drug small molecule samples, perturbed transcriptome data samples obtained by perturbing the cell population with the drug small molecule samples, and unperturbed transcriptome data samples of the cell population; the second acquisition module is used to obtain a fused vector representation of the drug small molecule samples and the unperturbed transcriptome data samples based on the identifiers of the drug small molecule samples and the unperturbed transcriptome data samples using a drug small molecule pre-trained model and an initial unperturbed encoder; the third acquisition module is used to obtain the transcriptome data based on the amount of transcriptome change between the perturbed transcriptome data samples and their corresponding unperturbed transcriptome data samples using an initial transcriptome change encoder. The training module is used to form positive sample pairs by combining the fused vector representation with the corresponding predicted vector representation of the transcriptome change sample, and negative sample pairs by combining the fused vector representation with the predicted vector representation of non-corresponding transcriptome change samples. Based on the contrastive loss function and the information bottleneck loss function, the pre-trained drug molecule model, the initial transcriptome change encoder, and the initial unperturbed encoder are trained using contrastive loss to perform the task of screening drug small molecules. The information bottleneck loss function is used to minimize the KL divergence between the distribution of the fused vector representation and the standard Gaussian distribution.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the small molecule virtual screening method based on information bottleneck training as described above.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the small molecule virtual screening method based on information bottleneck training as described above.
[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the small molecule virtual screening method based on information bottleneck training as described above.
[0017] This invention provides a small molecule virtual screening method and apparatus based on information bottleneck training. Utilizing a drug small molecule pre-trained model and an initial unperturbed encoder, a fused vector representation of the drug small molecule sample and the unperturbed transcriptome data sample is obtained based on the drug small molecule sample identifier and the unperturbed transcriptome data sample. Using the initial transcriptome change encoder, a predicted vector representation of the transcriptome change sample is obtained based on the transcriptome change sample between the perturbed transcriptome data sample and its corresponding unperturbed transcriptome data sample. The fused vector representation and its corresponding predicted vector representation of the transcriptome change sample are paired to form positive sample pairs, and the fused vector representation and the predicted vector representation of non-corresponding transcriptome change samples are paired to form negative sample pairs. Based on a contrastive loss function and an information bottleneck loss function, the drug small molecule pre-trained model, the initial transcriptome change encoder, and the initial unperturbed encoder are jointly trained. Through contrastive learning between the fused vector representation and the transcriptome change vector representation, the fused vector representation not only contains chemical structural information but also implicitly encodes its perturbation pattern on the cellular transcriptome. Furthermore, by introducing an information compression constraint through an information bottleneck loss function, the model participating in training is forced to learn a compact molecular representation, avoiding overfitting to inherent cell line properties or experimental noise. Therefore, the drug small molecule encoder, transcriptome change encoder, and unperturbed encoder trained using the scheme provided in this invention can accurately and robustly perform virtual screening of drug small molecules. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is a schematic flowchart of the training methods for the drug small molecule encoder, transcriptome change encoder, and undisturbed encoder provided by the present invention.
[0020] Figure 2 This is a flowchart illustrating the small molecule virtual screening method based on information bottleneck training provided by the present invention.
[0021] Figure 3 This is a schematic diagram of the training process of the drug small molecule encoder provided by the present invention.
[0022] Figure 4 This is a schematic diagram of the structure of the small molecule virtual screening device based on information bottleneck training provided by the present invention.
[0023] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0025] The following is combined with Figures 1-3 This invention describes a small molecule virtual screening method based on information bottleneck training.
[0026] Figure 1 This is a flowchart illustrating the training methods for the drug small molecule encoder, transcriptome change encoder, and unperturbed encoder provided by the present invention, as shown below. Figure 1 As shown, the method includes the following steps: Step 101: Obtain the training dataset, which includes the identifiers of drug small molecule samples, perturbed transcriptome data samples obtained by perturbing the cell population with drug small molecule samples, and unperturbed transcriptome data samples of the cell population.
[0027] Identifiers for small molecule drug samples, such as the string "SMILES" or a molecular diagram, are used to uniquely identify small molecule drugs.
[0028] The perturbation transcriptome data samples are gene expression profiles of cell populations after treatment with small molecule drug samples (such as RNA-seq data or microarray data).
[0029] The unperturbed transcriptome data samples are the gene expression profiles of the cell populations corresponding to the perturbed transcriptome data samples without drug treatment.
[0030] In the specific implementation process, the identification of drug small molecules, the transcriptome data after drug perturbation, and the unperturbed transcriptome data with the same cell line and the same batch information can be obtained from the high-throughput drug perturbation dataset to obtain the training dataset.
[0031] To avoid interference from the inherent properties of the cell line with the learning perturbation effect, each batch of sample data comes from the same cell line during the contrastive loss training process.
[0032] Step 102: Using the drug small molecule pre-trained model and the initial unperturbed encoder, based on the drug small molecule sample identifier and the unperturbed transcriptome data sample, obtain the fused vector representation of the drug small molecule sample and the unperturbed transcriptome data sample.
[0033] In the specific implementation process, a pre-trained model for drug small molecules can be used to obtain the predicted vector representation of the drug small molecule sample based on the identifier of the drug small molecule sample.
[0034] Drug small molecule pre-trained models are deep neural networks based on self-supervised learning, used to convert drug small molecule identifiers (such as SMILES strings and molecular graph structures) into low-dimensional dense molecular representation vectors. This model learns general molecular features through pre-training tasks without relying on labeled data, thus capturing molecular properties (such as bond lengths, angles, and charge distribution) across physical and biochemical scales. Drug small molecule pre-trained models can be built based on Transformers.
[0035] Pre-training tasks include: The 3D Position Recovery task is used to train models to learn the geometric conformations of molecules (such as bond angles and dihedral angles) and capture the effect of spatial arrangement on activity.
[0036] The Atom Masking Prediction task is used to train a model to randomly mask some atoms in a molecule (e.g., by labeling them with a mask) and predict the type of the masked atoms (e.g., carbon, oxygen, nitrogen) in order to learn the local chemical environment of the molecule (e.g., functional groups, chemical bond types).
[0037] In practice, the initial unperturbed encoder can be used to obtain the predicted vector representation of the unperturbed transcriptome data samples based on the unperturbed transcriptome data samples.
[0038] The initial unperturbed encoder is an unperturbed encoder that has not been trained with contrastive loss.
[0039] An unperturbed encoder is a trained machine learning model used to map unperturbed transcriptome data (such as gene expression matrices) into a fixed-dimensional vector representation. This vector contains key biological characteristics of cells in their basal state (such as gene regulatory patterns and cell type specificity) to allow for comparative analysis with transcriptome changes after drug perturbation, thereby assessing the perturbation effect of molecules on cells.
[0040] In practice, unperturbed encoders can be constructed based on various machine learning models, such as the Transformer, and are not limited to the descriptions in this specification. The input to the unperturbed transcriptome data is the unperturbed transcriptome data, and the output is a vector representation of the unperturbed transcriptome data.
[0041] In the specific implementation process, a variational encoder is used to fuse the predicted vector representations of small drug molecule samples and unperturbed transcriptome data samples to obtain a fused vector representation of the small drug molecule samples and the unperturbed transcriptome data samples. Specifically, the predicted vector representations of small drug molecule samples and unperturbed transcriptome data samples can be concatenated to obtain a concatenated vector; the variational encoder is then used to sample the latent space of the concatenated vector to obtain the fused vector representation.
[0042] A variational encoder is a type of generative neural network, a variant or simplified form of a variational autoencoder (VAE). Its core objective is to learn the probability distribution of input data in the latent space and generate more robust vector representations through resampling. The variational encoder outputs the mean and variance of the latent distribution and introduces noise through random sampling, which can effectively enhance the model's generalization ability.
[0043] Step 103: Using the initial transcriptome change encoder, obtain the predicted vector representation of the transcriptome change sample based on the transcriptome change sample between the perturbed transcriptome data sample and its corresponding unperturbed transcriptome data sample.
[0044] The initial transcriptome change encoder is a transcriptome change encoder that has not been trained with contrastive loss.
[0045] The transcriptome change encoder is a trained machine learning model used to map high-dimensional transcriptome changes into low-dimensional vector representations.
[0046] In practice, transcriptome change encoders can be constructed based on various machine learning models, such as Transformer, without being limited by the description in this specification.
[0047] The input data for the transcriptome change encoder is a vector of transcriptome changes, such as gene expression changes calculated using the interpolation or logarithmic ratio method. The output data for the transcriptome change encoder is a vector representation of the transcriptome changes.
[0048] Step 104: Form positive sample pairs by combining the fusion vector representation with the predicted vector representation of the corresponding transcriptome change sample, and form negative sample pairs by combining the fusion vector representation with the predicted vector representation of the non-corresponding transcriptome change sample. Based on the contrastive loss function and the information bottleneck loss function, perform contrastive loss training on the drug small molecule pre-training model, the initial transcriptome change encoder, and the initial unperturbed encoder, so as to use the trained drug small molecule encoder, transcriptome change encoder, and unperturbed encoder to perform the task of screening drug small molecules.
[0049] Using contrastive loss training can increase the cosine similarity between vector representations of positive sample pairs and decrease the cosine similarity between vector representations of negative sample pairs, so that the final fused vector representation determined by specific drug small molecules and specific cellular background can be consistent with the vector representation of the corresponding transcriptome changes.
[0050] In some embodiments, the comparison loss function As shown below: ; (1) in, N This represents the total number of positive and negative sample pairs. k The ratio is a positive integer. Loss function. By bidirectional contrastive loss function and get.
[0051] The first contrastive loss function is used to train and maximize the fused vector representation of sample pairs k. Its positive samples The probability of a match is calculated using the following formula: ; (2) in, For the first k The fusion vector representation of a pair of positive samples. For the first k The predicted vector representation of transcriptome change samples for each positive sample pair; This is a temperature coefficient used to control the sharpness of the distribution in the softmax function; For the first i The predicted vector representation of transcriptome change samples for each sample pair.
[0052] The second contrastive loss function maximizes the sample pairs through contrastive loss training. k Vector representation of transcriptome change samples Its positive samples The probability of a match is expressed by the following formula: ; (3) in, For the first i The fusion vector representation of a sample pair.
[0053] The information bottleneck loss function is used to minimize the KL divergence between the distribution of the fused vector representation and the standard Gaussian distribution.
[0054] In some embodiments, the information bottleneck loss function is as follows: ; (4) in, N The number of sample pairs participating in training; The dimension of the final fused vector representation. and The first The fusion vector representation of the nth sample pair Mean and standard deviation of each dimension; It follows a standard Gaussian distribution; The KL divergence between the final fusion vector representation and the standard Gaussian distribution.
[0055] In the embodiments provided by the present invention, an information compression constraint is introduced through an information bottleneck loss function, so that the model participating in training retains only the effective information related to the prediction target and filters out irrelevant perturbation factors.
[0056] In the specific implementation process, the joint loss function can be established according to formulas (1) to (4); the parameters of the small molecule pre-trained model, the initial unperturbed encoder and the initial transcriptome change encoder can be updated using a preset optimization algorithm (e.g., gradient descent optimization algorithm) to minimize the value of the joint loss function and directly reach the training termination condition (e.g., all models converge or reach the preset number of training times) to obtain the trained drug small molecule encoder, transcriptome change encoder and unperturbed encoder.
[0057] In the embodiments provided by this invention, a contrastive learning strategy is used to optimize the pre-trained model for small molecules, so that its output vector representation satisfies the following characteristics: Implicit modeling mechanism: Through comparative learning, the vector representation of drug small molecules output by the drug small molecule encoder not only contains chemical structure information, but also implicitly encodes its perturbation pattern on the cellular transcriptome (such as target binding and signaling pathway activation).
[0058] Cross-cell line generalization ability: Due to the cell line-based sampling strategy employed during training (avoiding interference from inherent cell line properties) and information bottleneck loss (filtering out irrelevant noise), the vector representation of the drug small molecule output by the drug small molecule encoder can accurately predict molecular effects in new cell lines. This vector representation can also align itself with the protein target vector representation, thereby implicitly learning drug-target interaction patterns.
[0059] Figure 2 This is a flowchart illustrating the small molecule virtual screening method based on information bottleneck training provided by the present invention, as shown below. Figure 2 As shown, the method includes the following steps: Step 201: Obtain the identifiers and unperturbed transcriptome data of the small molecules to be screened.
[0060] Small molecules to be screened are chemical molecules with therapeutic potential but whose activity has not yet been verified. Virtual screening technology is used to assess their impact on cell phenotype. For example, small molecules to be screened could be candidate compounds for Alzheimer's disease (such as β-secretase inhibitors) or potential small molecules for type 2 diabetes (such as PPARγ agonists).
[0061] The identifier for the small molecule drug to be screened is a unique identifier for that small molecule. For example, the identifier for a small molecule drug can be the string SMILES (such as "O=C(O)C1=CC=CC=C1" representing aspirin), a molecular diagram structure, etc.
[0062] Unperturbed transcriptome data consists of gene expression data from disease-associated cell lines without drug intervention, serving as a benchmark for assessing the effects of drug perturbation. For example, unperturbed transcriptome data could be the transcriptome sequencing data of a disease-associated cell line (HepG2) under drug-free conditions.
[0063] Step 202: Using a drug small molecule encoder, obtain the vector representation of the drug small molecule to be screened based on its identifier.
[0064] The drug small molecule encoder is a deep learning model that maps the identifiers of drug small molecules to be screened into low-dimensional vector representations through self-supervised pre-training tasks and contrastive learning fine-tuning. Its core objective is to capture the three-dimensional structural characteristics, physicochemical properties, and potential mechanisms of action of drug small molecules, thereby supporting subsequent modeling of transcriptome perturbation effects.
[0065] In practice, drug small molecule encoders can be constructed using various models, such as 3D-Transformer, without being limited by the description in this specification.
[0066] The input data for the drug small molecule encoder is the identifier of the drug small molecule to be screened.
[0067] Generate a fixed-dimensional vector representation (e.g., 512-dimensional) as the final encoding of the drug small molecule to be screened, and output the vector representation of the drug small molecule.
[0068] The vector representation of drug small molecules encodes the following information about the drug small molecules: three-dimensional structural characteristics, such as atomic spatial arrangement, bond length, and bond angle (learned through a three-dimensional position recovery task); physicochemical properties, such as hydrophobicity and charge distribution (learned through an atom masking prediction task); and knowledge of the mechanism of action, such as potential interaction patterns with protein targets (learned through comparative learning of aligned protein target representations).
[0069] Step 203: Using the unperturbed encoder, obtain the vector representation of the unperturbed transcriptome data based on the unperturbed transcriptome data.
[0070] For a detailed description of the unperturbed encoder, see [link to documentation]. Figure 1 The relevant content will not be repeated here.
[0071] Step 204: Using a variational encoder, the vector representation of the small molecule drug to be screened and the vector representation of the unperturbed transcriptome data are fused to obtain a fused vector representation.
[0072] For a detailed description of the variational encoder, see [link to documentation]. Figure 1 The relevant content will not be repeated here.
[0073] In some embodiments, the vector representations of the small molecule drug to be screened and the vector representations of the unperturbed transcriptome data can be concatenated to obtain a concatenated vector; a variational encoder is then used to sample the latent space of the concatenated vector to obtain a fused vector representation.
[0074] In practice, the vector representations of the small molecule drug to be screened and the vector representations of the undisturbed transcriptome data can be concatenated along the feature dimension to form an initial fusion vector, which serves as the input data for the variational encoder. The variational encoder outputs the mean (μ) and variance (σ²) vectors of the latent distribution. Latent distribution: Assuming the latent space follows a diagonal Gaussian distribution, a vector can be randomly sampled from the latent distribution as the final fusion vector representation.
[0075] The efficacy of small drug molecules depends not only on their own structure but also on inherent cellular properties such as cell type and metabolic state. In the embodiments provided by this invention, a variational encoder is used to fuse the vector representation of the small drug molecule to be screened with the vector representation of the undisturbed transcriptome data, which can effectively and implicitly capture molecular-cell interactions (such as target binding and metabolic regulation).
[0076] Step 205: Based on the fusion vector representation, perform drug screening on the small molecule drug to be screened.
[0077] In practice, various drug screening tasks can be performed based on the fusion vector representation of the small molecule drug to be screened, without being limited by the description in this manual.
[0078] In some embodiments, target perturbation transcriptome data is acquired; the amount of transcriptome change between the target perturbation transcriptome data and the unperturbed transcriptome data is determined; a transcriptome change encoder is used to obtain a vector representation of the amount of transcriptome change based on the amount of transcriptome change; the similarity between the fusion vector representation and the vector representation of the amount of transcriptome change is determined; and drug screening is performed on the small molecule drug to be screened based on the similarity.
[0079] Post-target perturbation transcriptome data refers to the data on changes in gene expression levels in cells after treatment with small molecule drugs. It is usually expressed as a matrix of gene expression levels (such as FPKM values of RNA-seq, fluorescence intensity of microarrays) or transcript abundance.
[0080] Transcriptome variation is the difference between target-perturbed and unperturbed transcriptome data, used to quantify the expected perturbation of cellular gene expression by small drug molecules.
[0081] In practice, the amount of transcriptomic change between the perturbed and unperturbed transcriptomic data can be determined in several ways. For example, the gene-by-gene difference between the perturbed and unperturbed transcriptomic data can be calculated directly. Another example is the log-ratio between the perturbed and unperturbed transcriptomic data.
[0082] For a detailed description of the transcriptome change encoder, see [link to documentation]. Figure 1 The relevant content will not be repeated here.
[0083] In practice, the similarity between the fusion vector representation and the vector representation of transcriptome changes can be determined in various ways, such as the cosine similarity algorithm, without being limited by the description in this specification.
[0084] In the specific implementation process, after obtaining the similarity of multiple drug molecules to be screened through the above methods, the similarity of multiple drug molecules to be screened is sorted from high to low, and the Top-k molecules are selected as potential effective drugs.
[0085] In some embodiments, such as Figure 3As shown, the fusion vector representation can be input into the drug response prediction head to obtain the area under the dose-response curve corresponding to the small molecule of drug to be screened; based on the area under the dose-response curve, drug screening is performed on the small molecule of drug to be screened.
[0086] The drug response prediction head is a regression model used to predict the area under the dose-response curve (AUC) of the small molecule drug to be screened in a specific cell line or disease model.
[0087] In practice, the drug response prediction head can be implemented in various ways, without being limited to the description in this manual. For example, it can be based on a multilayer perceptron (MLP) or a regression neural network.
[0088] A dose-response curve describes the relationship between drug concentration and cellular response (such as survival rate), and is typically an S-shaped curve. AUC is the integral value under this curve, used to quantify the overall effect of the drug.
[0089] In the specific implementation process, after obtaining the AUC of various small molecules of drugs to be screened through the above methods, the candidate molecules can be sorted from high to low or from low to high according to the needs of the task, and the Top-k molecules can be selected as potential effective drugs.
[0090] The following describes the small molecule virtual screening device based on information bottleneck training provided by the present invention. The small molecule virtual screening device based on information bottleneck training described below can be referred to in correspondence with the small molecule virtual screening method based on information bottleneck training described above.
[0091] Figure 4 This is a schematic diagram of the small molecule virtual screening device based on information bottleneck training provided by the present invention. Figure 4 As shown, the small molecule virtual screening device based on information bottleneck training includes the following modules.
[0092] The first acquisition module 410 is used to acquire a training dataset; wherein the training dataset includes the identifier of the drug small molecule sample, the perturbed transcriptome data sample obtained by perturbing the cell population with the drug small molecule sample, and the undisturbed transcriptome data sample of the cell population.
[0093] The second acquisition module 420 is used to obtain a fusion vector representation of the drug small molecule sample and the unperturbed transcriptome data sample by using a drug small molecule pre-trained model and an initial unperturbed encoder, based on the identifier of the drug small molecule sample and the unperturbed transcriptome data sample.
[0094] The third acquisition module 430 is used to obtain a predicted vector representation of the transcriptome change sample based on the transcriptome change sample between the perturbed transcriptome data sample and its corresponding unperturbed transcriptome data sample using the initial transcriptome change encoder.
[0095] The training module 440 is used to form positive sample pairs by combining the fusion vector representation with the predicted vector representation of the corresponding transcriptome change sample, and to form negative sample pairs by combining the fusion vector representation with the predicted vector representation of the non-corresponding transcriptome change sample. Based on the contrastive loss function and the information bottleneck loss function, the module performs contrastive loss training on the drug small molecule pre-training model, the initial transcriptome change encoder, and the initial unperturbed encoder, so as to use the trained drug small molecule encoder, transcriptome change encoder, and unperturbed encoder to perform the task of screening drug small molecules. The information bottleneck loss function is used to minimize the KL divergence between the distribution of the fusion vector representation and the standard Gaussian distribution.
[0096] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a small molecule virtual screening method based on information bottleneck training. This method includes: acquiring a training dataset; wherein the training dataset contains identifiers of drug small molecule samples, perturbed transcriptome data samples obtained by perturbing a cell population with the drug small molecule samples, and unperturbed transcriptome data samples of the cell population; using a drug small molecule pre-trained model and an initial unperturbed encoder, obtaining a fusion vector representation of the drug small molecule samples and the unperturbed transcriptome data samples based on the identifiers of the drug small molecule samples and the unperturbed transcriptome data samples; using an initial transcriptome change encoder, based on the perturbed transcriptome data samples and their corresponding unperturbed transcriptome data samples... Transcriptome change samples are used to obtain predicted vector representations of transcriptome change samples. Positive sample pairs are formed by combining the fused vector representation with the predicted vector representations of their corresponding transcriptome change samples, and negative sample pairs are formed by combining the fused vector representation with the predicted vector representations of non-corresponding transcriptome change samples. Based on contrastive loss and information bottleneck loss functions, the drug small molecule pre-trained model, the initial transcriptome change encoder, and the initial unperturbed encoder are trained using contrastive loss to perform the task of screening drug small molecules. The information bottleneck loss function is used to minimize the KL divergence between the distribution of the fused vector representation and the standard Gaussian distribution.
[0097] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0098] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the small molecule virtual screening method based on information bottleneck training provided by the above methods. The method includes: acquiring a training dataset; wherein the training dataset includes the identifier of a drug small molecule sample, perturbed transcriptome data samples obtained by perturbing a cell population with the drug small molecule sample, and unperturbed transcriptome data samples of the cell population; using a drug small molecule pre-trained model and an initial unperturbed encoder, obtaining a fusion vector representation of the drug small molecule sample and the unperturbed transcriptome data samples based on the identifier of the drug small molecule sample and the unperturbed transcriptome data samples; and using the initial transcriptome change encoder... The encoder obtains a predicted vector representation of the transcriptome change sample based on the transcriptome change sample between the perturbed transcriptome data sample and its corresponding unperturbed transcriptome data sample. The fused vector representation is paired with the predicted vector representation of its corresponding transcriptome change sample to form a positive sample pair, and the fused vector representation is paired with the predicted vector representation of a non-corresponding transcriptome change sample to form a negative sample pair. Based on a contrastive loss function and an information bottleneck loss function, the drug small molecule pre-trained model, the initial transcriptome change encoder, and the initial unperturbed encoder are trained using contrastive loss to perform the task of screening drug small molecules. The information bottleneck loss function is used to minimize the KL divergence between the distribution of the fused vector representation and the standard Gaussian distribution.
[0099] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the small molecule virtual screening method based on information bottleneck training provided by the above methods. The method includes: acquiring a training dataset; wherein the training dataset includes identifiers of drug small molecule samples, perturbed transcriptome data samples obtained by perturbing a cell population with the drug small molecule samples, and unperturbed transcriptome data samples of the cell population; using a drug small molecule pre-trained model and an initial unperturbed encoder, obtaining a fusion vector representation of the drug small molecule samples and the unperturbed transcriptome data samples based on the identifiers of the drug small molecule samples and the unperturbed transcriptome data samples; and using an initial transcriptome change encoder, based on the perturbed transcriptome data samples... Based on the transcriptome change samples between the sample and its corresponding unperturbed transcriptome data sample, a predicted vector representation of the transcriptome change sample is obtained. The fused vector representation and its corresponding predicted vector representation of the transcriptome change sample are paired to form a positive sample pair, and the fused vector representation and the predicted vector representation of a non-corresponding transcriptome change sample are paired to form a negative sample pair. Based on the contrastive loss function and the information bottleneck loss function, the drug small molecule pre-training model, the initial transcriptome change encoder, and the initial unperturbed encoder are trained using contrastive loss to perform the task of screening drug small molecules. The information bottleneck loss function is used to minimize the KL divergence between the distribution of the fused vector representation and the standard Gaussian distribution.
[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A small molecule virtual screening method based on information bottleneck training, characterized in that, include: Obtain a training dataset; wherein the training dataset includes the identifiers of drug small molecule samples, perturbed transcriptome data samples obtained by perturbing the cell population with the drug small molecule samples, and unperturbed transcriptome data samples of the cell population; Using a drug small molecule pre-trained model and an initial unperturbed encoder, a fused vector representation of the drug small molecule sample and the unperturbed transcriptome data sample is obtained based on the identifier of the drug small molecule sample and the unperturbed transcriptome data sample. Using the initial transcriptome change encoder, a predicted vector representation of the transcriptome change sample is obtained based on the transcriptome change sample between the perturbed transcriptome data sample and its corresponding unperturbed transcriptome data sample. The fusion vector representation and its corresponding predicted vector representation of transcriptome change samples are used to form positive sample pairs, and the fusion vector representation and the predicted vector representation of non-corresponding transcriptome change samples are used to form negative sample pairs. Based on the contrastive loss function and the information bottleneck loss function, the drug small molecule pre-trained model, the initial transcriptome change encoder, and the initial unperturbed encoder are trained using contrastive loss, so as to use the trained drug small molecule encoder, transcriptome change encoder, and unperturbed encoder to perform the task of screening drug small molecules; wherein, the information bottleneck loss function is used to minimize the KL divergence between the distribution of the fusion vector representation and the standard Gaussian distribution.
2. The small molecule virtual screening method based on information bottleneck training according to claim 1, characterized in that, The process of using a drug small molecule pre-trained model and an initial unperturbed encoder to obtain a fused vector representation of the drug small molecule sample and the unperturbed transcriptome data sample based on the identifier of the drug small molecule sample and the unperturbed transcriptome data sample includes: Using the drug small molecule pre-trained model, a predicted vector representation of the drug small molecule sample is obtained based on the identifier of the drug small molecule sample; Using the initial unperturbed encoder, a predicted vector representation of the unperturbed transcriptome data sample is obtained based on the unperturbed transcriptome data sample; A variational encoder is used to fuse the predicted vector representation of the drug small molecule sample and the predicted vector representation of the unperturbed transcriptome data sample to obtain a fused vector representation of the drug small molecule sample and the unperturbed transcriptome data sample.
3. The small molecule virtual screening method based on information bottleneck training according to claim 2, characterized in that, The step of fusing the predicted vector representation of the drug small molecule sample and the predicted vector representation of the unperturbed transcriptome data sample using a variational encoder to obtain a fused vector representation of the drug small molecule sample and the unperturbed transcriptome data sample includes: The predicted vector representation of the drug small molecule sample and the predicted vector representation of the unperturbed transcriptome data sample are concatenated to obtain the predicted concatenated vector; The variational encoder is used to perform latent space sampling on the predicted splicing vector to obtain the fused vector representation.
4. The small molecule virtual screening method based on information bottleneck training according to claim 3, characterized in that, During the contrastive loss training process, each batch of sample data comes from the same cell line.
5. The small molecule virtual screening method based on information bottleneck training according to any one of claims 1 to 4, characterized in that, The contrastive loss function is as follows: ; in, N This represents the total number of positive and negative sample pairs. k It is a positive integer; The first contrastive loss function is used to train and maximize the fused vector representation of sample pairs k. Its positive samples The probability of a match is calculated using the following formula: ; in, For the first k The fusion vector representation of a pair of positive samples. For the first k The predicted vector representation of transcriptome change samples for each positive sample pair; Temperature coefficient; For the first i The predicted vector representation of transcriptome change samples for each sample pair; in, The second contrastive loss function maximizes the sample pairs through contrastive training. k Vector representation of transcriptome change samples Its positive samples The probability of a match is expressed by the following formula: in, For the first i The fusion vector representation of each sample pair.
6. The small molecule virtual screening method based on information bottleneck training according to claim 5, characterized in that, The information bottleneck loss function is as follows: in, N The number of sample pairs participating in training; The dimension of the final fused vector representation. and The first The fusion vector representation of the nth sample pair Mean and standard deviation of each dimension; It follows a standard Gaussian distribution; The KL divergence between the final fusion vector representation and the standard Gaussian distribution.
7. A small molecule virtual screening device based on information bottleneck training, characterized in that, include: The first acquisition module is used to acquire a training dataset; wherein, the training dataset includes the identifier of the drug small molecule sample, the perturbed transcriptome data sample obtained by perturbing the cell population with the drug small molecule sample, and the undisturbed transcriptome data sample of the cell population. The second acquisition module is used to obtain a fusion vector representation of the drug small molecule sample and the unperturbed transcriptome data sample by using a drug small molecule pre-trained model and an initial unperturbed encoder, based on the identifier of the drug small molecule sample and the unperturbed transcriptome data sample. The third acquisition module is used to obtain a predicted vector representation of the transcriptome change sample based on the transcriptome change sample between the perturbed transcriptome data sample and its corresponding unperturbed transcriptome data sample using the initial transcriptome change encoder. The training module is used to form positive sample pairs by combining the fused vector representation with the predicted vector representation of the corresponding transcriptome change sample, and negative sample pairs by combining the fused vector representation with the predicted vector representation of the non-corresponding transcriptome change sample. Based on the contrastive loss function and the information bottleneck loss function, the module performs contrastive loss training on the drug small molecule pre-training model, the initial transcriptome change encoder, and the initial unperturbed encoder, so as to use the trained drug small molecule encoder, transcriptome change encoder, and unperturbed encoder to perform the task of screening drug small molecules. The information bottleneck loss function is used to minimize the KL divergence between the distribution of the fused vector representation and the standard Gaussian distribution.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the small molecule virtual screening method based on information bottleneck training as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the small molecule virtual screening method based on information bottleneck training as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the small molecule virtual screening method based on information bottleneck training as described in any one of claims 1 to 6.