Drug small molecule virtual screening method and device based on transcriptome and electronic equipment

By employing a transcriptome-based virtual screening method for small molecule drugs, which utilizes an encoder to generate a fusion vector representation and combines transcriptome changes for drug screening, this approach addresses the problem of poor screening performance for complex diseases in existing technologies and enables accurate prediction of drug efficacy in new cell lines.

CN120998346APending Publication Date: 2025-11-21TSINGHUA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511043972.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies have limited effectiveness in virtual screening of small molecule drugs for complex diseases such as Alzheimer's disease and type 2 diabetes, and perform poorly in cross-cell line drug screening, reflecting the limitations of modeling molecular perturbation effects.

Method used

A transcriptome-based virtual screening method for small molecule drugs is adopted. By fusing a drug small molecule encoder, an unperturbed encoder, and a variational encoder, a fused vector representation is generated. Drug screening is performed in combination with transcriptome changes. The trained encoder is used for contrastive loss training to obtain accurate and robust drug screening results.

Benefits of technology

It enables precise and robust virtual screening of small drug molecules, accurately predicts molecular effects in new cell lines, and improves the screening effect for complex diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998346A_ABST
    Figure CN120998346A_ABST
Patent Text Reader

Abstract

The invention provides a transcriptome-based drug small molecule virtual screening method and device and electronic equipment, and relates to the technical field of artificial intelligence. The transcriptome-based drug small molecule virtual screening method comprises the following steps: acquiring an identifier of a drug small molecule to be screened and undisturbed transcriptome data to be matched; using a drug small molecule encoder to obtain vector representation of the drug small molecules to be screened according to the identifiers of the drug small molecules to be screened; obtaining a vector representation of the to-be-matched undisturbed transcriptome data by using an undisturbed encoder according to the to-be-matched undisturbed transcriptome data; fusing the vector representation of the drug small molecules to be screened and the vector representation of the undisturbed transcriptome data to be matched by using a variational encoder to obtain a fused vector representation; and performing drug screening on the to-be-screened drug small molecules based on the fusion vector representation. According to the method, virtual screening of the drug small molecules can be accurately and robustly carried out.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a drug small molecule virtual screening method and device based on transcriptome and electronic equipment. BACKGROUND

[0002] In recent years, target-based virtual screening technology has achieved remarkable results in drug development by identifying compounds that bind to disease-related targets. However, this method is only suitable for diseases with clear targets, and its screening effect is very limited for complex diseases such as neurodegenerative diseases (e.g. Alzheimer's disease) or multifactorial metabolic syndromes (e.g. type 2 diabetes).

[0003] With the rapid development of systems biology and omics technology, phenotype-based virtual screening technology has gradually become a research hotspot. This strategy does not rely on pre-set target information, but rather assesses the ability of molecules to reverse disease-related phenotypes at the cellular or tissue level to screen candidate compounds. Statistical data shows that between 1999 and 2008, more "first-in-class" drugs were approved using the phenotype strategy than the target strategy, highlighting its clinical translation potential. However, existing phenotype-based virtual screening technologies perform poorly in cross-cell line drug screening, reflecting their limitations in modeling molecular perturbation effects.

[0004] Therefore, how to accurately and robustly perform virtual screening of drug small molecules is a technical problem that needs to be solved. SUMMARY

[0005] The present application provides a drug small molecule virtual screening method and device based on transcriptome and electronic equipment to solve the above defects in the prior art, to accurately and robustly perform virtual screening of drug small molecules.

[0006] The present application provides a drug small molecule virtual screening method based on transcriptome, comprising the following steps.

[0007] Obtain the identification of the drug small molecule to be screened and the undisturbed transcriptome data to be matched; use a drug small molecule encoder to obtain the vector representation of the drug small molecule to be screened according to the identification of the drug small molecule to be screened; use an undisturbed encoder to obtain the vector representation of the undisturbed transcriptome data to be matched according to the undisturbed transcriptome data to be matched; use a variational encoder to fuse the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched, to obtain a fused vector representation; and perform drug screening on the drug small molecule to be screened based on the fused vector representation.

[0008] According to the application, a transcriptome-based drug small molecule virtual screening method is provided, and the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched are fused by using the variational encoder to obtain a fused vector representation, including: splicing the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched to obtain a spliced vector; the variational encoder is used for probabilistic modeling of the spliced vector, and the fused vector representation is obtained by sampling from the latent space of the probabilistic modeling.

[0009] According to the application, a transcriptome-based drug small molecule virtual screening method is provided, and the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched are fused by using the variational encoder to obtain a fused vector representation, including: splicing the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched to obtain a spliced vector; the variational encoder is used for probabilistic modeling of the spliced vector, and the fused vector representation is obtained by sampling from the latent space of the probabilistic modeling.

[0010] According to the application, a transcriptome-based drug small molecule virtual screening method is provided, and the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched are fused by using the variational encoder to obtain a fused vector representation, including: splicing the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched to obtain a spliced vector; the variational encoder is used for probabilistic modeling of the spliced vector, and the fused vector representation is obtained by sampling from the latent space of the probabilistic modeling.

[0011] According to the drug small molecule virtual screening method based on transcriptome provided by the application, the drug small molecule encoder, the transcriptome change encoder and the undisturbed encoder are obtained by the following training method: obtaining a training data set, wherein the training data set contains the identification of a drug small molecule sample, the disturbed transcriptome data sample obtained by disturbing a cell population by the drug small molecule sample and the undisturbed transcriptome data sample of the cell population; using a drug small molecule pre-training model to obtain the predicted vector representation of the drug small molecule sample according to the identification of the drug small molecule sample; using an initial undisturbed encoder to obtain the predicted vector representation of the undisturbed transcriptome data sample according to the undisturbed transcriptome data sample; using a variational encoder to fuse the predicted vector representation of the drug small molecule sample and the predicted vector representation of the undisturbed transcriptome data sample to obtain the predicted fusion vector representation of the drug small molecule sample and the undisturbed transcriptome data sample; using an initial transcriptome change encoder to obtain the predicted vector representation of the transcriptome change sample according to the transcriptome change sample between the disturbed transcriptome data sample and the corresponding undisturbed transcriptome data sample; taking the predicted fusion vector representation and the predicted vector representation of the corresponding transcriptome change sample as a positive sample pair, taking the predicted fusion vector representation and the predicted vector representation of the transcriptome change sample not corresponding thereto as a negative sample pair, and based on a loss function, performing contrastive loss training on the drug small molecule pre-training model, the initial transcriptome change encoder and the initial undisturbed encoder to obtain the drug small molecule encoder, the transcriptome change encoder and the undisturbed encoder.

[0012] According to the drug small molecule virtual screening method based on transcriptome provided by the application, in the contrastive loss training process, each batch of sample data is from the same cell line.

[0013] The application further provides a drug small molecule virtual screening device based on transcriptome, comprising the following modules: The first obtaining module is used to obtain the identification of a drug small molecule to be screened and undisturbed transcriptome data to be matched; the second obtaining module is used to obtain the vector representation of the drug small molecule to be screened according to the identification of the drug small molecule to be screened by using a drug small molecule encoder; the third obtaining module is used to obtain the vector representation of the undisturbed transcriptome data to be matched according to the undisturbed transcriptome data to be matched by using an undisturbed encoder; the fourth obtaining module is used to fuse the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched to obtain a fusion vector representation by using a variational encoder; and the screening module is used to perform drug screening on the drug small molecule to be screened based on the fusion vector representation.

[0014] The application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the transcriptome-based drug small molecule virtual screening method according to any one of the above when executing the computer program.

[0015] The application further provides a non-transitory computer-readable storage medium, having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the transcriptome-based drug small molecule virtual screening method according to any one of the above.

[0016] The application further provides a computer program product, comprising a computer program, wherein the computer program, when executed by a processor, implements the transcriptome-based drug small molecule virtual screening method according to any one of the above.

[0017] The application provides a transcriptome-based drug small molecule virtual screening method, device and electronic device, which utilizes a drug small molecule encoder to obtain a vector representation of a drug small molecule to be screened according to an identifier of the drug small molecule to be screened, utilizes an undisturbed encoder to obtain a vector representation of undisturbed transcriptome data to be matched according to the undisturbed transcriptome data to be matched, utilizes a variational encoder to fuse the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched to obtain a fused vector representation, and performs drug screening on the drug small molecule to be screened based on the fused vector representation. Since the fused vector representation contains the correlation between the drug small molecule and gene expression, the virtual screening of the drug small molecule can be accurately and robustly performed based on the fused vector representation. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0019] Figure 1 is a flowchart of the transcriptome-based drug small molecule virtual screening method provided by the application.

[0020] Figure 2 is a flowchart of the training method of the drug small molecule encoder provided by the application.

[0021] Figure 3 is a training process diagram of the drug small molecule encoder provided by the application.

[0022] Figure 4 is a structural diagram of the transcriptome-based drug small molecule virtual screening device provided by the application.

[0023] Figure 5 FIG. 1 is a structural schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION

[0024] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, but not all embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0025] The present application provides a method for virtual screening of small molecule drugs based on transcriptome. Figures 1-3 The present application provides a method for virtual screening of small molecule drugs based on transcriptome.

[0026] Figure 1 FIG. 2 is a flowchart of the method for virtual screening of small molecule drugs based on transcriptome provided by the present application, as shown in FIG. 2, the method comprises the following steps: Figure 1 Step 101, obtaining the identification of the small molecule drug to be screened and the undisturbed transcriptome data to be matched.

[0027] The small molecule drug to be screened is a chemical molecule with therapeutic potential but whose activity has not been verified, and it needs to be evaluated for its influence on cell phenotype by virtual screening technology. For example, the small molecule drug to be screened can be a candidate compound for Alzheimer's disease (such as β-secretase inhibitor), a potential small molecule for type 2 diabetes (such as PPARγ agonist), etc.

[0028] The identification of the small molecule drug to be screened is the unique identifier of the small molecule. For example, the identification of the small molecule can be a SMILES string (such as “O=C(O)C1=CC=CC=C1” representing aspirin), a molecular graph structure, etc.

[0029] The undisturbed transcriptome data to be matched is the gene expression data of the disease phenotype cell line when it is not intervened by the drug, which is used as a benchmark to evaluate the drug disturbance effect. For example, the undisturbed transcriptome data to be matched can be the transcriptome sequencing data of a certain disease-related phenotype cell line (HepG2) under the condition of no drug intervention.

[0030] Step 102, using a small molecule drug encoder to obtain the vector representation of the small molecule drug to be screened according to the identification of the small molecule drug to be screened.

[0031] ​The drug small molecule encoder is a deep learning model that can map the identity of a drug small molecule under screening to a low-dimensional vector representation through self-supervised pre-training tasks and contrastive learning fine-tuning. The core goal is to capture the three-dimensional structural characteristics, physical and chemical properties, and potential mechanisms of drug small molecules, thereby supporting subsequent modeling of transcriptional perturbation effects.

[0032] In the specific implementation process, the drug small molecule encoder can be constructed through various neural network model architectures, such as 3D-Transformer, which is not limited by the description in this specification.

[0033] The input data of the drug small molecule encoder is the identity of the drug small molecule under screening.

[0034] A fixed-dimensional vector representation (such as 512 dimensions) is generated as the final encoding of the drug small molecule under screening, and the output is the vector representation of the drug small molecule.

[0035] The vector representation of the drug small molecule encodes the following information of the drug small molecule: three-dimensional structural characteristics such as atomic spatial arrangement, bond length, and bond angle (learned through a three-dimensional position recovery task); physical and chemical properties such as hydrophobicity and charge distribution (learned through an atomic masking prediction task); and mechanism knowledge such as potential interaction patterns with protein targets (learned through contrastive learning of aligned protein target representations).

[0036] Step 103, using the unperturbed encoder, obtaining the vector representation of the matched unperturbed transcriptomic data according to the matched unperturbed transcriptomic data.

[0037] The unperturbed encoder is a trained machine learning model that maps unperturbed transcriptomic data (such as a gene expression matrix) to a fixed-dimensional vector representation. This vector contains key biological characteristics of the cell in its basal state (such as gene regulation patterns, cell type specificity) to facilitate comparative analysis with the transcriptional changes after drug perturbation, thereby evaluating the perturbation effect of molecules on cells.

[0038] In the specific implementation process, the unperturbed encoder can be constructed based on various neural network model architectures, such as Transformer, which is not limited by the description in this specification. The input of the unperturbed encoder is the matched unperturbed transcriptomic data, and the output is the vector representation of the matched unperturbed transcriptomic data.

[0039] Step 104, using the variational encoder, fusing the vector representation of the drug small molecule under screening and the vector representation of the matched unperturbed transcriptomic data to obtain a fused vector representation.

[0040] Variational encoder is a generative neural network, which is a variant or simplified form of variational autoencoder (VAE), and its core goal is to learn the probability distribution of input data in the latent space and generate more robust vector representation by re-sampling. The variational encoder outputs the mean and variance of the latent distribution, and introduces noise by random sampling, which can effectively enhance the generalization of the model.

[0041] In some embodiments, the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched can be spliced to obtain a spliced vector; the variational encoder is used to probabilistically model the spliced vector, and a fusion vector representation is obtained by sampling from the latent space obtained by probabilistic modeling.

[0042] In the specific implementation process, the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched can be spliced along the feature dimension to form an initial fusion vector, which is used as the input data of the variational encoder. The variational encoder outputs the mean (μ) and variance (σ²) vectors of the latent distribution. The latent distribution: assuming that the latent space obeys a diagonal Gaussian distribution, a vector can be randomly sampled from the latent distribution as the final fusion vector representation.

[0043] The effect of a drug small molecule depends not only on its own structure, but also on the inherent properties of the cell type, metabolic state, etc. In the embodiments provided by the present application, the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched are fused by using the variational encoder, which can effectively implicitly capture the molecular-cell interaction (such as target binding, metabolic regulation).

[0044] Step 105, based on the fusion vector representation, performing drug screening on the drug small molecule to be screened.

[0045] In the specific implementation process, a variety of drug screening tasks can be performed based on the fusion vector representation of the drug small molecule to be screened, which is not limited by the description in the specification.

[0046] In some embodiments, the target perturbed transcriptome data is obtained; the transcriptome change amount between the target perturbed transcriptome data and the undisturbed transcriptome data to be matched is determined; the vector representation of the transcriptome change amount is obtained by using the transcriptome change encoder according to the transcriptome change amount; the similarity between the fusion vector representation and the vector representation of the transcriptome change amount is determined; and the drug small molecule to be screened is screened based on the similarity.

[0047] The target perturbed transcriptome data refers to the change data of the cell at the gene expression level after the cell is expected to be treated with a drug small molecule, which is usually represented as a matrix of gene expression amount (such as FPKM value of RNA-seq, fluorescence intensity of microarray) or transcript abundance.

[0048] The transcriptomic change is a difference between the post-target perturbation transcriptomic data and the to-be-matched non-perturbation transcriptomic data, and is used to quantify the expected perturbation degree of the small molecule drug on the gene expression of the cell.

[0049] In the specific implementation, the transcriptomic change between the post-target perturbation transcriptomic data and the to-be-matched non-perturbation transcriptomic data can be determined in various ways. For example, the gene-by-gene difference between the post-target perturbation transcriptomic data and the to-be-matched non-perturbation transcriptomic data can be directly calculated. For another example, the log ratio between the post-target perturbation transcriptomic data and the to-be-matched non-perturbation transcriptomic data can be calculated.

[0050] The transcriptomic change encoder is a trained machine learning model, and is used to map the high-dimensional transcriptomic change to a low-dimensional vector representation.

[0051] In the specific implementation, the transcriptomic change encoder can be constructed based on various neural network model architectures, such as Transformer, and is not limited by the description in the specification.

[0052] The input data of the transcriptomic change encoder is the transcriptomic change vector, for example, the gene expression change calculated by the difference method or the log ratio method. The output data of the transcriptomic change encoder is the vector representation of the transcriptomic change.

[0053] In the specific implementation, the similarity between the fusion vector representation and the vector representation of the transcriptomic change can be determined in various ways, such as the cosine similarity algorithm, and is not limited by the description in the specification.

[0054] In the specific implementation, after obtaining the similarities of the various to-be-screened small molecule drugs in the above manner, the candidate molecules can be sorted according to the AUC value from high to low or from low to high according to the task requirement, and the Top-k molecules are selected as potential effective drugs.

[0055] In some embodiments, as shown in FIG. 2, the fusion vector representation can be input into the drug response prediction head to obtain the area under the dose-response curve corresponding to the to-be-screened small molecule drug; and the small molecule drug is screened based on the area under the dose-response curve. Figure 3

[0056] The drug response prediction head is a regression model, and is used to predict the area under the dose-response curve (AUC, Area Under the dose–response Curve) of the to-be-screened small molecule drug in a specific cell line or disease model.

[0057] ​In specific implementation, the drug response prediction head can be implemented in various ways without being limited by the description herein. For example, the drug response prediction head can be implemented based on a multi-layer perception (MLP) or a regression neural network.

[0058] The dose-response curve describes the relationship between drug concentration and cell response (such as survival rate), usually in an S-shaped curve. The AUC is the integral value under the curve, which is used to quantify the overall effect of the drug.

[0059] In specific implementation, after obtaining the AUC of a plurality of small molecule drug candidates by the above method, the candidate molecules can be sorted in descending order of AUC value, and the top-k molecules can be selected as potential effective drugs.

[0060] Figure 2 is a flowchart of the training method of the small molecule drug encoder provided by the present application, as shown in Figure 2 The method comprises the following steps: Step 201, obtaining a training data set, wherein the training data set comprises an identification of a small molecule drug sample, perturbed transcriptome data samples obtained by perturbing a cell population with the small molecule drug sample, and unperturbed transcriptome data samples of the cell population.

[0061] The identification of the small molecule drug sample, for example, a SMILES string or a molecular graph, is used to uniquely identify the small molecule drug.

[0062] The perturbed transcriptome data sample is the gene expression profile (such as RNA-seq data or microarray data) of the cell population after being treated with the small molecule drug sample.

[0063] The unperturbed transcriptome data sample is the gene expression profile of the cell population corresponding to the perturbed transcriptome data sample without drug treatment.

[0064] In specific implementation, the identification of the small molecule drug, the transcriptome data after drug perturbation, and the unperturbed transcriptome data of the same cell line and the same batch information can be obtained from the high-throughput drug perturbation data set to obtain the training data set.

[0065] To avoid the interference of cell line inherent properties on perturbation effect learning, each batch of sample data in the contrast loss training process comes from the same cell line.

[0066] Step 202, using a small molecule drug pre-training model to obtain a predicted vector representation of the small molecule drug sample according to the identification of the small molecule drug sample.

[0067] A drug small molecule pre-training model is a deep neural network based on self-supervised learning, which is used to convert the identification (such as a SMILES string, a molecular graph structure) of a drug small molecule into a low-dimensional dense molecular representation vector. The model learns the general features of the molecule through the pre-training task, without relying on labeled data, so as to capture the molecular properties (such as bond length, angle, charge distribution, etc.) across physical and biochemical scales. The drug small molecule pre-training model can be constructed based on the Transformer.

[0068] The pre-training task includes: A three-dimensional position recovery (3D Position Recovery) task is used to train the model to learn the geometric conformation (such as bond angle, dihedral angle) of the molecule, and capture the influence of spatial arrangement on activity.

[0069] An atom masking prediction task is used to train the model to randomly mask part of the atoms in the molecule (such as marked with a mask), and predict the type (such as carbon, oxygen, nitrogen) of the masked atoms, so as to learn the local chemical environment (such as functional group, bond type) of the molecule.

[0070] Step 203, using an initial unperturbed encoder, obtaining a predicted vector representation of the unperturbed transcriptome data sample according to the unperturbed transcriptome data sample.

[0071] The initial unperturbed encoder is an unperturbed encoder that has not been trained with a contrastive loss. For detailed description of the unperturbed encoder, please refer to the related content in Figure 1 , which will not be repeated here.

[0072] Step 204, using a variational encoder, fusing the predicted vector representation of the drug small molecule sample and the predicted vector representation of the unperturbed transcriptome data sample to obtain a predicted fusion vector representation of the drug small molecule sample and the unperturbed transcriptome data sample.

[0073] For detailed description of the variational encoder, please refer to the related content in Figure 1 , which will not be repeated here.

[0074] Step 205, using an initial transcriptome change encoder, obtaining a predicted vector representation of the transcriptome change quantity sample according to the transcriptome change quantity sample between the perturbed transcriptome data sample and its corresponding unperturbed transcriptome data sample.

[0075] The initial transcriptome change encoder is a transcriptome change encoder that has not been trained with a contrastive loss. For detailed description of the transcriptome change encoder, please refer to the related content in Figure 1 , which will not be repeated here.

[0076] Step 206: Using the predicted fusion vector and the predicted vector of the corresponding transcriptome change sample as positive sample pairs, and using the predicted fusion vector and the predicted vector of the transcriptome change sample that does not correspond to it as negative sample pairs, perform comparative loss training on the drug small molecule pre-trained model, the initial transcriptome change encoder, and the initial unperturbed encoder based on the loss function to obtain the drug small molecule encoder, the transcriptome change encoder, and the unperturbed encoder.

[0077] Using contrastive loss training can increase the cosine similarity between vector representations of positive sample pairs and decrease the cosine similarity between vector representations of negative sample pairs, so that the final fused vector representation determined by specific drug small molecules and specific cellular background can be consistent with the vector representation of the corresponding transcriptome changes.

[0078] In some embodiments, the loss function includes the contrastive loss function shown below. : ; (1) in, N This represents the total number of positive and negative sample pairs. k The ratio is a positive integer. Loss function. By bidirectional contrastive loss function and get.

[0079] The first contrastive loss function is used to train and maximize the fused vector representation of sample pairs k. Its positive samples The probability of a match is calculated using the following formula: ; (2) in, For the first k The predicted fusion vector representation of a pair of positive samples. For the first k The predicted vector representation of transcriptome change samples for each positive sample pair; This is a temperature coefficient used to control the sharpness of the distribution in the softmax function; For the first i The predicted vector representation of transcriptome change samples for each sample pair.

[0080] The second contrastive loss function maximizes the sample pairs through contrastive loss training. k Vector representation of transcriptome change samples Its positive samples The probability of a match is expressed by the following formula: (3) wherein, is the predicted fused vector representation of the i-th sample pair. i

[0081] In some embodiments, the loss function further comprises an information bottleneck loss function as shown below: (4) wherein, N is the number of sample pairs participating in training; is the dimension of the final fused vector representation, and are the mean and standard deviation of the d-th dimension of the predicted fused vector representation of the i-th sample pair, respectively; is a standard Gaussian distribution; is the KL divergence between the distribution of the final fused vector representation and the standard Gaussian distribution.

[0082] In the embodiments provided by the present application, the information compression constraint is introduced through the information bottleneck loss function, so that the model participating in training only retains the effective information related to the prediction target, and irrelevant disturbance factors are filtered.

[0083] In the specific implementation process, the joint loss function can be established according to formulas (1) to (4); the parameters of the small molecule pre-training model, the initial undisturbed encoder and the initial transcriptome change encoder are updated by using a preset optimization algorithm (for example, a gradient descent optimization algorithm) to minimize the value of the joint loss function, and the training end condition (for example, all models converge or reach a preset training number) is directly reached, so as to obtain the trained drug small molecule encoder, transcriptome change encoder and undisturbed encoder.

[0084] In the embodiments provided by the present application, through comparative training, the small molecule pre-training model is optimized into a drug small molecule encoder, and the vector representation output by the drug small molecule encoder satisfies the following characteristics: implicit modeling mechanism, through comparative learning, the vector representation of the drug small molecule output by the drug small molecule encoder not only contains chemical structure information, but also implicitly encodes the disturbance mode (such as target point binding, signal pathway activation) of the drug small molecule on the cell transcriptome; cross-cell line generalization ability: due to the use of a cell line-based sampling strategy (avoiding interference of cell line inherent attributes) and an information bottleneck loss (filtering irrelevant noise) in training, the vector representation of the drug small molecule output by the drug small molecule encoder can accurately predict the molecular effect in a new cell line, align the vector representation of the drug small molecule with the vector representation of the protein target, and thus implicitly learn the drug-target interaction mode.

[0085] ​​​The transcriptional-based drug small molecule virtual screening device provided by the present application is described below, and the transcriptional-based drug small molecule virtual screening device described below can be correspondingly referred to the transcriptional-based drug small molecule virtual screening method described above.

[0086] Figure 4 is a structural schematic diagram of the transcriptional-based drug small molecule virtual screening device provided by the present application. As shown in Figure 4 The transcriptional-based drug small molecule virtual screening device includes the following modules.

[0087] The first acquisition module 410 is configured to acquire an identifier of a drug small molecule to be screened and undisturbed transcriptional data to be matched.

[0088] The second acquisition module 420 is configured to obtain a vector representation of the drug small molecule to be screened by using a drug small molecule encoder according to the identifier of the drug small molecule to be screened.

[0089] The third acquisition module 430 is configured to obtain a vector representation of the undisturbed transcriptional data to be matched by using an undisturbed encoder according to the undisturbed transcriptional data to be matched.

[0090] The fourth acquisition module 440 is configured to fuse the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptional data to be matched by using a variational encoder to obtain a fused vector representation.

[0091] The screening module 450 is configured to perform drug screening on the drug small molecule to be screened based on the fused vector representation.

[0092] Figure 5 An example of an entity structure schematic diagram of an electronic device is shown in Figure 5As shown, the electronic device can include a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 complete mutual communication through the communications bus 540. The processor 510 can invoke a logic instruction in the memory 530 to execute a transcript-based drug small molecule virtual screening method, which includes: obtaining an identification of a drug small molecule to be screened and undisturbed transcriptome data to be matched; using a drug small molecule encoder to obtain a vector representation of the drug small molecule to be screened according to the identification of the drug small molecule to be screened; using an undisturbed encoder to obtain a vector representation of the undisturbed transcriptome data to be matched according to the undisturbed transcriptome data to be matched; using a variational encoder to fuse the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched to obtain a fused vector representation; and performing drug screening on the drug small molecule to be screened based on the fused vector representation.

[0093] In addition, the logic instruction in the memory 530 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0094] In another aspect, the present application also provides a computer program product comprising a computer program, which can be stored on a non-transitory computer readable storage medium, and the computer program, when executed by a processor, enables a computer to perform the transcriptome-based drug small molecule virtual screening method provided by the above-mentioned methods, and the method comprises: obtaining an identifier of a drug small molecule to be screened and undisturbed transcriptome data to be matched; using a drug small molecule encoder to obtain a vector representation of the drug small molecule to be screened according to the identifier of the drug small molecule to be screened; using an undisturbed encoder to obtain a vector representation of the undisturbed transcriptome data to be matched according to the undisturbed transcriptome data to be matched; using a variational encoder to fuse the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched to obtain a fused vector representation; and performing drug screening on the drug small molecule to be screened based on the fused vector representation.

[0095] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, and the computer program, when executed by a processor, implements a transcriptome-based drug small molecule virtual screening method provided by the above-mentioned methods, and the method comprises: obtaining an identifier of a drug small molecule to be screened and undisturbed transcriptome data to be matched; using a drug small molecule encoder to obtain a vector representation of the drug small molecule to be screened according to the identifier of the drug small molecule to be screened; using an undisturbed encoder to obtain a vector representation of the undisturbed transcriptome data to be matched according to the undisturbed transcriptome data to be matched; using a variational encoder to fuse the vector representation of the drug small molecule to be screened and the vector representation of the undisturbed transcriptome data to be matched to obtain a fused vector representation; and performing drug screening on the drug small molecule to be screened based on the fused vector representation.

[0096] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0097] Those skilled in the art can clearly understand the implementation of the various embodiments by means of software and necessary general hardware platforms through the description of the above embodiments, and of course, the implementation can also be through hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the method described in each embodiment or some parts of the embodiment.

[0098] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for transcriptome-based virtual screening of small drug molecules, characterized in that, The method comprises the following steps: obtaining the identification of a drug small molecule to be screened and un-perturbed transcriptome data to be matched; using a drug small molecule encoder to obtain a vector representation of the drug small molecule to be screened according to the identification of the drug small molecule to be screened; using an un-perturbed encoder to obtain a vector representation of the un-perturbed transcriptome data to be matched according to the un-perturbed transcriptome data to be matched; using a variational encoder to fuse the vector representation of the drug small molecule to be screened and the vector representation of the un-perturbed transcriptome data to be matched to obtain a fused vector representation; based on the fused vector representation, performing drug screening on the drug small molecule to be screened.

2. The transcriptome-based small molecule drug virtual screening method according to claim 1, characterized in that, The method comprises the following steps: splicing the vector representation of the drug small molecule to be screened and the vector representation of the un-perturbed transcriptome data to be matched to obtain a spliced vector; using the variational encoder to probabilistically model the spliced vector and sampling from a latent space obtained from the probabilistic modeling to obtain the fused vector representation.

3. The transcriptome-based small molecule drug virtual screening method according to claim 1, characterized in that, The method comprises the following steps: obtaining target perturbed transcriptome data; determining the transcriptome change between the target perturbed transcriptome data and the un-perturbed transcriptome data to be matched; using a transcriptome change encoder to obtain a vector representation of the transcriptome change according to the transcriptome change; determining the similarity between the fused vector representation and the vector representation of the transcriptome change; based on the similarity, performing drug screening on the drug small molecule to be screened.

4. The transcriptome-based small molecule drug virtual screening method according to any one of claims 1, wherein, The method comprises the following steps: inputting the fused vector representation into a drug response prediction head to obtain an area under the dose-response curve corresponding to the drug small molecule to be screened; based on the area under the dose-response curve, performing drug screening on the drug small molecule to be screened.

5. The transcriptome-based small molecule drug virtual screening method according to any one of claims 1 to 4, characterized in that, The drug small molecule encoder, the transcriptome change encoder, and the un-perturbed encoder are trained in the following manner: obtaining a training data set, wherein the training data set contains the identification of a drug small molecule sample, perturbed transcriptome data samples obtained by perturbing a cell population of the drug small molecule sample, and un-perturbed transcriptome data samples of the cell population; using a drug small molecule pre-training model to obtain a predicted vector representation of the drug small molecule sample according to the identification of the drug small molecule sample; using an initial un-perturbed encoder to obtain a predicted vector representation of the un-perturbed transcriptome data sample according to the un-perturbed transcriptome data sample; using a variational encoder to fuse the predicted vector representation of the drug small molecule sample and the predicted vector representation of the un-perturbed transcriptome data sample to obtain a predicted fused vector representation of the drug small molecule sample and the un-perturbed transcriptome data sample; obtaining a predicted vector representation of the transcriptomic change amount sample according to a transcriptomic change amount sample between the perturbed transcriptome data sample and its corresponding unperturbed transcriptome data sample by using an initial transcriptomic change encoder; using the predicted fusion vector representation and the predicted vector representation of the transcriptomic change amount sample corresponding thereto as a positive sample pair, using the predicted fusion vector representation and the predicted vector representation of the transcriptomic change amount sample not corresponding thereto as a negative sample pair, and based on a loss function, performing contrastive loss training on the drug small molecule pre-training model, the initial transcriptomic change encoder, and the initial unperturbed encoder to obtain the drug small molecule encoder, the transcriptomic change encoder, and the unperturbed encoder.

6. The transcriptome-based small molecule drug virtual screening method according to claim 5, characterized in that, In the contrastive loss training process, each batch of sample data is from the same cell line.

7. A device for transcriptome-based virtual screening of small drug molecules, characterized in that, The method comprises the following steps: a first obtaining module configured to obtain an identifier of a drug small molecule to be screened and unperturbed transcriptome data to be matched; a second obtaining module configured to obtain a vector representation of the drug small molecule to be screened according to the identifier of the drug small molecule to be screened by using a drug small molecule encoder; a third obtaining module configured to obtain a vector representation of the unperturbed transcriptome data to be matched according to the unperturbed transcriptome data to be matched by using an unperturbed encoder; a fourth obtaining module configured to fuse the vector representation of the drug small molecule to be screened and the vector representation of the unperturbed transcriptome data to be matched to obtain a fusion vector representation by using a variational encoder; a screening module configured to perform drug screening on the drug small molecule to be screened based on the fusion vector representation.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor executes the computer program to implement the method for virtual screening of drug small molecules based on transcriptome according to any one of claims 1 to 6. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method for virtual screening of drug small molecules based on transcriptome according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the method for virtual screening of drug small molecules based on transcriptome according to any one of claims 1 to 6.