Virtual screening algorithm based on gene expression profile and contrast learning

By using a virtual screening algorithm based on gene expression profiles and contrastive learning, and utilizing a dual-feature encoder and a joint loss function optimization model, the problems of high cost, long cycle, and low computational accuracy in drug development are solved, achieving efficient and rapid drug screening.

CN120808882APending Publication Date: 2025-10-17NANJING TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510919511.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In existing technologies, drug development is costly, time-consuming, and relies on single-modality data, resulting in low computational accuracy and making it difficult to efficiently screen potential compounds.

Method used

A virtual screening algorithm based on gene expression profiles and contrastive learning is adopted. The joint representation of drugs and gene expression profiles is learned through a dual-feature encoder. The cosine similarity scoring function and joint loss function are used to optimize the model to achieve cross-modal matching of drugs and gene expression.

Benefits of technology

It achieves efficient and rapid drug screening, improves calculation accuracy and adaptability, and reduces the cost and time of drug research and development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808882A_ABST
    Figure CN120808882A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, in particular to an application of a deep learning technology in drug research and development, and particularly relates to a virtual screening algorithm based on a gene expression profile and comparative learning, which comprises the following steps: on the basis of comparative learning, defining a drug and a matched expression profile as positive samples and other pairs in batches as negative samples; a double-feature encoder is adopted, and the cosine similarity is used as a scoring function to carry out similarity measurement. And the two loss function training models are combined, so that the performance of virtual drug screening is improved. In a word, the method marks an important step of virtual drug screening.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of deep learning, and is an application of deep learning technology in drug research and development, in particular to a virtual screening algorithm based on gene expression profiles and contrast learning. BACKGROUND

[0002] Drug research and development is high in investment, high in risk and long in cycle. Traditional drug research and development experiments are high in cost. For example, high-throughput screening technology is used to screen out promising bioactive compounds from a determined compound library. A low-cost computing method appears, virtual screening is used to screen out compounds with high possibility of binding to a target target from a super large compound library. There is structure-based virtual screening, but it needs to be matched with or protein 3D structure information, and it is difficult to capture the conformational changes of dynamic binding sites. The commonly used technology is molecular docking, but the key step needs to determine the scoring function for screening, the scoring function is very complex, and a large amount of computing resources are needed. There is also a ligand-based virtual screening method, which is highly dependent on molecular structure similarity and has limited performance in new compounds. With the development of machine learning and deep learning technologies, machine learning methods have been successfully applied to computer-aided drug design, but the calculation accuracy is low, and a single modality data is used, which is not suitable for complex biological scenes. With the development of deep learning technology, methods based on convolutional neural networks and recurrent neural networks have been successfully used for drug screening, but they depend on reliable data labels. SUMMARY

[0003] The purpose of the application is to solve the problems of difficult data label acquisition and time-consuming screening in the prior art. We propose a virtual screening algorithm based on gene expression profiles and contrast learning.

[0004] The idea of the application is based on contrast learning, which realizes the screening of drug-like molecules by learning the joint representation of gene expression profiles and drug molecules. 1) Data preparation, collect molecular data set and corresponding gene expression profile data set after drug perturbation under cell line. 2) Drug molecules as the input of drug encoder, and drug perturbed gene expression profile and drug acting cell line as the input of expression profile encoder. 3) The drug encoder adopts GAT to extract the features of the drug molecules, and obtains the drug embedding vector. 4) The gene expression profile encoder adopts multilayer perception to extract the expression profile features, and obtains the expression profile embedding vector. 5) Sample generation mechanism, batch sample construction, drug and its corresponding gene expression profile as positive sample, and the rest of the batch is paired as negative sample. 6) Model architecture fuses two feature encoders, and uses a similarity function as a feature pairing scoring function. 7) Joint use of two kinds of loss functions, expression profile to drug loss and drug to expression profile loss.

[0005] The specific technical scheme of the present application is: a virtual screening algorithm based on gene expression profile and contrast learning, comprising the following steps:

[0006] The data required by the model are drug molecules, expression profiles disturbed by drugs and corresponding cell lines, gene knockout expression profiles of target proteins, and known ligand molecule library of target proteins.

[0007] The drug encoder receives drug molecules as input, extracts drug molecule features through GAT, and obtains drug embedding vectors.

[0008] The model fuses the two feature encoders to calculate the cosine similarity of the drug embedding vectors and the expression profile embedding vectors.

[0009] The present application has the following beneficial effects:

[0010] 1. A new contrast learning framework is used, two feature encoders are used to learn the joint representation of gene expression profiles and molecules, and cross-modal matching between drugs and gene expression is realized.

[0011] 2. The present application designs a bidirectional contrast learning loss, which can better learn the matching between drugs and expression profiles by jointly using two losses.

[0012] 3. The present application has the advantages of fast screening speed, high efficiency and good performance. DETAILED DESCRIPTION

[0013] Figure 1 The model algorithm framework diagram is shown in Figure 1;

[0014] Figure 2 The ROC curve of the model is shown in Figure 2; DETAILED DESCRIPTION

[0015] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments.

[0016] Referring to Figure 1 A virtual screening algorithm based on gene expression profile and contrast learning, comprising the following steps:

[0017] S1: Data collection: Adopt the drug molecule and its corresponding drug perturbation expression profile dataset of cell lines to divide the training set, validation set and test set in the ratio of 6:2:2.

[0018] S2: Input drug molecule SMILES into drug encoder. Input cell lines and corresponding drug perturbation expression profile into gene expression profile encoder.

[0019] S3: Drug encoder adopts double-layer GAT, and the last layer is global max pooling layer. Expression profile encoder adopts three-layer fully connected layer.

[0020] S4: The model adopts cosine similarity as the scoring function. In a batch, the drug and its expression profile after drug action as positive sample pair, and the rest one-to-one corresponding drug expression profile pair as negative sample pair. The cosine similarity formula is as follows:

[0021]

[0022] Wherein represents the expression profile, represents the drug, and cell is the corresponding drug-affected cell line. θ and are the expression profile encoder and drug encoder and their corresponding parameters, respectively.

[0023] S5: The model jointly uses two loss functions for contrastive learning, expression profile to drug loss and drug to expression profile loss. The expression profile to drug loss function is:

[0024]

[0025] Wherein x i is the drug-perturbed expression profile, s i is the drug, and x j is all expression profiles. Its meaning is the possibility of ordering the given expression profile before other expression profiles.

[0026] The drug to expression profile loss is:

[0027]

[0028] Wherein x i is the drug-perturbed expression profile, s i is the drug, and s j is all drugs. Its meaning is the possibility of ordering the given drug before other drugs.

[0029] The total loss is:

[0030]

[0031] Network model:

[0032] Drug encoder: extract features from drug molecules.

[0033] Expression profile encoder: combine cell line encoding to extract features from perturbed expression profiles.

[0034] Scoring function: cosine similarity as positive and negative sample scoring function.

[0035] Contrastive learning framework: joint two loss functions, drug to expression profile loss and expression profile to drug loss.

[0036] In this embodiment, a total of 46066 small molecule drugs are used in the expression profiles of 42 cell lines. Figure 2 The prediction performance of the model is good.

[0037] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art within the technical range disclosed by the present application, according to the technical scheme and the inventive concept of the present application, equivalent replacement or change, should be covered within the protection scope of the present application.

Claims

1. A virtual screening algorithm based on gene expression profiling and comparative learning, characterized by: Including steps: 1) Data screening and preparation: The data required for the model are drug molecules, drug-perturbed expression profiles and corresponding cell lines, gene knockout expression profiles of target proteins, and a library of known ligand molecules for target proteins. 2) Drug molecules are input into the drug encoder, and the model extracts drug molecular features through the graph attention network (GAT). The drug-perturbed expression profile and the corresponding cell line are input into the expression profile encoder, and the model extracts expression profile features through a multi-layer perceptron. 3) The model uses two feature encoders to calculate the similarity of the two obtained embedding vectors (drug embedding vector and expression profile embedding vector) and uses it as the pairing score function. 4) Using a contrastive learning framework, two loss functions are used jointly to optimize the scoring function. In step 1), only small molecule drugs are considered. The perturbation expression profile was obtained after 24 h of drug exposure at 10 μM.

2. A virtual screening algorithm based on gene expression profiling and comparative learning according to claim 1, characterized in that: In step 2), the drug molecule encoder architecture uses GAT, ReLU activation function, and global maximum pooling layer. The expression spectrum encoder architecture uses a multi-layer perceptron.

3. The virtual screening algorithm based on gene expression profiling and comparative learning according to claim 1, characterized in that: In step 3), the similarity is cosine similarity.

4. The virtual screening algorithm based on gene expression profiling and comparative learning according to claim 1, characterized in that: In step 4), the comparative learning framework is that the positive sample pair is the drug and its corresponding cell line perturbation expression profile, and the negative sample pair is other drug and expression profile pairs in the batch.

5. The virtual screening algorithm based on gene expression profiling and comparative learning according to claim 1, characterized in that: In step 4), two contrastive loss functions are used jointly: drug-to-encoder loss and expression profile-to-drug loss.