Method for predicting disturbance reaction of single-cell drug
By constructing a single-cell drug perturbation prediction model, integrating drug properties and cell state information, and utilizing causal classifiers and optimal transport theory, the problem of insufficient analysis of drug action mechanisms in existing technologies is solved, achieving efficient prediction and improved interpretability of single-cell drug perturbation responses.
Patent Information
- Application Number
- CN202511571384.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-10-30
AI Technical Summary
Existing technologies are insufficient for precisely analyzing drug action pathways at the single-cell level, and are inadequate in predicting the effects of unknown drugs on specific cells and understanding drug action mechanisms, failing to effectively capture heterogeneous cellular responses.
A single-cell drug perturbation prediction model was constructed. Single-cell data were integrated through a drug attribute adapter and a causal classifier to generate perturbation embedding vectors and perform encoding and decoding processing. The potential distribution of drugs and cells was separated by optimal transport theory to predict the transcriptome expression state after drug perturbation.
It significantly improves the predictive ability of single-cell drug perturbation responses, enhances the interpretability of drug effects and training efficiency, and demonstrates excellent generalization ability on multiple datasets, especially in the task of predicting unknown drugs and pathways.
Smart Images

Figure CN121459913A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of single-cell drug perturbation, and particularly relates to a single-cell drug perturbation response prediction method. BACKGROUND
[0002] Single-cell drug perturbation response research uses single-cell technology to deeply explore the mechanism of drug action on cells and its influence. Traditional drug response research relies on population-level analysis methods, which cannot reveal the heterogeneity characteristics between cells. However, with the development of single-cell technology, this limitation has been overcome - especially the rapid development of high-throughput screening (HTS) and single-cell RNA sequencing (scRNAseq) technology, which enables researchers to analyze the influence of drugs on the transcriptome in different microenvironments (such as various drug dose combinations). Capturing the subtle changes of drug-induced cells or cell lines through single-cell RNA sequencing lays an important foundation for understanding cell heterogeneity and drug action mechanism.
[0003] However, the high dimensionality and strong heterogeneity of single-cell data pose new computational challenges for analyzing drug perturbation response. The rapid development of computational methods provides a breakthrough solution. Among them, deep learning has become a core tool with its powerful pattern recognition ability: through nonlinear mapping and multi-layer feature extraction technology, it can accurately capture the subtle differences between complex cells. Through biological space analysis, this method reveals the personalized influence of drug perturbation on heterogeneous cells, and has become the mainstream computational paradigm for analyzing single-cell drug perturbation response. In recent years, the academic community has developed various modeling methods: models based on encoder-decoder architecture build the basic computational framework; optimal transport theory further expands cell correlation analysis (realizes effective pairing of observed cells and unobserved cells); with further research, deep generative models have attracted attention because of their ability to capture multiple nonlinear relationships in gene expression and cell heterogeneity.
[0004] Although the above methods have promoted the development of drug research and improved the prediction ability of single-cell drug perturbation response, there are still significant limitations between experiments and computational methods: at the experimental level, it is difficult to finely analyze drug action pathways at the single-cell level, and the ability to effectively screen potential effective drugs is still insufficient; at the computational level, the complexity of drug perturbation (such as the interwoven relationship between direct effects and indirect influences) has not been fully analyzed, and there is still room for improvement in capturing heterogeneous cell responses. SUMMARY
[0005] The present application aims at the above-mentioned deficiencies in the prior art, and provides a single-cell drug perturbation reaction prediction method to solve the problems that the existing method is still difficult to fully analyze the mechanism of drug action in cells, and cannot capture the heterogeneous reaction of cells to different drugs, and the current research exploration is still insufficient in the key fields of predicting the influence of unknown drugs on specific cells, understanding the mechanism of drug action, and researching cell specificity.
[0006] To achieve the above-mentioned purpose, the technical solution adopted by the present application is: A single-cell drug perturbation reaction prediction method, comprising the following steps: S1, constructing a single-cell data set; S2, constructing a single-cell drug perturbation prediction model based on a drug attribute adapter and a causal classifier; S3, generating a perturbation embedding vector through the drug attribute adapter for the drug attribute data in the single-cell data set; S4, integrating the perturbation embedding vector with the undisturbed single-cell transcriptome expression profile data in the single-cell data set to obtain input data; S5, inputting the input data into the single-cell drug perturbation prediction model for encoding and decoding processing to predict the transcriptome expression state after drug perturbation.
[0007] Further, in S1, the single-cell data set is represented as: , In the formula, represents the single-cell data set; represents the transcriptome expression profile data; represents a drug attribute set; represents a sample; represents the maximum number of samples; represents a drug structure attribute; represents a drug dose attribute; represents a cell type.
[0008] Further, S3 specifically comprises: After standardizing the drug structure attribute , the drug dose attribute , and the cell type , the input vector of the drug attribute adapter is integrated, After deep neural network processing of the input vector , a perturbation embedding vector is generated in the latent space, which is specifically represented as: In the formula, represents the perturbation embedding vector; denotes a deep neural network; denotes a weight matrix of the drug property adapter layer during training; denotes a drug linear network function; denotes a hidden layer during training, denotes a trainable parameter; denotes a bias vector of the drug property adapter layer during training.
[0009] Further, the total loss function of the drug property adapter is: wherein, denotes a total loss of the drug property adapter; denotes a target data loss; denotes a weight hyperparameter of the divergence regularized loss; denotes a divergence regularized loss.
[0010] Further, the divergence regularized loss is denoted as: wherein, denotes a Kullback-Leibler divergence; is a normal distribution; denotes a mean of the latent variable; denotes a variance of the latent variable; denotes an identity matrix; the target data loss is denoted as: wherein: wherein, denotes an observed gene expression count; denotes a cell index or a sample index, denotes a gene index, denotes a number of genes; denotes a probability mapping function, denotes a sigmoid activation function, denotes a weight matrix for the probability mapping, denotes a bias vector for the probability mapping; denotes a distribution parameter of a gene j ; and denotes a Gamma function; denotes a factorial.
[0011] Further, the S5 specifically comprises: The input data is processed by the encoder in the single-cell drug perturbation prediction model, which encodes drug properties and single-cell transcriptome expression profile data into latent vectors. The causal classifier will classify the latent vectors The data is divided into two subspaces. After decoding, the two subspaces are processed by the decoder to output the transcriptome expression state after drug perturbation.
[0012] Furthermore, based on optimal transport theory, the causal classifier maps the cell state distribution of the control group and the drug-treated group, calculates the causal effect, and separates the drug... With cell lines The potential distribution is then analyzed, and the causal effect matrix is finally output. Among them, latent variable data of the control group were obtained. Potential variables of drug treatment group Define the optimal transmission cost: In the formula, Indicates the optimal transmission distance; Indicates the infimum; Indicates the transmission plan; Let γ represent the set of all transmission plans that satisfy the marginal distribution constraint; This represents an element in the transport plan matrix, specifically referring to the coupling probability quality from the i-th control group cell to the j-th treatment group cell; This represents the set of potential embedding vectors for the control group cells; This represents the set of potential embedding vectors for the cells in the drug-treated group; The causal effect matrix is represented as follows: in: In the formula, Represents the causal effect matrix; This represents a drug causal classifier; This represents a cell line causal classifier; This indicates the predicted causal effect of the drug on cell i; This indicates the predictive causal effect of cell line i itself; σ ( · ) represents the probability mapping function; For shared parameters Multilayer neural network transformation.
[0013] Furthermore, the loss function of the causal classifier is expressed as: in: wherein, denotes the total loss function of the causal classifier; denotes the classification loss; denotes the weight coefficient of the optimal transport loss term; denotes the optimal transport distance; denotes the weight coefficient of the L2 regularization term; denotes the true label of the causal effect of the drug on cell i; denotes the true label of the causal effect of the cell line i itself.
[0014] Further, in the S5, the cell state in the real environment is simulated by adding Gaussian noise, and then the reconstruction loss between the undisturbed transcriptome expression profile data and the disturbed transcriptome expression profile data is calculated, which is represented as: wherein, denotes the reconstruction loss; denotes the mean of the expression distribution reconstructed by the decoder for gene j in cell i; denotes the variance of the expression distribution reconstructed by the decoder for gene j in cell i.
[0015] The single-cell drug perturbation response prediction method provided by the present application has the following beneficial effects: 1、The present application can effectively integrate drug information from perturbed single-cell transcriptome data, thereby robustly capturing drug treatment effects and cell-specific responses. By combining causal reasoning and optimal transport strategies, the present application can distinguish between direct drug effects and indirect cell influences, thereby enhancing the explainability of the potential dimensions of single-cell drug perturbation. A large number of experiments on L1000 and sci-Plex3 data sets show that scDPR exhibits excellent generalization ability for unseen drugs and pathways. In addition, scDPR significantly improves the training efficiency, and compared with the latest method (PRnet), the training time is reduced by 56.42%.
[0016] 2、Compared with the baseline model and related models, the scDPR of the present application not only shows excellent performance in drug effect and cell line separation, but also achieves a high level in drug perturbation explainability, understanding of cell heterogeneity and training efficiency.
[0017] 3、The drug attribute adapter of the present application separates the cell line, drug and dose from the transcriptome information, which is used to guide the scDPR to learn the drug perturbation pattern during the training process.
[0018] 4、The present application uses a causal reasoning method to explore the drug perturbation pattern and evaluate the influence and action mode of different drugs on cell lines. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 Figure 1 is a framework diagram of a single-cell drug perturbation prediction model (scDPR framework) of an embodiment of the present application, wherein, Figure 1 (a) in Figure 1 is the overall architecture of scDPR, Figure 1 (b) in Figure 1 is a drug property adapter architecture diagram, Figure 1 (c) in Figure 1 is a causal classifier architecture diagram.
[0020] Figure 2 Figure 2 is a graph showing the effect of scDPR in predicting OOD data in an embodiment of the present application, wherein Figure 2 (a) in Figure 2 is a comparison of the average gene expression prediction R2 score of unseen compounds (OOD) in two data sets; Figure 2 (b) in Figure 2 is a comparison of the average gene expression prediction R2 score of unseen pathways (OOD) on the sci-Plex3 data set.
[0021] Figure 3 Figure 3 is a potential space UMAP (188 drugs, 3 cancer cell lines) of different methods on the sci-Plex3 data set in an embodiment of the present application, which from left to right are: original untreated data, scDPR and chemCPA best prediction results of coloring effect.
[0022] Figure 4 Figure 4 is a multi-drug treatment pathway heat map of an embodiment of the present application, showing the relevance of flavonoid amide alcohol, TAK-901, and tan spirochaete to signal pathways, diseases (prostate cancer, glioma), and cellular processes (endocytosis, cell carcinogenesis); the color gradient represents statistical significance (-log10(P)), and the larger the value, the stronger the correlation.
[0023] Figure 5 Figure 5 is a comparison of training time in an embodiment of the present application, under the same experimental conditions, based on different perturbation division methods, the r2 score and training time of different methods in the training process are compared; wherein, Figure 5 (a) in Figure 5 is the division and training of the data set according to different drug perturbations; Figure 5 (b) in Figure 5 is the division and training of the data set according to different biological pathways.
[0024] Figure 6 Figure 6 is a flowchart of a single-cell drug perturbation response prediction method of the present application. DETAILED DESCRIPTION
[0025] The specific embodiments of the present application are described below to facilitate the understanding of the present application for those skilled in the art, but it should be clear that the present application is not limited to the scope of the specific embodiments, and for those skilled in the art, it is obvious that various changes are within the spirit and scope of the present application defined and determined by the appended claims, and all the inventions utilizing the concept of the present application are within the scope of protection.
[0026] Embodiment 1 The single-cell drug perturbation response prediction method of the present embodiment can integrate drug properties and cell state information to analyze single-cell drug perturbation responses. This method not only significantly improves the ability to analyze drug action mechanisms and cell-specific responses, but also exhibits better prediction performance than existing methods on multiple data sets. Especially in the generalization task involving unknown drugs and pathways: the embedding vectors learned by the single-cell drug perturbation prediction model (scDPR) have good interpretability and can effectively distinguish different cell lines, providing intuitive support for heterogeneity analysis; at the same time, its efficient training ability makes it more suitable for large-scale drug screening tasks, as shown in Figure 1 and Figure 6 , which specifically includes the following contents: S1, constructing a single-cell data set; The single-cell data set is represented as: , In the formula, represents a single-cell data set; represents n-dimensional transcriptome expression profile data; represents a set of drug properties; represents a sample; represents the maximum number of samples; When constructing a single-cell RNA sequencing (scRNA-seq) drug perturbation model (single-cell drug perturbation prediction model), the structural properties and dosage properties of the drug need to be considered comprehensively, i.e.: drug structural properties and drug dosage properties and cell types
[0027] S2, constructing a single-cell drug perturbation prediction model based on a drug property adapter and a causal classifier; Referring to (a) in Figure 1 , the single-cell drug perturbation prediction model includes a drug property adapter, an encoder, a decoder, and a causal classifier; The drug property adapter can effectively integrate drug information from perturbed single-cell transcriptome data, thereby robustly capturing drug treatment effects and cell-specific responses; the causal classifier combined with the optimal transmission strategy can distinguish direct drug effects from indirect cell effects, thereby enhancing the explainability of the potential dimensions of single-cell drug perturbation.
[0028] S3, drug property data in the single-cell data set generates a perturbation embedding vector through a drug property adapter; In the prior art, the emergence of single-cell RNA sequencing technology makes it possible to obtain high-resolution data, which is crucial for studying cell responses triggered by drug perturbation. However, its high dimensionality, sparsity and cell heterogeneity pose challenges to modeling the causal effects of drug-cell interactions. Existing methods often directly use drug molecules or transcriptome expression data, making it difficult to effectively integrate drug properties (such as molecular structure and dosage) and cell type information. This makes the feature description of complex perturbation responses insufficient, especially when predicting unknown drugs.
[0029] Based on this, the embodiment designs a drug property adapter, as shown in (b) of the formula (1) in the specification, which is used to capture the correlation between drug molecular structure, dosage information, cell type characteristics and single-cell transcriptome data, thereby improving the modeling accuracy of drug perturbation response. Figure 1
[0030] Specifically, for a sample i , the input parameters drug structure property , drug dosage property , and cell type are standardized and integrated into the input vector of the drug property adapter . After the input vector is processed by a deep neural network, a perturbation embedding vector is generated in the latent space, which is specifically represented as: In the formula, z represents the perturbation embedding vector; f represents the deep neural network; W represents the weight matrix of the drug property adapter layer during the training process; g represents the drug linear network function; h represents the hidden layer during the training process; b represents the trainable parameter; and b represents the bias vector of the drug property adapter layer during the training process.
[0031] In the process of generating the perturbation embedding in the latent space, divergence regularization is introduced to ensure the stability of the latent variable distribution, and it is assumed that Follows a standard normal distribution: In the formula, Indicates the Kullback-Leibler divergence; It follows a normal distribution; This represents the mean of the latent variables; Represents the variance of the latent variables; Represents the identity matrix; Due to the sparsity of single-cell drug perturbation data, this embodiment constructs a target data loss function based on the probability density of a negative binomial distribution to achieve better convergence of the drug adapter performance; wherein, the probability density of the negative binomial distribution and the loss function are as follows: In the formula, This represents the observed gene expression count; Indicates a cell index or sample index. Represents a gene index. Indicates the number of genes; Represents the probability mapping function. This represents the Sigmoid function. This represents the weight matrix used in the probability mapping layer. This represents the bias vector used in the probability mapping layer; Indicates gene j The distribution parameters; ! represents the Gamma function; ! represents factorial; The total loss function of the drug property adapter is: In the formula, This represents the total loss of the drug property adapter; Indicates the loss of target data; The weight hyperparameters representing the divergence regularization loss; This represents the divergence regularization loss.
[0032] S4. Integrate the perturbation embedding vector with the unperturbed single-cell transcriptome expression profile data in the single-cell dataset to obtain the input data; Specifically, the model's input includes single-cell transcriptome data: upper channel Using unperturbed single-cell transcriptome expression profile data, the lower channel Then, covariate information (drug type, dosage, and cell type) processed by the drug property adapter is used. By integrating these two channels, the output is... The input data of the encoder-decoder is the perturbed transcriptomic expression profile distribution.
[0033] S5, inputting the input data into the single-cell drug perturbation prediction model for encoding-decoding processing to predict the transcriptomic expression state after drug perturbation; The specific process is: inputting the input data The drug attribute and single-cell transcriptomic expression profile data are encoded into latent vectors by the encoder in the single-cell drug perturbation prediction model The causal classifier divides the latent vectors into two subspaces The two subspaces and The two subspaces and are decoded by the decoder, and the decoder maps these components to the transcriptomic expression state , , after drug perturbation i .
[0034] In the model optimization training process, the learning ability of the framework is optimized by calculating the reconstruction loss between the unperturbed transcriptomic expression profile data and the perturbed transcriptomic expression profile data, so as to realize the prediction of the transcriptomic expression profile under different perturbation environments. By adding Gaussian noise to simulate the cell state in the real environment, the perturbation conditions affecting the transcriptomic expression are learned more effectively.
[0035] wherein the reconstruction loss function is as follows: In the formula, represents the reconstruction loss; represents the mean of the gene expression distribution reconstructed by the decoder; represents the variance of the gene expression distribution reconstructed by the decoder.
[0036] By converting the transcriptional response prediction into a distribution problem based on different perturbation conditions, a causal classifier is used to learn the state distribution of the cell line under different drug perturbation conditions. This method can not only model the perturbation pattern, but also effectively analyze the mechanism of action (MoA) in multiple scenarios and quantify the perturbation effect of unseen drugs. Specifically, for the causal classifier; In drug perturbation analysis, traditional correlation-based methods are difficult to distinguish the direct causal effect of drugs on cell state and indirect confounding association. The lack of such causal relationship may lead to false correlation, confounding bias and mechanism misjudgment.
[0037] Therefore, in order to infer the causal effects of drug perturbations on single-cell transcriptomes through the distribution of potential spatial perturbations. To quantify the interaction between drugs and cell states and elucidate their mechanisms of action, and to design, for example... Figure 1 The causal classifier shown in (c) is based on optimal transport (OT) theory. It maps the cell state distributions of the control and drug-treated groups, calculates causal effects, and separates the drug... With cell lines The classifier calculates the optimal transport path between the latent variable distributions of the control and drug-treated groups to capture cell state changes induced by drug perturbation. This is achieved by acquiring the latent variable data from the control group. and latent variables of the treatment group The optimal transmission cost is defined as: In the formula, Indicates the optimal transmission distance; Indicates the infimum; Indicates the transmission plan; Let γ represent the set of all transmission plans that satisfy the marginal distribution constraint; This represents an element in the transport plan matrix, specifically referring to the coupling probability quality from the i-th control group cell to the j-th treatment group cell; This represents the set of potential embedding vectors for the control group cells; This represents the set of potential embedding vectors for the cells in the drug-treated group; The causal classifier outputs a causal effect matrix. ,in N d represents the number of drug types, and M represents the number of cell states. The elements of this matrix are derived from a drug causal classifier. and cell causal classifier Joint generation: Therefore, the causal effect matrix is represented as follows: In the formula, Represents the causal effect matrix; This represents a drug causal classifier; This represents a cell line causal classifier; This indicates the predicted causal effect of the drug on cell i; This indicates the inherent causal effect of cell line i itself; σ ( · ) represents the probability mapping function; For shared parameters multilayer neural network transformation.
[0038] The loss function of the causal classifier, i.e., the drug-target binding, optimal transport loss, and regularization term, is denoted as: The classification loss is denoted as: wherein, denotes the total loss function of the causal classifier; denotes the classification loss; denotes the weight coefficient of the optimal transport loss term; denotes the optimal transport distance; denotes the weight coefficient of the L2 regularization term; denotes the true label of the causal effect of the drug on cell i; denotes the true label of the causal effect of the cell line i itself.
[0039] Embodiment 2 This embodiment is based on the content in Embodiment 1, and corresponding experiments and analysis of the implementation results are carried out, which specifically include the following contents: 1. Experimental setup; The L1000 [1] and sci-Plex3 [2] datasets were selected to evaluate the impact of drug perturbation response and analyze the drug action mechanism.
[0040] Dataset: The L1000 dataset contains RNA transcriptome expression profiles of more than 1.3 million cells, covering the expression information of 10,000 genes. Its core advantage is that it records the transcriptome dynamic changes of 70 cell lines under the perturbation of 17,000 compounds, and each perturbation is systematically observed at more than 2,000 concentration gradients. This large-scale, multi-dimensional design makes it a benchmark research dataset for exploring the mechanism of action (MoA) of drugs.
[0041] Another key dataset, sci-Plex3, focuses on single-cell resolution research. This dataset contains about 500,000 single-cell RNA sequencing (scRNA-seq) samples, systematically recording the effects of 188 small molecule drugs on three human cancer cell lines under single compound perturbation at four concentration conditions. What is particularly important is that the Srivatsan team annotated all drugs in this dataset with 22 known mechanisms of action (or pathways). These annotations are directly related to the biological principles of drug therapeutic effects at the molecular or cellular level, providing important evidence for mechanism interpretation and verification of the model.
[0042] Data Preparation: To evaluate the performance of scDPR in the tasks of unknown drug prediction and pathway identification, a rigorous split-data strategy was adopted: according to the perturbation attributes (compounds and pathways), the dataset was divided into training set, test set and out-of-distribution (OOD) validation set. Among them, the out-of-distribution validation set is specifically used to evaluate the generalization ability of the model under two unobserved scenarios: (1) drug perturbation that does not appear in both training set and test set; (2) pathway perturbation that does not appear in both training set and test set
[13] . To comprehensively verify the robustness of the model, training and testing were performed on two different resolution datasets L1000[1] and sci-Plex3[2] respectively.
[0043] Evaluation Index: In the embodiments, to quantitatively evaluate the performance of scDPR on single-cell perturbation dataset, the commonly used measurement method was adopted, and the average coefficient of determination (R² score) and mean square error (MSE) between the true value and the predicted value of the gene expression of the perturbed cells in the test set were compared.
[0044] wherein, represents the true gene expression profile of the non-perturbed control cell ; represents the gene expression profile of the cell predicted by the model after drug perturbation; represents the total number of cells in the test set; To evaluate the scDPR method for predicting drug treatment effect, two indicators were used: individual treatment effect (TEi) and average treatment effect (TEavg) to quantify the treatment effect of the drug. ITE ATE
[0045] wherein, represents the actual observed gene expression profile of the cell after drug treatment in the real world; represents the total number of cells in the drug treatment group.
[0046] Comparative method: To qualitatively analyze the advantages of scDPR in predicting unknown drugs or pathways and its excellent characteristics in training time, it is necessary to conduct a competition among similar models. Therefore, the prediction experiments of unobserved drugs were carried out on sci-Plex3 [2] and L1000 [1] datasets, and the prediction experiments of unknown pathways were carried out on sci-Plex3 dataset, and the results were compared with scGen [4], CPA
[11] , chemCPA
[12] and PRnet
[13] and other models.
[0047] 2. scDPR performs well in predicting unobserved perturbation responses, as follows: The comprehensive performance comparison of scDPR and other single-cell perturbation response prediction models (including scGen [4], CPA
[11] , chemCPA
[12] and PRnet
[13] ) was carried out on multiple data sets and perturbation identification tasks. All models were run in a unified hardware environment (using the same CPU and GPU configuration) to ensure experimental fairness. The data were divided into training set, test set and distributed validation set in a fixed ratio of 6:2:2. During training, the maximum training rounds of all models were uniformly set to 100 rounds, and the early termination mechanism was used - when the validation loss did not decrease for 20 consecutive rounds, the training was automatically terminated. The Adam optimizer was used for training, with a fixed learning rate of 1e-3 and a weight decay coefficient of 1e-8 to ensure the consistency of hyperparameter configuration. In all performance comparison experiments and ablation experiments, the mean coefficient of determination R ² was selected as the core evaluation index to quantify the model performance.
[0048] As shown in Table 1, in all gene perturbation response prediction tasks of two datasets, scDPR achieved the best R ² and MSE results. To further evaluate the model's generalization ability to unobserved compounds and unobserved pathways in out-of-distribution scenarios, targeted experiments were carried out (results shown in Figure 2). The experiments showed that scDPR exhibited the best prediction performance for unknown perturbation responses on both sci-Plex3 and L1000 benchmark datasets. Taking the sci-Plex3 dataset as an example, in the test of different unknown perturbations, the R ² scores (0.8994, 0.9101) and mean square error values (0.0326, 0.0323) of scDPR were significantly better than all comparison methods.
[0049] Table 1: Performance comparison of scDPR and other methods ( R 2 and MSE) Furthermore, the contributions of each core module of scDPR were quantified through systematic knockout experiments (results are shown in Table II). Quantitative analysis showed that: (1) the drug property adapter module significantly improved the model's ability to characterize drug-cell interactions by effectively integrating drug molecular features; (2) the causal classifier module greatly enhanced the model's causal inference performance for drug perturbation effects; and (3) Gaussian-based noise modeling could more accurately capture the expression variation characteristics of single cells under drug perturbation. The synergistic effect of these key components enabled the model to more accurately simulate the dynamic changes of single-cell transcriptomes in complex drug environments, thereby achieving robust learning of drug action mechanisms.
[0050] Table 2: Comparison of ablation experiments of the scDPR core module ( R 2 and MSE) 3. scDPR learns interpretable latent embeddings, as follows: To accurately characterize the heterogeneity of gene expression in single cells under drug intervention, scDPR learns to reconstruct the disturbed transcriptome state and extracts biologically meaningful feature embeddings from the latent space. For example... Figure 3 As shown in the UMAP dimensionality reduction results based on the sci-Plex3 dataset, compared to the original unprocessed cell distribution, the low-dimensional embeddings learned by scDPR exhibit a significant cell line clustering structure—that is, samples from the same cell line are highly clustered in the latent space.
[0051] In contrast, existing methods such as chemCPA
[12] are difficult to effectively distinguish between different cell lines. scDPR not only achieves clearer cell line segmentation but also significantly enhances intra-cell cohesion among samples. This result clearly demonstrates scDPR's ability to decouple the inherent characteristics of cell lines even under complex drug interference.
[0052] 4. Verify the drug perturbation mechanism; To verify the effectiveness of scDPR in resolving the mechanism of drug action, three candidate drugs were selected: xanthohumol, TAK-901 and tanespimycin - the mechanisms of action of these drugs have been clearly and reliably verified by Sriwastava's research. In the experimental stage, the perturbation effects of the above drugs in three cancer cell lines (A549, MCF7 and K562) were quantitatively evaluated using scDPR, and the molecular pathways affected by the drugs were further revealed by KEGG pathway enrichment analysis. The specific operation process is as follows: scDPR first models the original single-cell transcriptome data to generate single-cell drug perturbation simulation data; based on this simulation data, the cell response score (effect score) after drug perturbation is calculated to quantify the direct influence of the drug on the cell phenotype (see Table III for results). It is worth noting that when the drug dose is 0 (i.e. no drug perturbation is applied), all effect scores are 0, which provides key verification for the baseline validity of the score.
[0053] At the same time, the KEGG enrichment analysis heatmap in Figure 4 is also constructed based on the single-cell drug perturbation simulation data generated by the model. Through this heatmap, the perturbation mechanisms of the three drugs can be systematically and intuitively analyzed. The experimental results show that xanthohumol and tanespimycin exhibit significant enrichment characteristics in cancer-related pathways, which is highly consistent with their high effect scores in the corresponding cancer cell lines. This indicates that these two drugs have potential application value in the treatment of breast cancer (MCF7 cell line) and leukemia (K562 cell line). In contrast, TAK-901 has relatively low pathway enrichment, reflecting its higher specificity in targeting mechanisms. Based on the above results, it can be seen that single-cell drug perturbation (scDPR) can effectively analyze the mechanism of drug action by mapping the distribution characteristics of different perturbations.
[0054] 5. Time complexity analysis; In addition to significantly improving prediction performance, scDPR also achieves a breakthrough in training efficiency. To ensure fairness, all models are set to a maximum of 100 training rounds, and an early termination mechanism is introduced based on the loss change curve - when the validation set loss value does not significantly decrease for multiple rounds in a row, the system will automatically terminate training to avoid overfitting and determine the best convergence point. As shown in Figure 5 Under uniform experimental settings, scDPR only takes 285.56 minutes (for compound tasks) and 439 minutes (for pathway tasks) to achieve optimal performance in training tasks for different perturbation attributes (compounds / pathways), with overall prediction accuracy superior to all baseline models.
[0055] Specifically, compared with the PRnet model, the training time of scDPR is shortened by 56.42%; compared with chemCPA and CPA, it is reduced by 69.17% and 26.31% respectively. Even compared with scGen with similar training time, scDPR still maintains a significant performance advantage. The above results confirm that scDPR can simulate single-cell drug perturbation response mechanism with super high efficiency, providing a tool with both precision and timeliness for large-scale drug mechanism research.
[0056] Although the specific embodiments of the application are described in detail with reference to the accompanying drawings, it should not be understood as limiting the scope of protection of the patent. Various modifications and variations made by those skilled in the art within the scope described in the claims are still within the scope of protection of the patent.
Claims
1. A method for predicting single-cell drug perturbation responses, characterized in that, Includes the following steps: S1. Construct a single-cell dataset; S2. Construct a single-cell drug perturbation prediction model based on drug property adapter and causal classifier; S3. Drug attribute data in single-cell datasets are used to generate perturbation embedding vectors through drug attribute adapters; S4. Integrate the perturbation embedding vector with the unperturbed single-cell transcriptome expression profile data in the single-cell dataset to obtain the input data; S5. Input the input data into the single-cell drug perturbation prediction model for encoding and decoding processing, and predict the transcriptome expression state after drug perturbation.
2. The method for predicting single-cell drug perturbation responses according to claim 1, characterized in that, In S1, the single-cell dataset is represented as follows: , In the formula, This represents a single-cell dataset; Represents transcriptome expression profile data; Represents a set of drug properties; Indicates a sample; Indicates the maximum number of samples; Indicates the structural properties of a drug; Indicates drug dosage attributes; Indicates cell type.
3. The method for predicting single-cell drug perturbation responses according to claim 1, characterized in that, S3 specifically includes: Drug structural properties Drug dosage attributes Cell types After standardization, the input vector is integrated into the drug property adapter. , input vector After processing by a deep neural network, a perturbation embedding vector is generated in the latent space, which is specifically represented as follows: In the formula, Represents the perturbation embedding vector; Represents a deep neural network; This represents the weight matrix of the drug property adapter layer during training; Represents a linear network function for drugs; This represents the hidden layer during the training process. Indicates trainable parameters; This represents the bias vector of the drug attribute adapter layer during training.
4. The method for predicting single-cell drug perturbation responses according to claim 3, characterized in that, The total loss function of the drug property adapter is: In the formula, This represents the total loss of the drug property adapter; Indicates the loss of target data; The weight hyperparameters representing the divergence regularization loss; This represents the divergence regularization loss.
5. The method for predicting single-cell drug perturbation responses according to claim 4, characterized in that, The divergence regularization loss Represented as: In the formula, Indicates the Kullback-Leibler divergence; It follows a normal distribution; This represents the mean of the latent variables; Represents the variance of the latent variables; Represents the identity matrix; The target data loss Represented as: in: In the formula, This represents the observed gene expression count; Indicates a cell index or sample index. Represents a gene index. Indicates the number of genes; Represents the probability mapping function. This represents the sigmoid activation function. This represents the weight matrix used for probability mapping. This represents the bias vector used for probability mapping; Indicates gene j The distribution parameters; ! represents the Gamma function; ! represents factorial.
6. The method for predicting single-cell drug perturbation responses according to claim 2, characterized in that, S5 specifically includes: The input data is processed by the encoder in the single-cell drug perturbation prediction model, which encodes drug properties and single-cell transcriptome expression profile data into latent vectors. The causal classifier will classify the latent vectors The data is divided into two subspaces. After decoding, the two subspaces are processed by the decoder to output the transcriptome expression state after drug perturbation.
7. The method for predicting single-cell drug perturbation responses according to claim 6, characterized in that, Based on optimal transport theory, the causal classifier maps the cell state distribution of the control group and the drug-treated group, calculates the causal effect, and separates the drug. With cell lines The potential distribution is then analyzed, and the causal effect matrix is finally output. Among them, latent variable data of the control group were obtained. Potential variables of drug treatment group Define the optimal transmission cost: In the formula, Indicates the optimal transmission distance; Indicates the infimum; Indicates the transmission plan; Let γ represent the set of all transmission plans that satisfy the marginal distribution constraint; This represents an element in the transport plan matrix, specifically referring to the coupling probability quality from the i-th control group cell to the j-th treatment group cell; This represents the set of potential embedding vectors for the control group cells; This represents the set of potential embedding vectors for the cells in the drug-treated group; The causal effect matrix is represented as follows: in: In the formula, Represents the causal effect matrix; This represents a drug causal classifier; This represents a cell line causal classifier; This indicates the predicted causal effect of the drug on cell i; This indicates the predictive causal effect of cell line i itself; σ ( · ) represents the probability mapping function; For shared parameters Multilayer neural network transformation.
8. The method for predicting single-cell drug perturbation responses according to claim 7, characterized in that, The loss function of the causal classifier is expressed as: in: In the formula, This represents the total loss function of the causal classifier; Indicates classification loss; The weighting coefficients represent the optimal transmission loss term; Indicates the optimal transmission distance; Represents the weight coefficients of the L2 regularization term; A true label representing the causal effect of a drug on cell i; This represents the true label of the causal effect of cell line i itself.
9. The method for predicting single-cell drug perturbation responses according to claim 5, characterized in that, In step S5, Gaussian noise is added to simulate the cell state in a real environment, and then the reconstruction loss between the unperturbed transcriptome expression profile data and the perturbed transcriptome expression profile data is calculated, which is expressed as: In the formula, Indicates the reconstruction loss; The decoder represents the mean of the expression distribution reconstructed by gene j in cell i; The decoder represents the variance of the expression distribution reconstructed by gene j in cell i.
Citation Information
Patent Citations
Single cell transcription reaction prediction algorithm based on cross-domain feature cross migration
CN119252336A
Single-cell drug causal gene discovery method and system based on anti-transfer learning
CN119446281A
Method for predicting transcriptional response to novel drug perturbation and virtual screening method and system
CN119763720A
Drug relocation method and system
CN120564891A
Designing Chemical or Genetic Perturbations using Artificial Intelligence
US20240047081A1