Method for predicting response of single cell gene expression to perturbation, model training method, system and terminal
Patent Information
- Application Number
- PCT/CN2025/086263
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-03-31
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025086263_01102026_PF_FP_ABST
Abstract
Description
Methods for predicting single-cell gene expression responses to disturbances, model training methods, systems, and terminals. Technical Field
[0001] This invention relates to the fields of biotechnology and artificial intelligence, and in particular to methods, model training methods, systems and terminals for predicting the response of single-cell gene expression to disturbances. Background Technology
[0002] With the rapid development of single-cell transcriptome sequencing (scRNA-seq) technology, researchers' understanding of cellular processes has been enhanced. They can now resolve the heterogeneity and dynamic state of individual cells in complex tissues at high resolution, providing unprecedented insights into human physiological functions by analyzing cellular heterogeneity and dynamic changes at the single-cell level. Despite these advances, comprehensively covering the responses of all possible perturbations under all cell types and conditions remains a highly challenging task. This is primarily due to the complexity of biological variability, the vast number of potential perturbations, and the high cost of experiments, making it difficult for researchers to obtain complete expression profiles across all cell types and conditions. Computational methods have become a mainstream alternative for studying cellular perturbation responses, utilizing machine learning and deep learning techniques to predict gene expression changes in cells under different perturbations, providing important tools for personalized medicine and genomics research. Currently, methods for predicting gene expression after single-cell perturbation can be broadly categorized into three types:
[0003] (1) Methods that do not decouple cell expression
[0004] These methods directly utilize gene regulatory network information or drug perturbation data for prediction without decoupling cellular expression information. For example, the DeepCE method predicts changes in gene expression by combining gene regulatory network information and drug perturbation data. However, the input to this method is the gene regulatory network, not the actual expression data of cells, which limits its ability to model cellular heterogeneity and makes it difficult to accurately predict the specific responses of different cell types under different perturbation conditions.
[0005] (2) Cell expression decoupling method based on latent spatial coding
[0006] These methods attempt to decouple cellular gene expression but fail to utilize additional external information. For example, scGen employs a variational autoencoder (VAE) to learn the latent representation of gene expression and calculates the mean expression change between the treatment and control groups to obtain a perturbation vector. However, this method, relying solely on mean expression values, may lead to significant prediction errors. scPARM uses optimal transport theory to find the optimal transport matrix in the latent space to predict perturbation responses for different cell types. Furthermore, CPA removes the influence of perturbations and cellular covariates in the latent space through adversarial loss to decouple cellular expression features. However, this method fails to effectively utilize cellular covariate and perturbation information during decoupling, relying solely on the adversarial module, resulting in limited decoupling effectiveness.
[0007] (3) Methods for decoupling cell expression using external information
[0008] This category of methods introduces additional external information during the decoupling of cell expression to improve prediction accuracy. For example, the VCI method, based on the variational Bayesian causal inference framework, utilizes the balance between individual cell characteristics and population gene expression distribution to construct a gene expression model under hypothesis perturbation, thus providing a statistically robust solution. However, this method has high computational complexity and high data quality requirements, limiting its application on larger-scale datasets.
[0009] It is evident that while existing methods can predict single-cell responses to perturbations to some extent, they suffer from the following limitations: insufficient decoupling ability of cell expression, weak generalization ability of predictions, inadequate utilization of external information, lack of biological rationality in prediction results, and a need to improve prediction accuracy.
[0010] Therefore, existing technologies still need to be improved and developed. Summary of the Invention
[0011] Based on the shortcomings of the prior art, the purpose of this invention is to provide a method, model training method, system and terminal for predicting the response of single-cell gene expression to disturbances, in order to solve the problems of insufficient cell expression decoupling ability, weak prediction generalization ability and need to be further improved in existing methods for predicting the response of single cells to disturbances.
[0012] The technical solution of the present invention is as follows:
[0013] In a first aspect, the present invention provides a method for training a prediction model of single-cell gene expression response to perturbation, wherein the prediction model includes an encoder, a latent space based on contrastive learning, and a decoder, and the method for training the prediction model includes the following steps:
[0014] Obtain gene expression data of single cells under known perturbation conditions and under known perturbation conditions, as well as cellular covariates of single cells;
[0015] Using the known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cell covariates of single cells, the encoder, the latent space based on contrastive learning, and the decoder are trained to obtain a trained single-cell gene expression response prediction model to perturbation.
[0016] Optionally, the steps of training the encoder, the contrastive learning-based latent space, and the decoder using the known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cellular covariates of single cells specifically include:
[0017] The known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cell covariates of single cells are input into the encoder, and then, after passing through the latent space based on contrastive learning, the first latent variable is obtained.
[0018] After combining the first latent variable with the known perturbation conditions, the data is input into the decoder to obtain the predicted single-cell gene expression value under the known perturbation conditions.
[0019] A loss function is obtained based on the predicted single-cell gene expression values under the known perturbation conditions and the single-cell gene expression data under the known perturbation conditions, and training is performed based on the loss function.
[0020] Optionally, the encoder includes a feedforward neural network, a self-attention mechanism, and a fully connected layer. The specific steps of training the encoder, the contrastive learning-based latent space, and the decoder using the known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cellular covariates of single cells specifically include:
[0021] The known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cell covariates of single cells are input into the encoder. After passing through the feedforward neural network, the first feature vector, the second feature vector, and the third feature vector are obtained respectively.
[0022] The first feature vector, the second feature vector, and the third feature vector are concatenated and input into a self-attention mechanism for fusion. The result is output through a fully connected layer to obtain a potential high-level feature vector.
[0023] The potential high-level feature vectors are input into the latent space based on contrastive learning to obtain the first latent variable;
[0024] After combining the first latent variable with the known perturbation conditions, the data is input into the decoder to obtain the predicted single-cell gene expression value under the known perturbation conditions.
[0025] A loss function is obtained based on the predicted single-cell gene expression values under the known perturbation conditions and the single-cell gene expression data under the known perturbation conditions, and training is performed based on the loss function.
[0026] Optionally, the step of inputting the latent high-level feature vector into the latent space based on contrastive learning to obtain the first latent variable specifically includes:
[0027] The latent high-level feature vectors are input into the latent space based on contrastive learning to form latent variables. Based on the equality of cellular covariates and the distance between latent variables in the latent space, positive and negative loss functions are introduced.
[0028] The positive and negative loss functions are summed to obtain the contrast loss function, and the first latent variable is output.
[0029] Optionally, the first latent variable is combined with the known perturbation conditions and input into the decoder to obtain the predicted single-cell gene expression value under the known perturbation conditions; a loss function is obtained based on the predicted single-cell gene expression value under the known perturbation conditions and the gene expression data of single cells under the known perturbation conditions, and the training step based on the loss function specifically includes:
[0030] The first latent variable is combined with the known perturbation conditions and input into the decoder to obtain the predicted value of single-cell gene expression under the known perturbation conditions. The reconstruction loss function is obtained based on the predicted value of single-cell gene expression under the known perturbation conditions and the gene expression data of single cells under the known perturbation conditions.
[0031] Simultaneously, it provides counterfactual perturbation conditions and gene expression data of single cells under counterfactual perturbation conditions. After combining the first latent variable with the counterfactual perturbation conditions, it inputs it into the decoder to obtain the predicted value of single-cell gene expression under counterfactual perturbation conditions. Based on the predicted value of single-cell gene expression under counterfactual perturbation conditions and the gene expression data of single cells under counterfactual perturbation conditions, it obtains the counterfactual loss function.
[0032] Training is performed based on the reconstruction loss function, the counterfactual loss function, and the contrastive loss function.
[0033] Optionally, the decoder is composed of a fully connected layer;
[0034] The known perturbation conditions include at least one of drug treatment perturbation conditions, gene perturbation conditions, and environmental perturbation conditions.
[0035] A second aspect of the present invention provides a method for predicting the response of single-cell gene expression to perturbation, comprising the following steps:
[0036] Acquire gene expression data of the single cell to be predicted without perturbation, preset perturbation conditions, and cellular covariates of the single cell to be predicted;
[0037] The gene expression data of the single cell to be predicted without perturbation, the preset perturbation conditions, and the cell covariates of the single cell to be predicted are input into the single cell gene expression response prediction model to perturbation trained by the training method described above in this invention, so as to obtain the gene expression data of the single cell under the preset perturbation conditions.
[0038] A third aspect of the present invention provides a system for predicting the response of single-cell gene expression to perturbations, comprising:
[0039] The data acquisition module is used to acquire gene expression data of the single cell to be predicted without perturbation, preset perturbation conditions, and cell covariates of the single cell to be predicted.
[0040] The response prediction module is used to input the gene expression data of the single cell to be predicted without perturbation, the preset perturbation conditions, and the cell covariates of the single cell to be predicted into the single cell gene expression response prediction model to perturbation trained by the training method described above, so as to obtain the gene expression data of the single cell to be predicted under the preset perturbation conditions.
[0041] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the training method or the prediction method described above.
[0042] A fifth aspect of the present invention provides a terminal comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the computer program is executed by the processor, implements the training method of the present invention as described above or implements the prediction method of the present invention as described above.
[0043] Beneficial Effects: This invention rethinks the perturbation prediction problem from a decoupling perspective. Based on contrastive learning and an encoder-decoder architecture, it decouples gene expression features to predict the response of single-cell gene expression to perturbations. The prediction model based on the encoder-decoder architecture, trained using the method provided by this invention, can capture the relationship between perturbation conditions and cellular responses more precisely and at a finer granular level, especially in complex biological data, predicting the response of single-cell gene expression to perturbations. It exhibits stronger accuracy, robustness, and generalization ability when processing different types of perturbation and cellular response data. The prediction model trained using the method provided by this invention can accurately predict the gene expression data of single cells under corresponding perturbations. Attached Figure Description
[0044] Figure 1 is a flowchart illustrating the training method of a single-cell gene expression response prediction model to perturbation in one embodiment of the present invention.
[0045] Figure 2 is a flowchart illustrating the training method of a single-cell gene expression response prediction model to perturbation in another embodiment of the present invention.
[0046] Figure 3 is a flowchart illustrating the training method of a single-cell gene expression response prediction model to perturbation in another embodiment of the present invention.
[0047] Figure 4 is a prediction model diagram of the response of single-cell gene expression to perturbation in an embodiment of the present invention.
[0048] Figure 5 is a flowchart illustrating the training method of a single-cell gene expression response prediction model to perturbation in another embodiment of the present invention.
[0049] Figure 6 is a flowchart of the method for predicting the response of single-cell gene expression to perturbation in an embodiment of the present invention.
[0050] Figure 7 is a detailed structural diagram of the encoder and decoder in the prediction model of single-cell gene expression response to perturbation in an embodiment of the present invention.
[0051] Figure 8 shows the gene expression results of CD8+ T cells under CD247 overexpression perturbation predicted by different prediction methods.
[0052] Figure 9 shows the gene expression results of CD8+ T cells under SLA2 overexpression perturbation predicted by different prediction methods.
[0053] Figure 10 shows the gene expression results of CD8+ T cells under AKAP12 overexpression perturbation predicted by different prediction methods.
[0054] Figure 11 is a structural block diagram of the single-cell gene expression response prediction system to perturbation in an embodiment of the present invention. Detailed Implementation
[0055] This invention provides a method for predicting the response of single-cell gene expression to perturbations, a model training method, a system, and a terminal. To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. The invention is further illustrated below with reference to specific embodiments.
[0057] When a single cell is perturbed, its gene expression may change; that is, gene expression responds differently to different perturbations, and different types of single cells may also respond differently to the same perturbation. Predicting these responses can provide technical support for personalized medicine and genomics research. However, due to the difficulty in comprehensively covering all possible perturbations under experimental conditions, and the complexity of biological variability and high experimental costs, researchers find it difficult to obtain complete expression profiles under all cell types and conditions. Based on this, this invention provides a training method for a single-cell gene expression response prediction model to perturbations. The single-cell gene expression response prediction model includes an encoder, a contrastive learning-based latent space, and a decoder, as shown in Figure 1. The training method for the single-cell gene expression response prediction model includes the following steps:
[0058] S1. Obtain gene expression data of single cells under known perturbation conditions and under known perturbation conditions, as well as cell covariates of single cells;
[0059] S2. Using the known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cell covariates of single cells, train the encoder, the latent space based on contrastive learning, and the decoder to obtain a trained single-cell gene expression response prediction model to perturbation.
[0060] This invention rethinks the perturbation prediction problem from a decoupling perspective. Based on contrastive learning and an encoder-decoder architecture, it decouples gene expression features to predict the response of single-cell gene expression to perturbations. The prediction model based on the encoder-decoder architecture, trained using the method provided by this invention, can capture the relationship between perturbation conditions and cellular responses more precisely and at a finer granular level, especially in complex biological data, predicting the response of single-cell gene expression to perturbations. It exhibits stronger accuracy, robustness, and generalization ability when processing different types of perturbation and cellular response data. The prediction model trained using the method provided by this invention can accurately predict the gene expression data of single cells under corresponding perturbations.
[0061] In step S1, the known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cell covariates of single cells are used as the training dataset for model training. The gene expression data of single cells under the known perturbation conditions are preprocessed using one-hot encoding. The samples used to collect this training dataset include several single-cell samples (i.e., the training dataset is taken from several single-cell samples, including single cells of the same type and single cells of different types). For each single cell in the several single-cell samples, the training dataset includes the perturbation conditions for that single cell, the gene expression data corresponding to the perturbation conditions (i.e., the actual gene expression under the perturbation conditions), and the cell covariates of that single cell. In other words, the training dataset used for model training includes data from several single-cell samples, and the data for each single-cell sample includes the known perturbation conditions, the gene expression data of the single cell under the known perturbation conditions, and the cell covariates of the single cell.
[0062] The known perturbation conditions include, but are not limited to, at least one of drug treatment perturbation conditions, gene perturbation conditions, and environmental perturbation conditions. Drug treatment perturbation conditions may involve stimulating single cells with compounds to induce perturbation. Gene perturbation conditions include, but are not limited to, gene mutations, such as gene addition, deletion (knockout), or substitution. Gene expression data of single cells under known perturbation conditions can be obtained through transcriptome sequencing.
[0063] Gene expression data of single cells under known perturbation conditions include gene expression data of the same and different types of single cells under the same known perturbation conditions, as well as gene expression data of the same and different types of single cells under different known perturbation conditions.
[0064] The single-cell cellular covariates include the cell type of the single cell.
[0065] In step S2, in some embodiments, as shown in Figure 2, the step of training the encoder, the contrastive learning-based latent space, and the decoder using the known perturbation conditions, the gene expression data of single cells under the known perturbation conditions, and the cellular covariates of single cells specifically includes:
[0066] S211. The known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cell covariates of single cells are input into the encoder, and then the first latent variable is obtained after passing through the latent space based on contrastive learning.
[0067] S212. After combining the first latent variable with the known perturbation conditions, input it into the decoder to obtain the predicted value of single-cell gene expression under the known perturbation conditions.
[0068] In this step, the known perturbation condition is the original known perturbation condition input into the encoder.
[0069] S213. Obtain a loss function based on the predicted single-cell gene expression value under the known perturbation conditions and the gene expression data of single cells under the known perturbation conditions, and train according to the loss function.
[0070] In this step, the model is trained according to the loss function until the loss function converges, resulting in a trained single-cell gene expression response prediction model to perturbations.
[0071] In this embodiment, a loss function is obtained based on the predicted single-cell gene expression value under known perturbation conditions and the gene expression data of single cells under known perturbation conditions (i.e., the true value of single-cell gene expression under known perturbation conditions). Training based on the loss function can ensure that the prediction model retains key covariates and heterogeneity information within the cell.
[0072] In other embodiments, as shown in Figure 3, the encoder includes a feedforward neural network, a self-attention mechanism, and a fully connected layer. The specific steps of training the encoder, the contrastive learning-based latent space, and the decoder using the known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cellular covariates of single cells include:
[0073] S221. The known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cell covariates of single cells are input into the encoder. After passing through the feedforward neural network, the first feature vector, the second feature vector, and the third feature vector are obtained respectively.
[0074] In this step, the known perturbation conditions, the gene expression data of single cells under the known perturbation conditions, and the cell covariates of single cells are input into the encoder. The perturbation conditions, the gene expression data of single cells under the known perturbation conditions, and the cell covariates of single cells are characterized by the feedforward neural network to obtain the first feature vector, the second feature vector, and the third feature vector, respectively.
[0075] S222. The first feature vector, the second feature vector, and the third feature vector are concatenated and input into a self-attention mechanism for fusion. The result is output through a fully connected layer to obtain a potential high-level feature vector.
[0076] In this step, the first feature vector, the second feature vector, and the third feature vector are concatenated and input into the self-attention mechanism for effective combination, extracting potential high-level feature vectors (learned by the model itself), and output through a fully connected layer.
[0077] S223. Input the potential high-level feature vector into the latent space based on contrastive learning to obtain the first latent variable.
[0078] S224. The first latent variable is combined with the known perturbation condition (i.e., the original known perturbation condition input into the encoder) and then input into the decoder to obtain the single-cell gene expression prediction value under the known perturbation condition; a loss function is obtained based on the single-cell gene expression prediction value under the known perturbation condition and the gene expression data of the single cell under the known perturbation condition (i.e., the true value of single-cell gene expression under the known perturbation condition), and training is performed based on the loss function.
[0079] In this embodiment, the encoder represents the gene expression information of a single cell under known perturbation conditions by combining a feedforward neural network and a self-attention mechanism. After three types of data are input into the encoder, they are represented by the feedforward neural network to obtain computer-recognizable vector representations. These vector representations are then concatenated and input into the self-attention mechanism to effectively combine the three types of data and extract potential high-level features. This design can address common problems in bioinformatics such as noise and missing data, handle high-dimensional and complex data, ensure the extraction of key biological features, and improve the robustness of the model. This embodiment utilizes the attention mechanism and contrastive learning to decouple gene expression features, fully leveraging cell-related feature data to achieve an efficient information encoding and decoding process, accurately predicting the single-cell gene expression response to perturbations with high prediction accuracy. For example, it can predict perturbation responses observed in other cells but not yet seen in specific target cells.
[0080] In some embodiments, as shown in FIG4, step S223 specifically includes:
[0081] The latent high-level feature vectors are input into the latent space based on contrastive learning to form latent variables. Based on the equality of cellular covariates and the distance between latent variables in the latent space, positive and negative loss functions are introduced.
[0082] The positive and negative loss functions are summed to obtain the contrast loss function, and the first latent variable is output.
[0083] In this embodiment, to enhance the encoder's ability to effectively decouple data while retaining as much cellular covariate information as possible, contrastive learning is employed. A similarity matrix is constructed based on the equality of cellular covariates, and positive and negative loss functions are defined by combining this with the distance between latent variables in the latent space, ultimately forming a contrastive loss function. The encoder parameters are then optimized based on this contrastive loss function (ensuring convergence of the contrastive loss function during training). This method ensures that latent variables corresponding to the same cellular covariate are more closely spaced (or closer) in the latent space, while latent variables corresponding to different cellular covariates are more distant. (That is, when the cellular covariate for a single cell is cell type, contrastive learning allows the prediction model to bring gene expression vectors of the same cell type closer together and expression vectors of different cell types further apart, enabling the prediction model to learn the commonalities in gene expression for each cell type.) This method enhances the encoder's ability to separate data features and retains more meaningful cellular covariate information, thereby improving the model's performance and discriminative ability.
[0084] Specifically, a similarity matrix is constructed based on the equality of cell covariates (whether they belong to the same category), and the distances between latent variables in the latent space are obtained. The positive loss function (Lp) is then derived by combining the similarity matrix and the distances between latent variables. ositive ) and negative loss function (L negative ), thus obtaining the contrastive loss function (L contrast ),in: L contrast =L positive +L negative ;
[0085] N is the total number of samples (i.e., the total number of single-cell samples), S ij Let S represent the similarity matrix, which quantifies the degree to which samples i and j are classified into the same category of cellular covariates. By comparing the cellular covariates of sample i and sample j, a difference vector is obtained between the cellular covariates of sample i and sample j. Then, the weight matrix is multiplied by this difference vector to obtain S. ij The similarity weights of different cellular covariates may vary, thus reflecting their respective impact on the analysis. ij This represents the distance between the latent variables corresponding to sample i and sample j in the latent space, specifically the Euclidean distance. ∈ is a constant, a very small value used to prevent division by zero, thus ensuring the stability and reliability of the calculation. margin is the edge value (margin = 1), max(0, margin - D) ij ) indicates the range between 0 and margin-D ij Take the maximum value between the two.
[0086] In some embodiments, as shown in Figures 4 and 5, step S224 specifically includes:
[0087] S2241. After combining the first latent variable with the known perturbation conditions, input it into the decoder to obtain the predicted value of single-cell gene expression under the known perturbation conditions. Obtain the reconstruction loss function based on the predicted value of single-cell gene expression under the known perturbation conditions and the gene expression data of single cells under the known perturbation conditions (i.e., the true value of single-cell gene expression under the known perturbation conditions).
[0088] Simultaneously, counterfactual perturbation conditions and single-cell gene expression data under counterfactual perturbation conditions are provided. The first latent variable is combined with the counterfactual perturbation conditions and input into the decoder to obtain the predicted value of single-cell gene expression under counterfactual perturbation conditions. The counterfactual loss function is obtained based on the predicted value of single-cell gene expression under counterfactual perturbation conditions and the gene expression data of single cells under counterfactual perturbation conditions (i.e., the true value of single-cell gene expression under counterfactual perturbation conditions).
[0089] In some implementations, the decoder consists of fully connected layers. Since the attention mechanism and contrastive learning method used by the encoder already result in sufficiently high model complexity, the decoder employs fully connected layers. The purpose of using fully connected layers is to reduce the risk of overfitting without increasing model complexity, while still enabling decoding based on the feature information output by the encoder.
[0090] S2242. Train according to the reconstruction loss function, the counterfactual loss function and the contrast loss function.
[0091] In this embodiment, the encoder integrates single-cell gene expression data under perturbation conditions, single-cell cellular covariates, and preset perturbation conditions to obtain fused information. This fused information is then mapped to a latent space, outputting a first latent variable. Within this latent space, as shown in Figure 4, the data is distributed along two distinct pathways: the first pathway combines with the original perturbation condition data (i.e., the first latent variable combines with known perturbation conditions); while the second pathway integrates with perturbation data from counterfactual cells sharing the same covariates (i.e., the first latent variable combines with counterfactual perturbation conditions). During the decoder stage, these two pathways (or data streams) are processed to reconstruct the actual gene expression of the cells and predict the gene expression of counterfactual cells.
[0092] In this embodiment, the reconstruction loss function ensures that the prediction model can retain key covariates and heterogeneity information within cells, while the counterfactual loss function requires the encoder to retain relevant cell data (the counterfactual loss function prevents the encoder from indiscriminately retaining all information about gene expression, but only retains relevant cell data), thereby enabling accurate prediction of gene expression under assumed perturbation conditions when fused with other perturbation condition data.
[0093] Step S2242, the step of training based on the reconstruction loss function, the counterfactual loss function, and the contrastive loss function specifically includes:
[0094] According to the reconstruction loss function (L) recon ) and the counterfactual loss function (L cf ), to obtain the target loss function (L target );
[0095] The model is trained using the backpropagation algorithm based on the target loss function and the contrast loss function. The parameters of the prediction model are optimized (or adjusted) until the target loss function and the contrast loss function converge, thus obtaining the prediction model. (That is, during the training process, the target loss function and the contrast loss function continue to decrease until they no longer decrease within 5 training rounds. At this point, training stops, and the training parameters at the lowest point of the target loss function and the contrast loss function are saved as the final optimized parameters.)
[0096] in, L target =λ1L reco n+λ2L cf
[0097] N is the number of samples (i.e., the total number of single-cell samples), y i and y i,cf These are the actual gene expression levels of the i-th sample (i.e., the i-th single-cell sample) under known perturbation and counterfactual perturbation conditions, respectively (corresponding to the gene expression data of single cells under known perturbation and counterfactual perturbation conditions mentioned above, i.e., the true gene expression values of single cells under known perturbation and counterfactual perturbation conditions); E(y i c i , t i ) is the process of processing gene expression y i Cellular covariate c i Given the disturbance condition t i The first latent variable obtained later; D(E(y) i c i , t i ), t i ) is under the known disturbance condition ti Decoder function for reconstructing gene expression; D(E(y) i c i , t i ), t i,cf ) is under the counterfactual perturbation condition t i,cf The decoder function for reconstructing gene expression; λ1 and λ2 are weighting coefficients that balance reconstruction importance and counterfactual loss, respectively.
[0098] The training method of the single-cell gene expression response prediction model to perturbation is explained in detail below with reference to Figure 4.
[0099] Known perturbation conditions, gene expression data of single cells under known perturbation conditions, and cell covariates of single cells are input into the encoder. After passing through the feedforward neural network, the first feature vector, the second feature vector, and the third feature vector are obtained respectively. The first feature vector, the second feature vector, and the third feature vector are concatenated and input into the self-attention mechanism for fusion. The result is output through a fully connected layer to obtain the potential high-level feature vector.
[0100] The potential high-level feature vectors are input into the latent space based on contrastive learning. In the latent space, a similarity matrix is constructed based on the equality of cell covariates (whether they belong to the same category), and the distance between latent variables in the latent space is calculated. By combining the similarity matrix and the distance between latent variables, positive loss function and negative loss function are obtained, and then contrastive loss function is obtained, and the first latent variable is output.
[0101] The first latent variable is combined with the known perturbation conditions and input into the decoder to obtain the predicted single-cell gene expression value under the known perturbation conditions. A reconstruction loss function is obtained based on the predicted single-cell gene expression value under the known perturbation conditions and the gene expression data of the single cells under the known perturbation conditions. Simultaneously, counterfactual perturbation conditions and gene expression data of the single cells under these conditions are provided. The first latent variable is combined with the counterfactual perturbation conditions and input into the decoder to obtain the predicted single-cell gene expression value under the counterfactual perturbation conditions. A counterfactual loss function is obtained based on the predicted single-cell gene expression value under the counterfactual perturbation conditions and the gene expression data of the single cells under these conditions.
[0102] The target loss function is obtained based on the constructed loss function and the counterfactual loss function;
[0103] The prediction model is trained using the contrastive loss function and the target loss function. The parameters in the prediction model are then adjusted and optimized until the two functions converge to obtain the prediction model. (The contrastive loss function is used in the training of the encoder, and the target loss function is used in the training of the entire prediction model. That is, the optimization of the decoder is based on the target loss function, and the optimization of the encoder is based on the contrastive loss function and the target loss function.)
[0104] This invention also provides a method for predicting the response of single-cell gene expression to perturbation, which, as shown in Figure 6, includes the following steps:
[0105] S11. Obtain gene expression data of the single cell to be predicted without perturbation, preset perturbation conditions, and cell covariates of the single cell to be predicted.
[0106] S12. Input the gene expression data of the single cell to be predicted without disturbance, the preset disturbance conditions, and the cell covariates of the single cell to be predicted into the single cell gene expression response prediction model to disturbance to obtain the gene expression data of the single cell under the preset disturbance condition.
[0107] In some implementations, step S12 specifically includes: inputting the gene expression data of the single cell to be predicted without perturbation, the preset perturbation conditions, and the cell covariates of the single cell to be predicted into a single-cell gene expression response prediction model to perturbation, passing it sequentially through an encoder and a latent space based on contrastive learning to obtain a second latent variable; combining the second latent variable with the preset perturbation conditions and inputting it into the encoder to obtain the gene expression data of the single cell to be predicted under the preset perturbation conditions.
[0108] The training method for the single-cell gene expression response prediction model to perturbation is described above. For example, the single-cell gene expression response prediction model to perturbation includes an encoder, a latent space based on contrastive learning, and a decoder. The decoder includes a feedforward neural network, a self-attention mechanism, and a fully connected layer. The decoder is composed of a fully connected layer, and the prediction model is generated using the following method:
[0109] The known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cell covariates of single cells (i.e. cell type of single cells) are input into the encoder. After passing through the feedforward neural network, the first feature vector, the second feature vector, and the third feature vector are obtained respectively.
[0110] The first feature vector, the second feature vector, and the third feature vector are concatenated and input into a self-attention mechanism for fusion. The result is output through a fully connected layer to obtain a potential high-level feature vector.
[0111] The potential high-level feature vectors are input into the latent space based on contrastive learning. In the latent space, a similarity matrix is constructed based on the equality of cell covariates (whether they belong to the same category), and the distance between latent variables in the latent space is calculated. Positive loss function and negative loss function are obtained by combining the similarity matrix and the distance between latent variables. Contrast loss function is obtained by summing loss function and negative loss function, and the first latent variable is output.
[0112] The first latent variable is combined with the known perturbation condition and input into the decoder to obtain the predicted value of single-cell gene expression under the known perturbation condition. The reconstruction loss function is obtained based on the predicted value of single-cell gene expression under the known perturbation condition and the true value of single-cell gene expression under the known perturbation condition.
[0113] Simultaneously, the first latent variable is combined with the counterfactual perturbation and input into the encoder to obtain the predicted value of single-cell gene expression under the counterfactual perturbation condition. The counterfactual loss function is obtained based on the predicted value of single-cell gene expression under the counterfactual perturbation condition and the true value of single-cell gene expression under the counterfactual perturbation condition. The target loss function is obtained based on the reconstruction loss function and the counterfactual loss function.
[0114] The prediction model is trained using the contrastive loss function and the target loss function, and the parameters are adjusted until the two loss functions converge to obtain the trained prediction model.
[0115] Specifically, as shown in Figure 7, Y represents gene expression data (in prediction, it corresponds to the gene expression data of the single cell to be predicted without perturbation), C represents cell covariates, T represents perturbation conditions, and Z represents latent variables output by the latent space (in prediction, it corresponds to the second latent variable).
[0116] When performing response prediction, that is, when predicting gene expression of a single cell after perturbation, the input is the gene expression data of the single cell to be predicted without perturbation, the cell covariates of the single cell to be predicted, and the preset perturbation conditions.
[0117] Then, the three types of input data are represented by a feedforward neural network, that is, converted into vector form that the computer can understand. These are then fused using a self-attention mechanism and output through a fully connected layer. After passing through a latent space based on contrastive learning (not shown in Figure 7), a second latent variable is output. This second latent variable is combined with the current perturbation (corresponding to the preset perturbation condition during prediction) and the counterfactual perturbation, respectively, and then passed through a decoder with the same structure to make the final prediction, obtaining the single-cell gene expression prediction value and the single-cell gene expression prediction value under the counterfactual perturbation condition.
[0118] To verify the accuracy of the single-cell gene expression response prediction method (prediction model based on contrastive learning and attention mechanism) provided in this embodiment of the invention, the effects of three perturbations in CD8+ T cells—CD247 overexpression, SLA2 overexpression, and AKAP12 overexpression—on CD8+ T cell gene expression were analyzed. For each perturbation, differential gene expression (DE) under different prediction methods was plotted. The control group (i.e., the unperturbed single-cell group, corresponding to "control" in the figure), the prediction results obtained using GraphVCI (Graph Variational Contrastive Identifiers) (corresponding to "GraphVCI" in the figure), and the prediction results obtained using the prediction method provided in this embodiment of the invention (corresponding to "scCADE" in the figure) were compared with the true gene expression levels of perturbed CD8+ T cells (corresponding to "True" in the figure). The results are shown in Figures 8, 9, and 10. The overexpression results of CD247 and SLA2 showed that there were significant differences between the true perturbed expression and control expression of genes such as LSP1, KIF3A, TRAV26-2, RPS6KA5, HMGB2, NFIA, GLCCI1, and CD300A. In addition, in the AKAP12 overexpression results, the true disordered expression of the P2RY10 gene also showed a significant difference compared with the control group.
[0119] Furthermore, by comparing the prediction results obtained using GraphVCI (corresponding to GraphVCI in the figure) and the prediction results obtained using the prediction method provided in this embodiment of the invention (corresponding to scCADE in the figure) with the actual gene expression of CD8+ T cells under perturbation (corresponding to True in the figure), it was found that, compared with the prediction generated by GraphVCI, the prediction results obtained using the prediction method provided in this embodiment of the invention are more consistent with the actual gene expression of CD8+ T cells under perturbation, and have a larger deviation from the gene expression of the control group, highlighting the superior accuracy of the prediction method provided by this invention in simulating gene perturbation. Compared with GraphVCI, the prediction method provided by this invention shows a significant advantage in accurately predicting the response of single cells to gene perturbation. The prediction model provided by this invention is particularly evident in its ability to closely match the actual perturbation expression, highlighting its effectiveness.
[0120] This invention also provides a single-cell gene expression response to perturbation prediction system, wherein, as shown in Figure 11, the single-cell gene expression response to perturbation prediction system includes:
[0121] Data acquisition module 1 is used to acquire gene expression data of the single cell to be predicted without perturbation, preset perturbation conditions, and cell covariates of the single cell to be predicted;
[0122] The response prediction module 2 is used to input the gene expression data of the single cell to be predicted without disturbance, the preset disturbance conditions, and the cell covariates of the single cell to be predicted into the single cell gene expression response prediction model to disturbance trained by the training method described above in the embodiments of the present invention, so as to obtain the gene expression data of the single cell to be predicted under the preset disturbance conditions.
[0123] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the training method or the prediction method described above in this invention.
[0124] The computer-readable medium described in this embodiment may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0125] The present invention also provides a terminal, which includes a memory and a processor. The memory stores a computer program that can run on the processor. When the computer program is executed by the processor, it implements the training method or the prediction method described above in the embodiments of the present invention.
[0126] In this embodiment, the memory can be volatile memory, such as random access memory; the memory can also be non-volatile memory, such as read-only memory, flash memory, hard disk, or optical disk. The processor can be a central processing unit, controller, microcontroller, microprocessor, or other data processing chip.
[0127] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A training method for a single-cell gene expression response prediction model to perturbation, characterized in that, The single-cell gene expression response prediction model to perturbation includes an encoder, a latent space based on contrastive learning, and a decoder. The training method of the single-cell gene expression response prediction model to perturbation includes the following steps: Obtain gene expression data of single cells under known perturbation conditions and under known perturbation conditions, as well as cellular covariates of single cells; Using the known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cell covariates of single cells, the encoder, the latent space based on contrastive learning, and the decoder are trained to obtain a trained single-cell gene expression response prediction model to perturbation.
2. The training method according to claim 1, characterized in that, The specific steps for training the encoder, the contrastive learning-based latent space, and the decoder using the known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cellular covariates of single cells include: The known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cell covariates of single cells are input into the encoder, and then, after passing through the latent space based on contrastive learning, the first latent variable is obtained. After combining the first latent variable with the known perturbation conditions, the data is input into the decoder to obtain the predicted single-cell gene expression value under the known perturbation conditions. A loss function is obtained based on the predicted single-cell gene expression values under the known perturbation conditions and the single-cell gene expression data under the known perturbation conditions, and training is performed based on the loss function.
3. The training method according to claim 1, characterized in that, The encoder includes a feedforward neural network, a self-attention mechanism, and a fully connected layer. The specific steps for training the encoder, the contrastive learning-based latent space, and the decoder using the known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cellular covariates of single cells include: The known perturbation conditions, gene expression data of single cells under the known perturbation conditions, and cell covariates of single cells are input into the encoder. After passing through the feedforward neural network, the first feature vector, the second feature vector, and the third feature vector are obtained respectively. The first feature vector, the second feature vector, and the third feature vector are concatenated and input into a self-attention mechanism for fusion. The result is output through a fully connected layer to obtain a potential high-level feature vector. The potential high-level feature vectors are input into the latent space based on contrastive learning to obtain the first latent variable; After combining the first latent variable with the known perturbation conditions, the data is input into the decoder to obtain the predicted single-cell gene expression value under the known perturbation conditions. A loss function is obtained based on the predicted single-cell gene expression values under the known perturbation conditions and the single-cell gene expression data under the known perturbation conditions, and training is performed based on the loss function.
4. The training method according to claim 3, characterized in that, The specific steps of inputting the latent high-level feature vectors into the latent space based on contrastive learning to obtain the first latent variable include: The latent high-level feature vectors are input into the latent space based on contrastive learning to form latent variables. Based on the equality of cellular covariates and the distance between latent variables in the latent space, positive and negative loss functions are introduced. The positive and negative loss functions are summed to obtain the contrast loss function, and the first latent variable is output.
5. The training method according to claim 4, characterized in that, The first latent variable is combined with the known perturbation condition and input into the decoder to obtain the predicted single-cell gene expression value under the known perturbation condition. A loss function is obtained based on the predicted single-cell gene expression value under the known perturbation condition and the gene expression data of the single cell under the known perturbation condition. The training step based on the loss function specifically includes: The first latent variable is combined with the known perturbation conditions and input into the decoder to obtain the predicted value of single-cell gene expression under the known perturbation conditions. The reconstruction loss function is obtained based on the predicted value of single-cell gene expression under the known perturbation conditions and the gene expression data of single cells under the known perturbation conditions. Simultaneously, it provides counterfactual perturbation conditions and gene expression data of single cells under counterfactual perturbation conditions. After combining the first latent variable with the counterfactual perturbation conditions, it inputs it into the decoder to obtain the predicted value of single-cell gene expression under counterfactual perturbation conditions. Based on the predicted value of single-cell gene expression under counterfactual perturbation conditions and the gene expression data of single cells under counterfactual perturbation conditions, it obtains the counterfactual loss function. Training is performed based on the reconstruction loss function, the counterfactual loss function, and the contrastive loss function.
6. The training method according to claim 1, characterized in that, The decoder consists of a fully connected layer; The known perturbation conditions include at least one of drug treatment perturbation conditions, gene perturbation conditions, and environmental perturbation conditions.
7. A method for predicting the response of single-cell gene expression to perturbation, characterized in that, Includes the following steps: Acquire gene expression data of the single cell to be predicted without perturbation, preset perturbation conditions, and cellular covariates of the single cell to be predicted; The gene expression data of the single cell to be predicted without perturbation, the preset perturbation conditions, and the cell covariates of the single cell to be predicted are input into the single cell gene expression response prediction model to perturbation trained by the training method described in any one of claims 1-6, to obtain the gene expression data of the single cell under the preset perturbation conditions.
8. A system for predicting the response of single-cell gene expression to perturbations, characterized in that, include: The data acquisition module is used to acquire gene expression data of the single cell to be predicted without perturbation, preset perturbation conditions, and cell covariates of the single cell to be predicted. The response prediction module is used to input the gene expression data of the single cell to be predicted without disturbance, the preset disturbance conditions, and the cell covariates of the single cell to be predicted into the single cell gene expression response prediction model to disturbance trained by the training method described in any one of claims 1-6, so as to obtain the gene expression data of the single cell to be predicted under the preset disturbance conditions.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the training method according to any one of claims 1-6 or the prediction method according to claim 7.
10. A terminal, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program that can run on the processor, and when the computer program is executed by the processor, it implements the training method of any one of claims 1-6 or the prediction method of claim 7.