Single-cell drug sensitivity prediction algorithm based on semi-supervised transfer learning
Through a semi-supervised transfer learning algorithm, using cell lines and a small amount of labeled single-cell data, combined with shared feature encoder and adversarial learning, the problem of scarcity of data in single-cell drug sensitivity prediction is solved, significantly improving prediction accuracy and reducing experimental costs.
Patent Information
- Application Number
- CN202510053111.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-16
AI Technical Summary
When predicting the sensitivity of single-cell drug, the prior art is difficult to effectively overcome the problem of insufficient single-cell domain training data due to the scarcity and cost of data.
A single-cell drug sensitivity prediction algorithm for semi-supervised transfer learning is proposed. By utilizing cell line drug sensitivity data and a small amount of labeled single-cell data, combined with shared feature encoder, 12 normalization and adversarial learning, semi-supervised domain adaptation is carried out to improve prediction performance.
It significantly improves the accuracy of single-cell drug sensitivity prediction, and can efficiently predict single-cell sensitivity to drugs on multiple single-cell data sets, reducing the cost of experimental verification.
Smart Images

Figure CN120015117A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, which is an application of deep technology in drug research and development, and in particular to a single-cell drug sensitivity prediction algorithm based on semi-supervised transfer learning. Background Art
[0002] High-throughput drug screening technologies have generated large-scale drug response data covering thousands of tumor cell lines, facilitating the development of computational prediction of cell line drug response. Despite its great potential, cell line gene expression profiles cannot discern changes in transcriptional programs and regulatory mechanisms within individual cells, thereby masking intratumor heterogeneity. As a result, previous models perform poorly on out-of-distribution samples, including patients and single cells. Advances in single-cell sequencing technologies have enabled us to explore the complexity and variability between individual cells, providing us with an opportunity to gain a deeper understanding of intratumor heterogeneity. However, due to cost and technical limitations, currently available single-cell-level drug response data only cover a few cancer types and drugs, which poses a challenge to the development of computational methods for predicting single-cell drug responses. At present, some work has modeled drug-induced cell line sequencing data, using deep transfer learning to transfer cell line drug response knowledge to single cells, which has become an effective means to overcome the lack of training data in the single-cell domain. However, they either rely on manually designed scores to rank drugs or rely on prior knowledge to identify cell populations of interest.
[0003] In recent years, semi-supervised domain adaptation methods have achieved better performance by leveraging a small number of labeled samples in the target domain. Through semi-supervised domain adaptation, we transfer drug responses from the cell line level to the single cell level. Summary of the invention
[0004] The purpose of the present invention is to solve the shortcomings of the prior art and reduce the cost of experimental verification. The present invention proposes a single-cell drug sensitivity prediction algorithm based on semi-supervised transfer learning, which provides a new idea for predicting single-cell drug sensitivity using cell line drug sensitivity.
[0005] The specific technical solution of the present invention is: a single-cell drug sensitivity prediction algorithm based on semi-supervised transfer learning, comprising the following steps:
[0006] 1) Data preprocessing:
[0007] 1-1 Prepare data sets: The data sets required for the model are cell line gene expression data sets, single cell gene expression data sets, and drug sensitivity data sets; the single cell gene expression data sets include a small number of labeled single cell gene expression data sets and a large number of unlabeled single cell gene expression data sets; the drug sensitivity data sets include cell line drug sensitivity data sets and a small number of drug sensitivity data sets corresponding to labeled single cells.
[0008] 1-2 Standardization: Gene expression data with the same genes in the cell line gene expression dataset and the single cell gene expression dataset were extracted respectively, and the cell line gene expression dataset and the single cell gene expression dataset with the same genes were standardized in the same way. The standardized cell line gene expression dataset includes: data with row names as gene names and column names as cell line names; the standardized single cell gene expression dataset includes: data with row names as gene names and column names as single cell names; the drug sensitivity dataset includes: data with row names as cell names and column names as 0 or 1 of labels.
[0009] In step 1-2, the z-score method was used as the data normalization method.
[0010] 1-3 Split the dataset: 80% of the standardized cell line gene expression dataset and the standardized unlabeled single-cell gene expression dataset are used as the model training dataset, and 20% are used as the validation dataset; the standardized labeled single-cell gene expression dataset is divided in half and added to the model training dataset and model validation dataset respectively.
[0011] 2) Drug sensitivity prediction with labeled data:
[0012] 2-1 Input data contains N s Source domain data There are labeled samples and N t Labeled target domain data where x s represents the gene expression data of cell lines, y s represents the sensitivity label of the drug to the cell line, x t represents single-cell gene expression data, y t It represents the sensitivity label of the drug to a single cell. We process the dataset using methods 1-2 and 1-3 and then input it into the model.
[0013] 2-2 Use the shared feature encoder to map cell line and labeled single-cell gene feature data into an l-dimensional feature vector
[0014] 2-3 In order to obtain more reliable output of the model, make the directions of features from the same class closer to each other, and separate different classes, we use l2 normalization to obtain the feature vector obtained by 2-2 to obtain an l-dimensional feature vector
[0015] 2-4 A weight vector W C A linear layer composed of a layer is used as a classifier to evaluate the correlation between cell line pharmacogenomics information and drug response at the cell line level. According to 2-3 We use it as the input of the classifier, followed by a softmax layer with a temperature parameter T, to predict the drug sensitivity label at the cell line level.
[0016] In 2-4, the model uses an unbiased linear layer with 128 input dimensions and 2 output dimensions as a classifier.
[0017] 2-5 In order to make the model less sensitive to small changes in the input, we add a small amount of perturbation near the input data, and the perturbation is in the direction of the loss gradient increase.
[0018] 3) Semi-supervised domain adaptation:
[0019] 3-1 Input data is N u Unlabeled target domain data Unlabeled samples, where x u represents single-cell gene expression data, and N u <<N t , we process the data set using methods 1-2 and 1-3, and then input it into the model;
[0020] 3-2 Use the same method as 2-2 and 2-3 in step 2) to obtain the single cell l-dimensional feature vector
[0021] 3-3 For semi-supervised domain adaptation, we update the classifier by maximizing the entropy of unlabeled target domain samples and minimize the entropy of unlabeled target domain samples to update the feature extractor, thus forming an adversarial relationship.
[0022] 3-4 Use adversarial learning to jointly train and update all the models in the previous steps;
[0023] 4) Single cell drug sensitivity prediction:
[0024] 4-1 Test dataset preparation: the standardized single-cell gene expression dataset in step 1;
[0025] 4-2 will obtain the fully trained feature extractor and classifier assembly, and then input the test data set into the model to predict the sensitivity label of single cells to drugs.
[0026] The entire framework input can be divided into two parts: labeled data: including all cell gene expression datasets and a small amount of labeled single-cell gene expression data, which we use to initially train the model so that the model can correctly classify the labeled dataset, and use adversarial training to clarify the decision boundary of the model; unlabeled data: including a large number of unlabeled single-cell gene expression profiles, we update the classifier by maximizing the entropy of unlabeled target domain samples, and minimize the entropy of unlabeled target domain samples to update the feature extractor, and conduct adversarial training. Not only does it narrow the distribution gap between cell lines and single cells, but it also learns discriminative features specific to classification tasks, thereby significantly improving the performance of single-cell drug sensitivity prediction.
[0027] The present invention has the following beneficial effects:
[0028] 1. A semi-supervised transfer learning algorithm for predicting single-cell drug sensitivity. A small amount of labeled single-cell datasets are used in training. In reality, there may be a small amount of labeled data. Our model can significantly improve the prediction of single-cell drug sensitivity by using only one or two labeled single-cell data.
[0029] 2. A single-cell drug sensitivity prediction algorithm based on semi-supervised transfer learning can predict whether a single cell is sensitive to a certain drug with extremely high accuracy on multiple single-cell data sets. It can be used in related tumor heterogeneity research and provide a certain reference for experimenters to save manpower and material resources and reduce experimental costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a model framework diagram of the present invention;
[0031] Figure 2 The ROC curve and AUC index of the present invention in predicting the sensitivity of melanoma single cell 451Lu to the drug PLX4720; DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0033] The GDSC database cell line gene expression data and GDSC database cell line drug sensitivity data were downloaded from https: / / www.cancerrxgene.org / downloads / bulk_download; the single-cell gene expression data were downloaded from the GEO database, and the GEO database website is https: / / www.ncbi.nlm.nih.gov / geo / ;
[0034] Figure 1 , a single-cell drug sensitivity prediction algorithm based on semi-supervised transfer learning, including the following steps:
[0035] The following steps are involved:
[0036] 1) Data preprocessing:
[0037] 1-1 Prepare data sets: The data sets required for the model are cell line gene expression data sets, single cell gene expression data sets, and drug sensitivity data sets; the single cell gene expression data sets include a small number of labeled single cell gene expression data sets and a large number of unlabeled single cell gene expression data sets; the drug sensitivity data sets include cell line drug sensitivity data sets and a small number of drug sensitivity data sets corresponding to labeled single cells.
[0038] In step 1-1, the data sets are cell line gene expression data from the GDSC database, cell line drug sensitivity data from the GDSC database, and single cell gene expression data from the GEO database;
[0039] 1-2 Standardization: Gene expression data with the same genes in the cell line gene expression dataset and the single cell gene expression dataset were extracted respectively, and the cell line gene expression dataset and the single cell gene expression dataset with the same genes were standardized in the same way. The standardized cell line gene expression dataset includes: data with row names as gene names and column names as cell line names; the standardized single cell gene expression dataset includes: data with row names as gene names and column names as single cell names; the drug sensitivity dataset includes: data with row names as cell names and column names as 0 or 1 of labels.
[0040] In step 1-2, the z-score method was used as the data normalization method.
[0041] 1-3 Split the dataset: 80% of the standardized cell line gene expression dataset and the standardized unlabeled single-cell gene expression dataset are used as the model training dataset, and 20% are used as the validation dataset; the standardized labeled single-cell gene expression dataset is divided in half and added to the model training dataset and model validation dataset respectively.
[0042] 2) Drug sensitivity prediction with labeled data:
[0043] 2-1 Input data contains N s Source domain data There are labeled samples and N t Labeled target domain data where x s represents the gene expression data of cell lines, y s represents the sensitivity label of the drug to the cell line, x t represents single-cell gene expression data, y t It represents the sensitivity label of the drug to a single cell. We process the dataset using methods 1-2 and 1-3 and then input it into the model.
[0044] 2-2 Use the denoising autoencoder as a shared feature encoder to map cell line and labeled single-cell gene feature data to an l-dimensional feature vector The loss function is:
[0045]
[0046] In formula (1), N s +N t is the number of cells, x i is the cell line gene expression data, x′ i is the gene expression data of cell lines after adding noise, q D (x′ i |x i ) is the noise distribution, E is the encoding network, and D is the decoding network. Here we use the binomial distribution.
[0047] 2-3 In order to obtain more reliable output of the model, make the directions of features from the same class closer to each other, and separate different classes, we use l2 normalization to obtain the feature vector obtained by 2-2 to obtain an l-dimensional feature vector
[0048] 2-4 A weight vector W C A linear layer composed of a layer is used as a classifier to evaluate the correlation between cell line pharmacogenomics information and drug response at the cell line level. According to 2-4 We use it as the input of the classifier, followed by a softmax layer with a temperature parameter T, to predict the drug sensitivity label at the cell line level, with the loss function being:
[0049]
[0050] In formula (2), σ is the softmax activation function and y is the true label.
[0051] In 2-4, the model uses an unbiased linear layer with 128 input dimensions and 2 output dimensions as a classifier.
[0052] 2-5 In order to make the model less sensitive to small changes in the input, we add a small amount of adversarial perturbation near the input data. The perturbation is oriented in the direction of the loss gradient increase. The adversarial loss function is:
[0053]
[0054] In formula (3), D[·,·] represents the Kullback-Leibler divergence, which is used to measure the difference between two distributions, p(y|x) represents the probability distribution of predicting label y given input x, q(y) is the true distribution of labels, which is usually approximated by the one-hot vector of y, g is the gradient that can be efficiently calculated using backpropagation, and ε is the magnitude of the adversarial perturbation.
[0055] 3) Semi-supervised domain adaptation:
[0056] 3-1 Input data is N u Unlabeled target domain data Unlabeled samples, where x represents single-cell gene expression data, and N u <<N t , we process the data set using methods 1-2 and 1-3, and then input it into the model;
[0057] 3-2 Use the same method as 2-2 and 2-3 in step 2) to obtain the single cell l-dimensional feature vector
[0058] 3-3 For semi-supervised domain adaptation, we update the classifier by maximizing the entropy of unlabeled target domain samples and minimize the entropy of unlabeled target domain samples to update the feature extractor, thus forming an adversarial relationship. The entropy is calculated as:
[0059]
[0060] In formula (4), p(y|x) represents the probability distribution of predicting label y given input x.
[0061] Here, we use a gradient reversal layer to implement adversarial learning:
[0062]
[0063] In formula (5), Θ and Φ are the parameters of the classifier and feature extractor respectively.
[0064] 3-4 Use adversarial learning to jointly train and update all models in the previous steps. The total loss training target is:
[0065]
[0066] In formula (6), λ is a hyperparameter used to balance the minimum maximum entropy loss and the classification task loss on labeled samples.
[0067] 4) Unlabeled data drug sensitivity prediction:
[0068] 4-1 Test dataset preparation: the standardized single-cell gene expression dataset in step 1;
[0069] 4-2 will obtain the fully trained feature extractor and classifier assembly, and then input the test data set into the model to predict the sensitivity label of single cells to drugs.
[0070] In the network model:
[0071] Labeled data: contains all cell gene expression datasets and a small amount of labeled single-cell gene expression data. We use this to perform initial training on the model so that the model can correctly classify the labeled dataset, and use adversarial training to further clarify the decision boundary of the model.
[0072] Unlabeled data: Contains a large number of unlabeled single-cell gene expression profiles. We perform adversarial training by maximizing the entropy of unlabeled target domain samples to update the classifier and minimizing the entropy of unlabeled target domain samples to update the feature extractor. This not only narrows the distribution gap between cell lines and single cells, but also learns discriminative features specific to classification tasks, thereby significantly improving the performance of single-cell drug sensitivity prediction.
[0073] In this embodiment, the cell line gene expression data in the GDSC database includes 834 cell line gene expression data and 2856 gene features; the single cell gene expression data includes 155 single cell gene expression data and 2856 gene features; the cell line drug sensitivity data downloaded from the GDSC database includes the drug PLX4720 response data for 834 cell lines; the drug sensitivity data of the drug PLX4720 for 155 single cells are downloaded from the GEO database. For the labeled single cell data set, we randomly selected 1, 2, 5, and 10 samples from each type of these 155 single cells as labeled data, conducted 4 experiments, and used the remaining data as unlabeled data sets.
[0074] A denoising autoencoder with binomial noise is used as a shared feature encoder, and an unbiased linear layer with 128 input dimensions and 2 output dimensions is used as a classifier. The shared feature encoder in the model encodes the data into 128-dimensional features, and then outputs the prediction results through the classifier. Adagrad is used as the optimizer in the training process, and the encoder is updated with a learning rate of 0.001. Then, the graph attention neural network, denoising autoencoder, cross attention mechanism and classifier trained by adversarial learning are used as test models. The experiment uses 80% of the data as the training set and 20% of the data for verification. In the test phase of the model, the model inputs 155 melanoma single-cell gene expression data in the target domain, and a total of 100 epochs are trained.
[0075] The results of predicting whether melanoma single cell 451Lu is sensitive to the drug PLX4720 are as follows Figure 2 shown.
[0076] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A single-cell drug sensitivity prediction algorithm based on semi-supervised transfer learning, characterized by: Includes steps: 1) Data preprocessing: 1-1 Prepare data sets: The data sets required for the model are cell line gene expression data sets, single cell gene expression data sets, and drug sensitivity data sets; the single cell gene expression data sets include a small number of labeled single cell gene expression data sets and a large number of unlabeled single cell gene expression data sets; the drug sensitivity data sets include cell line drug sensitivity data sets and a small number of drug sensitivity data sets corresponding to labeled single cells. 1-2 Standardization: Gene expression data with the same genes in the cell line gene expression dataset and the single cell gene expression dataset are extracted respectively, and the cell line gene expression dataset and the single cell gene expression dataset with the same genes are standardized in the same way. The standardized cell line gene expression dataset includes: The row names are gene names, and the column names are cell line names; The standardized single-cell gene expression dataset includes data with row names as gene names and column names as single cell names; the drug sensitivity dataset includes data with row names as cell names and column names as labels, which are 0 or 1. 1-3 Split the dataset: 80% of the standardized cell line gene expression dataset and the standardized unlabeled single-cell gene expression dataset are used as the model training dataset, and 20% are used as the validation dataset; the standardized labeled single-cell gene expression dataset is divided in half and added to the model training dataset and model validation dataset respectively. 2) Drug sensitivity prediction with labeled data: 2-1 Input data contains N s Source domain data There are labeled samples and N t Labeled target domain data where x s represents the gene expression data of cell lines, y s represents the sensitivity label of the drug to the cell line, x t represents single-cell gene expression data, y t It represents the sensitivity label of the drug to a single cell. We process the dataset using methods 1-2 and 1-3 and then input it into the model. 2-2 Use the shared feature encoder to map cell line and labeled single-cell gene feature data into an l-dimensional feature vector 2-3 In order to obtain more reliable output of the model, make the directions of features from the same class closer to each other, and separate different classes, we use l2 normalization to obtain the feature vector obtained by 2-2 to obtain an l-dimensional feature vector 2-4 A weight vector W C A linear layer composed of a layer is used as a classifier to evaluate the correlation between cell line pharmacogenomics information and drug response at the cell line level. According to 2-3 We use it as the input of the classifier, followed by a softmax layer with a temperature parameter T, to predict the drug sensitivity label at the cell line level. 2-5 In order to make the model less sensitive to small changes in the input, we add a small amount of perturbation near the input data, and the perturbation is in the direction of the loss gradient increase. 3) Semi-supervised domain adaptation 3-1 Input data is N u Unlabeled target domain data Unlabeled samples, where x represents single-cell gene expression data, and N u < <N t , we process the data set using methods 1-2 and 1-3, and then input it into the model; 3-2 Use the same method as 2-2 and 2-3 in step 2) to obtain the l-dimensional feature vector of the unlabeled data 3-3 For semi-supervised domain adaptation, we update the classifier by maximizing the entropy of unlabeled target domain samples and minimize the entropy of unlabeled target domain samples to update the feature extractor, thus forming an adversarial relationship. 3-4 Use adversarial learning to jointly train and update all the models in the previous steps; 4) Unlabeled data drug sensitivity prediction: 4-1 Test dataset preparation: the standardized single-cell gene expression dataset in step 1; 4-2 will obtain the fully trained feature extractor and classifier assembly, and then input the test data set into the model to predict the sensitivity label of single cells to drugs.
2. The single-cell drug sensitivity prediction algorithm based on semi-supervised transfer learning according to claim 1, characterized in that: In step 1), the z-score method is used as the data normalization method.
3. The single-cell drug sensitivity prediction algorithm based on semi-supervised transfer learning according to claim 1, characterized in that: In step 2), the model uses an unbiased linear layer with 128 input dimensions and 2 output dimensions as a classifier.