Interpretable single cell annotation method, device and equipment and storage medium
By using graph interpretive technology based on information bottlenecks in single-cell annotation, predicting and interpreting the single-cell data is solved, and a more stable and reliable interpretive analysis is achieved.
Patent Information
- Application Number
- CN202411981310.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-02
AI Technical Summary
The prior art explains the results in single-cell annotations that are unstable and unreliable, making it difficult to provide a reliable understanding of model prediction logic.
Using graph interpretive technology based on information bottlenecks, cell type tag prediction is performed on the target cancer data set through the first module of the annotation model, and interpretability analysis is performed on the interpreted cancer data set through the second module to generate interpretation results.
While ensuring the accuracy of prediction, it improves the stability and reliability of explanations, avoids inconsistencies caused by relying on post-fact explanation methods, and helps understand the prediction logic of the model.
Smart Images

Figure CN119920329A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cell type annotation, and in particular relates to an interpretable single-cell annotation method, device, equipment and storage medium. Background Art
[0002] Cell type annotation (i.e., determining the cell type labels of unknown cells) is a key step in single-cell data analysis and a prerequisite for studying heterogeneous cell populations. By assigning cell type labels to unknown cells, annotation can reveal cellular heterogeneity between tissues, developmental stages, and different organisms, and help discover regulatory genes that drive heterogeneity, thereby deepening our understanding of cell and gene function. The process of cell type annotation is essentially to assign cells in each single-cell RNA Sequencing (scRNA-seq) data to known cell types. This step is crucial in the downstream analysis and interpretation of single-cell data, and it is also the most time-consuming step. At present, the main methods of cell type annotation include the following:
[0003] (1) Cell type annotation based on known marker genes: Due to the similarity of gene expression patterns between cells, the same cell type is often clustered into a cluster, but the cell type corresponding to each cluster is unknown. This method first clusters the single-cell RNA sequencing data, and then uses the marker genes specifically expressed in each cell cluster to identify the cell identity. Usually, we can use the relevant information about cell type-specific marker genes in the literature and databases, combined with visualization methods such as violin plots and feature expression plots, to manually assign each cluster to the corresponding cell type. Some new annotation tools have also been developed to improve the degree of automation and annotation accuracy. However, manual annotation still has certain challenges, especially differences in marker gene selection may lead to inconsistency in cell type labels in different data sets. This method has high accuracy in clarifying cell types or specific biological systems, and can be combined with specific disease contexts to annotate cell types or states. However, this method is still difficult to handle large-scale data and has limited ability to identify new or rare cell types. As the scale of sequencing data increases, the task of annotation through marker genes becomes increasingly arduous and time-consuming.
[0004] (2) Similarity-based cell type annotation: This method maps query cells to an annotated reference dataset by measuring the similarity between query cells and reference cells. Unlike marker gene methods, similarity-based annotation methods use the expression profile of the entire transcriptome to migrate the labels of reference cells or cell clusters to query cells or clusters with similar expression profiles. This method does not need to rely on marker genes as prior knowledge and can use the entire transcriptome information to provide more accurate annotations. However, most automatic annotation methods use linear models and can only map reference labels to query cells. It is difficult to fully extract the complex nonlinear relationships in gene expression data and cannot fully mine the potential information in gene expression data. In addition, these automatic annotation methods are highly dependent on the quality and diversity of reference data, making it difficult to identify cell types that are not in the reference dataset. At the same time, similarity-based cell type annotation methods cannot overcome batch effects and can only use cell type annotation information in the reference data, ignoring useful information in the target data.
[0005] (3) Cell type annotation based on machine learning and deep learning: This method uses machine learning and deep learning models to learn features from annotated data and automatically perform cell type annotation. Automatic annotation methods based on machine learning usually compare or integrate the query cells to be annotated with the annotated reference cells, and use the model to migrate the labels of the reference cells or cell clusters to the query cells or cell clusters with similar expression characteristics. Single-cell data are high-dimensional and sparse, and neural networks are particularly good at extracting information-rich features and patterns from noisy, heterogeneous, and high-dimensional single-cell RNA sequencing data. Therefore, they perform well in large-scale datasets and scenarios with diverse cell types, showing strong generalization and automation capabilities. Existing methods usually train classifier models on annotated reference cells to predict query cells and achieve automatic annotation. However, the model is easily affected by batch effects and may ignore the unique information in the query data. Therefore, the ideal cell type annotation method should be able to comprehensively utilize the gene expression characteristics of reference data and query data to adapt to changes in different data sources and cell types. In addition, most existing graph neural networks (GNNs) only utilize local features and do not integrate the structure of the global network to obtain richer graph structure information;
[0006] In the field of deep learning, an effective way to enhance trust in cell type annotations is to provide explanations that support them. These explanations clarify the model's predictions for human understanding and can be generated by various methods, such as identifying important substructures in the input data, providing additional examples from the training data, or constructing counterfactual examples by perturbing the input to produce different predictions (providing a contrasting example that changes the prediction to explain the prediction). Existing graph interpretation methods are often susceptible to interference from pseudo-correlated features and may extract features that are falsely correlated with the task goal, resulting in inaccurate model interpretation results. Traditional graph interpretation methods (such as GNNExplainer, GraphMask, etc.) may mistake features that are irrelevant to the task goal but are correlated in the data as key features when identifying key features. These pseudo-correlated features not only affect the model's decision-making, but also increase the model's sensitivity to data bias. In addition, the high nonlinearity and complexity of the GNN model itself lead to a certain conflict between prediction accuracy and interpretability. In order to improve interpretability, it is usually necessary to sacrifice a certain amount of prediction accuracy; while excessive pursuit of prediction accuracy may lead to insufficient interpretability and limit researchers' understanding of the basis for model decision-making. Traditional post-hoc explanation methods rely on trained models, so the explanation effect is greatly affected by the initial model, which easily leads to inconsistent explanation results. This instability reduces the generalization ability of the model on different data sets and affects the ability of researchers to obtain stable and reliable explanation results. Summary of the invention
[0007] The purpose of the present invention is to provide an interpretable single-cell annotation method, device, equipment and storage medium, aiming to solve the problem that the interpretation results of single-cell annotation are unstable and unreliable due to the existing technology.
[0008] In one aspect, the present invention provides an interpretable single-cell annotation method, the method comprising the following steps:
[0009] Based on the source cancer dataset, predicting the cell type label of the target cancer dataset through the first module of the annotation model to obtain the predicted label of each cell in the target cancer dataset, wherein the annotation model is constructed by adopting the graph interpretation technology based on the information bottleneck;
[0010] Based on the healthy data set and the predicted labels, an interpretability analysis is performed on the cancer data set to be interpreted through the second module of the annotation model to obtain an interpretation result, wherein the cancer data set to be interpreted is composed of the cancer cell gene expression matrix of the target cancer data set and the predicted labels.
[0011] Preferably, based on the healthy data set and the predicted label, the step of performing interpretability analysis on the cancer data set to be interpreted by the second module of the annotation model comprises:
[0012] Graphing the healthy data set and the cancer data set to be explained by the sampler of the second module to obtain a graph data set with node labels and graph labels;
[0013] Calculating the attention weight of each graph in the graph dataset by the feature extractor of the second module;
[0014] According to the attention weight, using a random attention mechanism, the subgraph generator of the second module perturbs each graph in the graph dataset to obtain a perturbed graph dataset;
[0015] The perturbation map data set is encoded and predicted by the predictor of the second module to obtain the explanation result.
[0016] Preferably, the step of mapping the healthy data set and the cancer data set to be interpreted by the sampler of the second module comprises:
[0017] Preprocessing the healthy data set and the target cancer data set respectively to obtain a healthy cell gene expression matrix and a cancer cell gene expression matrix;
[0018] Combining the cancer cell gene expression matrix with the prediction label to form the cancer data set to be interpreted;
[0019] According to the cell type and the preset sampling times, the healthy cell gene expression matrix and the cancer data set to be interpreted are randomly sampled with replacement to obtain a plurality of cell sets;
[0020] The K nearest neighbor algorithm is used to construct a map of each of the cell sets to obtain the map data set.
[0021] Preferably, the feature extractor of the second module includes a graph neural network and a multi-layer perceptron, and the step of calculating the attention weight of each graph in the graph data set by the feature extractor of the second module includes:
[0022] Encoding each graph in the graph dataset by using the graph neural network to generate node feature embedding of each graph in the graph dataset;
[0023] According to the node feature embedding, calculating the attention score of each edge in each graph in the graph dataset by the multi-layer perceptron;
[0024] The attention score is transformed using logistic regression to obtain the attention weight.
[0025] Preferably, according to the attention weight, the step of perturbing each graph in the graph dataset by using a random attention mechanism through the subgraph generator of the second module comprises:
[0026] Constructing a Bernoulli distribution according to the attention weights;
[0027] randomly drawing a weight variable from the Bernoulli distribution;
[0028] Each graph in the graph data set is perturbed according to the weight variable to obtain a corresponding perturbed subgraph, and the perturbed subgraphs constitute the perturbed graph data set.
[0029] Preferably, the first module includes a data processing submodule, a domain adaptation feature extraction submodule and an adversarial domain adaptation submodule, wherein the domain adaptation feature extraction submodule includes a shared encoder, a source encoder and a target encoder, and the adversarial domain adaptation submodule includes a domain discriminator, a node classifier, a source decoder and a target decoder.
[0030] Preferably, the shared encoder, the source encoder and the target encoder each include a local graph convolutional neural network and a global graph convolutional neural network, the local graph convolutional neural network includes three graph convolutional layers, the global graph convolutional neural network includes two graph convolutional layers, and the domain discriminator includes a gradient reversal layer and two linear layers.
[0031] In another aspect, the present invention provides an interpretable single-cell annotation device, comprising:
[0032] A label prediction unit, configured to predict cell type labels of a target cancer dataset based on a source cancer dataset by using a first module of an annotation model to obtain a predicted label of each cell in the target cancer dataset, wherein the annotation model is constructed by using a graph interpretative technology based on an information bottleneck;
[0033] The label interpretation unit is used to perform an interpretability analysis on the cancer dataset to be interpreted based on the healthy dataset and the predicted label through the second module of the annotation model to obtain an interpretation result, wherein the cancer dataset to be interpreted is composed of the cancer cell gene expression matrix of the target cancer dataset and the predicted label.
[0034] On the other hand, the present invention also provides a computing device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps described in the above-mentioned interpretable single-cell annotation method are implemented.
[0035] On the other hand, the present invention also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the steps described in the above-mentioned interpretable single-cell annotation method.
[0036] The present invention is based on a source cancer dataset, and predicts cell type labels for a target cancer dataset through a first module of an annotation model to obtain a predicted label for each cell in the target cancer dataset. The annotation model is constructed using a graph interpretative technology based on an information bottleneck. Based on a healthy dataset and predicted labels, an interpretability analysis is performed on a cancer dataset to be interpreted through a second module of the annotation model to obtain an interpretation result, wherein the cancer dataset to be interpreted is composed of a cancer cell gene expression matrix and predicted labels of the target cancer dataset. Thus, by integrating an explanatory mechanism into the model architecture, while ensuring prediction accuracy, the stability and reliability of the interpretation are improved, and the inconsistency caused by reliance on post-interpretation methods is avoided, thereby facilitating understanding of the prediction logic of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is a flowchart for implementing the explainable single-cell annotation method provided in Example 1 of the present invention;
[0038] Figure 2 is a schematic diagram of the structure of an interpretable single-cell annotation device provided in the second embodiment of the present invention;
[0039] Figure 3 This is a schematic diagram of a preferred structure of an interpretable single-cell annotation device provided in the second embodiment of the present invention;
[0040] Figure 4 It is a schematic diagram of the structure of a computing device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0042] It should be understood that, unless otherwise specified, the term "plurality" in the present invention refers to two or more than two, and other quantifiers are similar.
[0043] The terms "first", "second", "third", etc. in the present invention are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, for example, they can be implemented in an order other than those given in the diagrams or descriptions of the embodiments of the present disclosure.
[0044] The specific implementation of the present invention is described in detail below in conjunction with specific embodiments:
[0045] Embodiment 1:
[0046] Figure 1 The implementation process of the explainable single-cell annotation method provided in the first embodiment of the present invention is shown. For the convenience of explanation, only the part related to the embodiment of the present invention is shown, which is described in detail as follows:
[0047] In step S101, based on the source cancer dataset, the first module of the annotation model is used to predict the cell type label of the target cancer dataset to obtain a predicted label for each cell in the target cancer dataset.
[0048] The embodiment of the present invention is applicable to computing devices, such as personal computers, servers, etc. In the embodiment of the present invention, the source cancer dataset is used as a source domain, which is a dataset with labeled cell types, that is, each cell in the dataset corresponds to a known cell type label, which is used to train the annotation model. The target cancer dataset is used as a target domain, which is a dataset without labeled cell types. The source cancer dataset and the target cancer dataset are input into the annotation model. The first module of the annotation model predicts the cell type label of the target cancer dataset based on the source cancer dataset to obtain the predicted label of each cell in the target cancer dataset. The annotation model is constructed using a graph interpretability technology based on Information Bottleneck (IB), thereby improving the interpretability of the annotation model and making the annotations generated by the model easier to understand. At the same time, by compressing redundant information, the model can focus more on features related to the target task, thereby improving the accuracy and efficiency of the annotation.
[0049] In a feasible embodiment, the first module includes a data processing submodule, a domain adaptation feature extraction submodule and an adversarial domain adaptation submodule, wherein the domain adaptation feature extraction submodule includes a shared encoder, a source encoder and a target encoder, and the adversarial domain adaptation submodule includes a domain discriminator, a node classifier, a source decoder and a target decoder. By introducing an adversarial domain adaptation mechanism into the model, the model can effectively migrate shared features between the source domain and the target domain while retaining the unique features of each domain, thereby realizing the knowledge transfer of the labeled information of the source network to the target network, and at the same time overcoming the distribution differences between different data sets, improving the generalization ability of the model on different data sets, and enabling the annotation model to show higher accuracy and robustness on cancer data, which is particularly suitable for batch effects and heterogeneity problems commonly found in biological data.
[0050] In yet another feasible embodiment, prediction of the cell type label of the target cancer dataset is achieved by the following steps:
[0051] ① Process the source cancer dataset and the target cancer dataset through the data processing submodule to obtain the source graph structure and the target graph structure;
[0052] In an embodiment of the present invention, the data processing submodule first preprocesses the received source cancer dataset and target cancer dataset respectively. The preprocessing process includes but is not limited to removing noise, standardizing, unifying cell and gene types, and selecting 2000 highly variable genes with biological significance. These preprocessing measures are intended to improve data quality and ensure the accuracy and reliability of subsequent analysis. After preprocessing, the single-cell gene expression matrix corresponding to each dataset (i.e., the source single-cell gene expression matrix and the target single-cell gene expression matrix) is obtained. Then, the gene expression characteristics of each cell in the single-cell gene expression matrix corresponding to each dataset are used to calculate the similarity between cells (such as Euclidean distance or cosine similarity). Based on the similarity between cells, the K-Nearest Neighbors (KNN) algorithm is used to convert the gene expression information in the source single-cell gene expression matrix and the target single-cell gene expression matrix into corresponding graph structures. In this process, the K value (generally 5) is selected to determine the connection between each cell and its most similar K neighbors, thereby generating the corresponding KNN graph (i.e., source graph structure and target graph structure). In this way, each cell is regarded as a node in the graph, and the similarity relationship between cells is represented by edges. At the same time, the single-cell gene expression matrix, as the attribute information of the node, provides rich biological features for each node in the graph. In addition, during the KNN graph construction process, the data processing submodule will also generate label information for the nodes (i.e., cells) in the graph based on the cell type of each cell in the single-cell gene expression matrix corresponding to each data set. These label information plays an important role in subsequent analysis. Specifically, the label information of the source domain will be used to train the model, while the label information of the target domain will not participate in the training process and is only used to calculate the entropy loss of the classifier to evaluate the performance of the model.
[0053] ② The source graph structure and the target graph structure are extracted through the domain adaptation feature extraction submodule to obtain the source potential representation, the target potential representation and the shared potential representation;
[0054] In the embodiment of the present invention, the domain adaptation feature extraction submodule includes three encoders, namely a shared encoder and two private encoders (these two private encoders are source encoder and target encoder), which work together to extract cross-domain shared features and unique features of each graph structure (source graph structure and target graph structure), thereby realizing the separation of domain private and shared features. Specifically, the shared encoder extracts features that can be shared across the source domain and the target domain according to the source graph structure and the target graph structure, and obtains a more comprehensive shared latent representation. The source encoder extracts specific features of the source domain according to the source graph structure and obtains the source potential representation The target encoder extracts specific features of the target domain according to the target graph structure and obtains the target potential representation.
[0055] In a feasible embodiment, each encoder in the domain adaptation feature extraction submodule includes a local graph convolutional neural network (Graph Convolutional Networks, GCN) and a global GCN, which are respectively used to extract the local consistency and global consistency of the nodes, wherein the local GCN is a variational autoencoder (VariationalAutoencoder, VAE) model based on GCN. Specifically, the local GCN includes three graph convolution layers, denoted as gc1, gc2 and gc3, respectively. gc1 is used to convert the input feature dimension into the hidden layer dimension to extract the primary node representation, and gc2 and gc3 are used to generate the mean (mu) and logarithmic variance (Logarithmic VAE) in VAE, respectively. Variance, LogVar), provides the necessary parameters for the output of the encoder for subsequent reparameterized sampling. The global GCN contains two graph convolutional layers, denoted as gc1′ and gc′2, to extract high-level feature representations of nodes. Specifically, gc1′ is used to perform preliminary feature extraction first, and then activated and randomly deactivated (dropout) to enhance the robustness of the model. gc′2 is used to provide different feature representations for further use in tasks. Therefore, through this dual GCN architecture, a more comprehensive node representation can be obtained.
[0056] In another feasible embodiment, when the source graph structure and the target graph structure are feature extracted by the domain adaptation feature extraction submodule, the domain adaptation feature sub-extraction module generates an adjacency matrix for local GCN and a positive pointwise mutual information (PPMI) matrix for global GCN according to the source graph structure and the target graph structure, wherein the adjacency matrix can reflect the local relationship between nodes, and the PPMI matrix is used to measure the global association strength between nodes, which helps to extract the global relationship information of the nodes. Thus, through the combination of the adjacency matrix and the PPMI matrix, the model can more accurately represent the node relationship in the network, better separate nodes of different categories during the domain adaptation process, and improve the accuracy of node classification or clustering tasks.
[0057] In another feasible embodiment, both the local GCN and the global GCN include a reparameterization function to generate latent representations, but they have different focuses. Specifically, the reparameterization function in the local GCN focuses on generating latent space representations of nodes through the mean and variance of the latent space, which is suitable for variational methods that aim to generate random latent variables for node features. The reparameterization function in the global GCN generates latent representations by combining the outputs of two graph convolutional layers, which is suitable for scenarios that require a balance between different feature representations.
[0058] In another feasible embodiment, both the local GCN and the global GCN use an attention mechanism (Attention module) to perform weighted summation of different feature representations through adaptive weights, thereby improving the model's ability to utilize information at different levels and enhancing the model's flexibility and accuracy.
[0059] ③ Based on the source latent representation, target latent representation and shared latent representation, the cell type of each cell in the target cancer dataset is predicted through the adversarial domain adaptation submodule to obtain the predicted label of the cell type.
[0060] In the embodiment of the present invention, the adversarial domain adaptation submodule is intended to reduce the distribution difference between the source domain and the target domain, thereby realizing cross-domain graph structure data transfer learning. The submodule includes a domain discriminator, a node classifier, and two decoders (i.e., a source decoder for the source domain and a target decoder for the target domain). The domain discriminator is used to distinguish the shared features of the source domain and the target domain through an adversarial training strategy (i.e., the shared encoder obtains the shared potential representation). ), thereby forcing the shared encoder to generate domain-invariant features and reduce the distribution difference between the source and target networks. The source decoder is used to compute the private latent representation and shared latent representation The difference between them is used to reconstruct the adjacency matrix of the source graph structure, thereby ensuring that the features extracted by the source encoder maintain the topological structure of the original network corresponding to the source cancer dataset. The target decoder is used to calculate the private potential representation and shared latent representation The difference between them is used to reconstruct the adjacency matrix of the target graph structure, thereby ensuring that the features extracted by the target encoder maintain the topological structure of the original network corresponding to the target cancer dataset. The node classifier uses a simple linear layer to directly map the input high-dimensional features to the number of categories to achieve label prediction for cells in the target domain. Specifically, the node classifier is first trained on the labeled nodes in the source domain to achieve label prediction capability. Then, by applying transfer learning on the unlabeled data in the target domain, this classification capability is transferred to the target domain, and finally the label prediction of the target domain nodes is achieved, and the predicted label of the cell type of each cell in the target domain is obtained.
[0061] In a feasible embodiment, the domain discriminator includes a gradient reversal layer (GRL) and two linear layers, aiming to distinguish the domain of samples through adversarial training. Specifically, first, the gradient is reversed by GRL to promote the shared encoder to learn domain-indistinguishable features. Then, the input features are mapped to a 10-dimensional latent space through the first linear layer, and nonlinearity is introduced through the ReLU activation function. After the first linear layer, a dropout regularization process of 0.1 is applied to reduce overfitting and improve the generalization ability of the model. After that, the features after dropout processing are compressed to 2 dimensions through the second linear layer for binary classification to determine whether the sample comes from the source domain or the target domain. Through the above structure of the domain discriminator, the model can better learn domain-invariant features, thereby improving the performance of the model in adversarial training.
[0062] Through the above steps ① to ③, the cell type labels of the target cancer dataset are predicted, thereby overcoming the distribution differences between different datasets, improving the generalization ability of the model on different datasets, and enabling the annotation model to show high accuracy and robustness on cancer data, which is particularly suitable for batch effects and heterogeneity problems common in biological data.
[0063] In step S102, based on the healthy data set and the predicted labels, the second module of the annotation model performs an interpretability analysis on the cancer data set to be interpreted to obtain an interpretation result, wherein the cancer data set to be interpreted is composed of the cancer cell gene expression matrix and the predicted labels of the target cancer data set.
[0064] In an embodiment of the present invention, the healthy data set, the target cancer data set and the predicted label of each cell in the target cancer data set obtained in step S101 are input into the second module of the annotation model. The second module first combines the predicted label with the cancer cell gene expression matrix of the target cancer data set to form a cancer data set to be explained. Then, based on the healthy data set, the cancer data set to be explained is subjected to an interpretability analysis to obtain an explanation result. The second module includes a sampler, a feature extractor, a subgraph generator and a predictor.
[0065] In a feasible embodiment, interpretability analysis of the cancer dataset to be interpreted is performed by the second module of the annotation model through the following steps:
[0066] (1) The sampler of the second module constructs a graph of the healthy dataset and the cancer dataset to be explained, and obtains a graph dataset with node labels and graph labels;
[0067] In the implementation of the present invention, when the healthy data set and the cancer data set to be interpreted are mapped by the sampler of the second module, preferably, in the sampler, the healthy data set and the target cancer data set are first preprocessed respectively to obtain a healthy cell gene expression matrix and a cancer cell gene expression matrix, and then the cancer cell gene expression matrix is combined with the prediction label to constitute the cancer data set to be interpreted. After that, the healthy cell gene expression matrix and the cancer data set to be interpreted are randomly sampled with replacement according to the cell type and the preset number of sampling times to obtain several cell sets. Finally, the K nearest neighbor algorithm is used to map each cell set to obtain a graph data set.
[0068] In the embodiment of the present invention, the sampler first performs preprocessing of removing noise, standardizing and unifying cell and gene types on the healthy data set and the target cancer data set, respectively. In order to ensure interpretability, in the preprocessing process, the selection of 2000 highly variable genes is no longer performed. After preprocessing, the healthy cell gene expression matrix corresponding to the healthy data set and the cancer cell gene expression matrix corresponding to the target cancer data set are obtained. Then, the sampler combines the cancer cell gene expression matrix with the predicted label to form a cancer data set to be interpreted. Subsequently, the sampler performs random replacement sampling on the healthy cell gene expression matrix and the cancer data set to be interpreted according to the cell type and the preset sampling number to generate a certain number of different cell sets. Finally, the sampler uses KNN to construct these cell sets into different subsets. Graph, these subgraphs constitute a graph dataset as the input of subsequent interpretability tasks. At the same time, in the process of using KNN to construct the graph, in order to enhance the interpretability of the graph, the known cell type label information is used to assign a corresponding label to each node (i.e., cell). In this way, each node not only represents a cell, but also carries its type information, so that the graph structure can more intuitively reflect the type relationship between cells. Furthermore, additional label information is added to the graph according to the source of the dataset. Specifically, the cancer dataset and the healthy dataset are given different graph labels to distinguish their different sources, which helps to more clearly identify which nodes (cells) come from the cancer dataset and which come from the healthy dataset during the analysis process, further improving the interpretability and accuracy of the analysis.
[0069] (2) Calculate the attention weight of each graph in the graph dataset through the feature extractor of the second module;
[0070] In an embodiment of the present invention, the graph data set obtained by the sampler is used as the input of the feature extractor, and the attention weight of each graph in the graph data set is calculated by the feature extractor, wherein the feature extractor includes a graph neural network (GNN) and a multi-layer perceptron (MLP).
[0071] When calculating the attention weight of each graph in the graph dataset through the feature extractor of the second module, preferably, each graph in the graph dataset is encoded through a graph neural network to generate a node feature embedding of each graph in the graph dataset, and based on the node feature embedding, the attention score of each edge in each graph in the graph dataset is calculated through a multi-layer perceptron, and the attention score is transformed by logistic regression to obtain the attention weight.
[0072] In an embodiment of the present invention, when the feature extractor receives a graph data set input by the sampler, each graph in the graph data set is encoded by GNN to generate node feature embeddings for each graph. Then, these node feature embeddings are processed by MLP to calculate the attention score of each edge (i.e., node pair) in the graph. After that, the attention score is transformed by Sigmoid mapping to obtain the attention weight of each edge. These weights reflect the importance of the edge in the graph, and provide useful information for subsequent tasks such as graph analysis or node classification.
[0073] (3) According to the attention weight, the subgraph generator of the second module uses a random attention mechanism to perturb each graph in the graph dataset to obtain a perturbed graph dataset;
[0074] In an embodiment of the present invention, the subgraph generator of the second module perturbs each graph in the graph data set according to the attention weight and through a random attention mechanism to obtain a perturbed graph data set.
[0075] When the subgraph generator of the second module uses a random attention mechanism to perturb each graph in the graph dataset, preferably, a Bernoulli distribution is constructed according to the attention weight, and weight variables are randomly extracted from the Bernoulli distribution. According to the weight variables, each graph in the graph dataset is perturbed to obtain a corresponding perturbed subgraph, and the perturbed subgraphs constitute a perturbed graph dataset.
[0076] In the embodiment of the present invention, the subgraph generator first constructs a Bernoulli distribution according to the attention weight. The Bernoulli distribution is expressed as Bern(P uv ), where P uv is the attention weight corresponding to the node u and v. Then, Bern(P uv ) as the prior distribution, and randomly extract the weight variable α from it. Then, the subgraph generator uses the weight variable α to perform element-wise multiplication of the adjacency matrix A of each original graph G in the graph dataset (expressed as α⊙A), thereby extracting the perturbed subgraph G after perturbation s At the same time, in order to measure the difference between the original graph G and the perturbed subgraph G s The subgraph generator uses the KL divergence (Kullback-Leibler divergence) to calculate the mutual information I(G,G s ) to encourage the sampling distribution of each edge learned by the model to be as close as possible to the perturbation G s Bernoulli distribution, in other words, minimizing the KL divergence will make the original graph G and the perturbed subgraph G s Minimize the difference between them while keeping the perturbed subgraph G sIn addition, the subgraph generator also uses the Gumbel-Softmax technique for reparameterization, so that the model can perform gradient optimization on the process of random sampling from the Bernoulli distribution, so that the subgraph generator can update its weight parameters and generate a new and more optimized perturbation subgraph G s .
[0077] (4) The perturbation graph dataset is encoded and predicted through the predictor of the second module to obtain the explanation result.
[0078] In an embodiment of the present invention, the predictor of the second module includes a GNN and an MLP. When the predictor receives the perturbation graph data set input by the subgraph generator, the GNN in the predictor encodes each perturbation subgraph in the perturbation graph data set. Specifically, the GNN maps each perturbation subgraph to a high-dimensional latent space by capturing the relationship between nodes and edges in the perturbation subgraph, and extracts node features through multi-layer graph convolution operations, wherein each layer of graph convolution is connected to ReLU activation and dropout operations to increase the nonlinear expression ability of the model and prevent overfitting. After that, global pooling is used. The node features of each perturbed subgraph are summed up by pooling to generate a graph-level embedding. Subsequently, the graph-level embedding obtained by GNN encoding is passed to MLP. MLP further extracts useful information in the graph-level embedding through multiple layers of nonlinear transformations. The output of MLP is normalized by the Softmax layer to obtain the classification prediction label of the node or graph. Finally, the predictor generates an explanation result containing the most important substructure of each perturbed subgraph according to the classification prediction label. Specifically, the predictor determines the most important substructure in each perturbed subgraph by analyzing the changes in edge weights during GNN encoding. That is, the higher the edge weight, the more important the edge is in the perturbed subgraph and the greater its contribution to the prediction result. The explanation result can be presented in a visual form (such as highlighting important edges) or in a text form (listing important edges and their weights) to help users intuitively understand which substructures have an important impact on the prediction result.
[0079] The above steps (1) to (4) effectively eliminate pseudo-correlated features through the random attention mechanism and retain the key features that are truly relevant to the task results, so that the model can provide decision-making basis while generating annotation results, which improves the interpretability of the model and the credibility of the results, and helps to understand the logic behind the annotation results. At the same time, the embedded explanatory mechanism enhances the stability of the annotation results and avoids the inconsistency caused by relying on post-explanation methods.
[0080] In an embodiment of the present invention, based on a source cancer dataset, a cell type label prediction is performed on a target cancer dataset through a first module of an annotation model to obtain a predicted label for each cell in the target cancer dataset. The annotation model is constructed using a graph interpretative technology based on an information bottleneck. Based on a healthy dataset and predicted labels, an interpretability analysis is performed on the cancer dataset to be interpreted through a second module of the annotation model to obtain an interpretation result, wherein the cancer dataset to be interpreted consists of a cancer cell gene expression matrix and predicted labels of the target cancer dataset. Thus, by integrating the explanatory mechanism into the model architecture, while ensuring the accuracy of the prediction, the stability and reliability of the interpretation are improved, and the inconsistency caused by relying on post-interpretation methods is avoided, which helps to understand the prediction logic of the model.
[0081] Embodiment 2:
[0082] Figure 2 The structure of an interpretable single-cell annotation device provided by the second embodiment of the present invention is shown. For the convenience of description, only the part related to the embodiment of the present invention is shown, including:
[0083] A label prediction unit 21 is used to predict the cell type label of the target cancer dataset based on the source cancer dataset by using the first module of the annotation model to obtain a predicted label of each cell in the target cancer dataset, wherein the annotation model is constructed by using a graph interpretation technology based on an information bottleneck;
[0084] The label interpretation unit 22 is used to perform interpretability analysis on the cancer dataset to be interpreted based on the healthy dataset and the predicted labels through the second module of the annotation model to obtain an interpretation result, wherein the cancer dataset to be interpreted is composed of the cancer cell gene expression matrix and the predicted labels of the target cancer dataset.
[0085] Preferably, if Figure 3 As shown, the label interpretation unit 22 includes:
[0086] A data mapping unit 221 is used to map the healthy data set and the cancer data set to be interpreted through the sampler of the second module to obtain a graph data set with node labels and graph labels;
[0087] A weight calculation unit 222, used to calculate the attention weight of each graph in the graph data set through the feature extractor of the second module;
[0088] The original image perturbation unit 223 is used to perturb each image in the image data set by using a random attention mechanism through the sub-image generator of the second module according to the attention weight to obtain a perturbed image data set;
[0089] The result obtaining unit 224 is used to encode and predict the perturbation map data set through the predictor of the second module to obtain an explanation result.
[0090] In the embodiment of the present invention, each unit of the interpretable single cell annotation device can be implemented by a corresponding hardware or software unit, and each unit can be an independent software or hardware unit, or can be integrated into a software or hardware unit, which is not intended to limit the present invention. Specifically, the implementation of each unit can refer to the description of the aforementioned embodiment 1, which will not be repeated here.
[0091] Embodiment three:
[0092] Figure 4 The structure of a computing device provided in the third embodiment of the present invention is shown. For the convenience of description, only the part related to the embodiment of the present invention is shown.
[0093] The computing device 4 of the embodiment of the present invention includes a processor 40, a memory 41, and a computer program 42 stored in the memory 41 and executable on the processor 40. When the processor 40 executes the computer program 42, the steps of the above-mentioned interpretable single-cell annotation method embodiment are implemented, for example Figure 1 Alternatively, when the processor 40 executes the computer program 42, the functions of each unit in the above-mentioned device embodiments are realized, for example Figure 2 Function of the unit shown.
[0094] In an embodiment of the present invention, based on a source cancer dataset, a cell type label prediction is performed on a target cancer dataset through a first module of an annotation model to obtain a predicted label for each cell in the target cancer dataset. The annotation model is constructed using a graph interpretative technology based on an information bottleneck. Based on a healthy dataset and predicted labels, an interpretability analysis is performed on the cancer dataset to be interpreted through a second module of the annotation model to obtain an interpretation result, wherein the cancer dataset to be interpreted consists of a cancer cell gene expression matrix and predicted labels of the target cancer dataset. Thus, by integrating the explanatory mechanism into the model architecture, while ensuring the accuracy of the prediction, the stability and reliability of the interpretation are improved, and the inconsistency caused by relying on post-interpretation methods is avoided, which helps to understand the prediction logic of the model.
[0095] The computing device of the embodiment of the present invention may be a personal computer or a server. The steps implemented when the processor 40 in the computing device 4 executes the computer program 42 to implement an interpretable single cell annotation method can be referred to the description of the aforementioned method embodiment, which will not be repeated here.
[0096] Embodiment 4:
[0097] In an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned interpretable single-cell annotation method embodiment are implemented, for example, Figure 1 Alternatively, when the computer program is executed by a processor, the functions of each unit in the above-mentioned device embodiments are realized, for example Figure 2 Function of the unit shown.
[0098] In an embodiment of the present invention, based on a source cancer dataset, a cell type label prediction is performed on a target cancer dataset through a first module of an annotation model to obtain a predicted label for each cell in the target cancer dataset. The annotation model is constructed using a graph interpretative technology based on an information bottleneck. Based on a healthy dataset and predicted labels, an interpretability analysis is performed on the cancer dataset to be interpreted through a second module of the annotation model to obtain an interpretation result, wherein the cancer dataset to be interpreted consists of a cancer cell gene expression matrix and predicted labels of the target cancer dataset. Thus, by integrating the explanatory mechanism into the model architecture, while ensuring the accuracy of the prediction, the stability and reliability of the interpretation are improved, and the inconsistency caused by relying on post-interpretation methods is avoided, which helps to understand the prediction logic of the model.
[0099] The computer-readable storage medium of the embodiment of the present invention may include any entity or device or recording medium capable of carrying computer program code, for example, ROM / RAM, magnetic disk, optical disk, flash memory and other memories.
[0100] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. An interpretable single-cell annotation method, characterized in that: The method comprises the following steps: Based on the source cancer dataset, predicting the cell type label of the target cancer dataset through the first module of the annotation model to obtain the predicted label of each cell in the target cancer dataset, wherein the annotation model is constructed by adopting the graph interpretation technology based on the information bottleneck; Based on the healthy data set and the predicted labels, an interpretability analysis is performed on the cancer data set to be interpreted through the second module of the annotation model to obtain an interpretation result, wherein the cancer data set to be interpreted is composed of the cancer cell gene expression matrix of the target cancer data set and the predicted labels.
2. The method according to claim 1, characterized in that Based on the healthy data set and the predicted label, the step of performing an interpretability analysis on the cancer data set to be explained by the second module of the annotation model comprises: Graphing the healthy data set and the cancer data set to be explained by the sampler of the second module to obtain a graph data set with node labels and graph labels; Calculating the attention weight of each graph in the graph dataset by the feature extractor of the second module; According to the attention weight, using a random attention mechanism, the subgraph generator of the second module perturbs each graph in the graph dataset to obtain a perturbed graph dataset; The perturbation map data set is encoded and predicted by the predictor of the second module to obtain the explanation result.
3. The method according to claim 2, characterized in that The step of mapping the healthy data set and the cancer data set to be interpreted by the sampler of the second module comprises: Preprocessing the healthy data set and the target cancer data set respectively to obtain a healthy cell gene expression matrix and a cancer cell gene expression matrix; Combining the cancer cell gene expression matrix with the prediction label to form the cancer data set to be interpreted; According to the cell type and the preset sampling times, the healthy cell gene expression matrix and the cancer data set to be interpreted are randomly sampled with replacement to obtain a plurality of cell sets; The K nearest neighbor algorithm is used to construct a map of each of the cell sets to obtain the map data set.
4. The method according to claim 2, characterized in that The feature extractor of the second module includes a graph neural network and a multi-layer perceptron, and the step of calculating the attention weight of each graph in the graph data set by the feature extractor of the second module includes: Encoding each graph in the graph dataset by using the graph neural network to generate node feature embedding of each graph in the graph dataset; According to the node feature embedding, calculating the attention score of each edge in each graph in the graph dataset by the multi-layer perceptron; The attention score is transformed using logistic regression to obtain the attention weight.
5. The method according to claim 2, characterized in that According to the attention weight, the step of perturbing each graph in the graph dataset by using a random attention mechanism through the subgraph generator of the second module comprises: Constructing a Bernoulli distribution according to the attention weights; randomly drawing a weight variable from the Bernoulli distribution; Each graph in the graph data set is perturbed according to the weight variable to obtain a corresponding perturbed subgraph, and the perturbed subgraphs constitute the perturbed graph data set.
6. The method according to claim 1, characterized in that The first module includes a data processing submodule, a domain adaptation feature extraction submodule and an adversarial domain adaptation submodule, wherein the domain adaptation feature extraction submodule includes a shared encoder, a source encoder and a target encoder, and the adversarial domain adaptation submodule includes a domain discriminator, a node classifier, a source decoder and a target decoder.
7. The method according to claim 6, characterized in that The shared encoder, the source encoder and the target encoder each include a local graph convolutional neural network and a global graph convolutional neural network, the local graph convolutional neural network includes three graph convolutional layers, the global graph convolutional neural network includes two graph convolutional layers, and the domain discriminator includes a gradient reversal layer and two linear layers.
8. An interpretable single-cell annotation device, characterized in that: The device comprises: A label prediction unit, configured to predict cell type labels of a target cancer dataset based on a source cancer dataset by using a first module of an annotation model to obtain a predicted label of each cell in the target cancer dataset, wherein the annotation model is constructed by using a graph interpretative technology based on an information bottleneck; The label interpretation unit is used to perform an interpretability analysis on the cancer dataset to be interpreted based on the healthy dataset and the predicted label through the second module of the annotation model to obtain an interpretation result, wherein the cancer dataset to be interpreted is composed of the cancer cell gene expression matrix of the target cancer dataset and the predicted label.
9. A computing device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Cell type annotation method and annotation system
CN120932748A