An intelligent recognition method for intercellular communication in spatial transcriptome data
By combining single-cell and spatial transcriptome data and using graph attention neural networks, the accurate positioning and analysis of intercellular communication at single-cell resolution is solved, efficient identification and understanding of intercellular communication is achieved, and the research depth of biological processes is improved.
Patent Information
- Application Number
- CN202410747250.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-06-11
AI Technical Summary
The prior art is difficult to accurately identify intercellular communication in single-cell resolution, especially to accurately locate and distinguish different cell types on spatial scale, analyze molecular signaling and interaction patterns between cells, resulting in limited understanding of biological processes.
Combining single-cell transcriptome and spatial transcriptome data, a graph attention neural network and self- and cross-attention mechanism are used to construct an intercellular communication recognition model through preprocessing, feature extraction and model optimization, and a graph attention neural network is used to perform deep learning and prediction of intercellular relationships.
Accurate identification of intercellular communication at single-cell resolution enhances understanding of coordinated intercellular activities, provides important insights into cell signaling networks and regulatory mechanisms, and supports new therapeutic strategies and precise medical applications.
Smart Images

Figure CN118711671B_ABST
Abstract
Description
Technical field
[0001] The present invention relates to the technical field of bioinformatics, and specifically, to an intelligent method for identifying cell - cell communication for spatial transcriptome data. Background technique
[0002] Complex multicellular organisms rely on coordinated mechanisms inside and outside tissues to maintain balance and respond appropriately to internal and external disturbances. In this coordination process, cell - cell communication plays a key role. Cell - cell communication can be achieved through multiple mechanisms, including biochemical signals and physical signals. The transmission of biochemical and physical signals between cells directly affects the phenotype and function of cells. Biochemical signals include molecular signal substances such as hormones, neurotransmitters, and cytokines, while physical signals cover cell - contact signals. Cell - cell communication plays an important role in physiological and pathological processes, including cell differentiation, tissue development, immunity, and the occurrence and development of diseases. In - depth study of cell - cell communication mechanisms helps to reveal the functions of multicellular organisms and disease mechanisms, and provides new directions for disease prevention, diagnosis, and treatment.
[0003] Traditionally, the study of cell - cell communication has been limited to selected genes at the resolution of a few cell types and cell populations. The emergence of single - cell transcriptomics has enabled the study of tissues with unprecedented genomic coverage. Currently, some computational tools have been developed to infer the intensity of cell - cell communication using single - cell transcriptome data. However, these methods have limitations in capturing cell spatial information, resulting in a high false - positive rate in the inferred cell - cell communication. Therefore, in order to more accurately study the impact of cell - cell communication on various physiological processes, it is necessary to incorporate spatial information into the analysis of cell - cell interactions. Spatial transcriptomics can measure the spatial gene expression of two - dimensional or three - dimensional tumor tissue samples, thus providing important spatial information for studying cell - cell communication. However, current methods based on spatial transcriptome data have not overcome the limitations in gene throughput and spatial resolution, and still have difficulty in inferring cell - cell communication at the single - cell resolution. From a biological perspective, cell - cell communication does not function at the level of cell populations; instead, this interaction occurs between individual cells. This limitation affects the understanding of the coordinated activities of cells in various biological processes. Therefore, it is necessary to further develop improved methods to overcome the current technical limitations and truly achieve accurate inference and interpretation of cell - cell communication at the single - cell level to comprehensively understand the complexity of various biological processes.
[0004] Currently, the challenge in identifying cell - cell communication at single - cell resolution lies in accurately localizing and differentiating different cell types at the spatial scale, and parsing the molecular signal transduction and interaction patterns between cells. This information is crucial for comprehensively understanding cell - cell communication in biological processes. Integrating spatial transcriptomic data and single - cell transcriptomic data using artificial intelligence can compensate for the limitations of current data and enable accurate identification of cell - cell communication at single - cell resolution. Through the intelligent identification of cell - cell communication, it is possible to better understand the coordinated activities between different cells, thereby providing researchers with important insights into cell signal transduction networks and regulatory mechanisms. This helps in developing new treatment strategies and precision medicine applications, such as immunotherapy in cancer treatment. In addition, through in - depth understanding of cell - cell communication, it is also possible to reveal the underlying mechanisms of important biological processes such as cell differentiation, development, and dysfunction, and provide new directions and breakthroughs for life science research. Summary of the Invention
[0005] Aiming at the defects existing in the above - mentioned cell - cell communication recognition method, the present invention aims to provide an efficient and accurate intelligent recognition method for cell - cell communication in spatial transcriptomic data by using spatial transcriptomic data and combining artificial intelligence algorithms. Through the combination of bioinformatics, computational biology, and artificial intelligence technologies, a comprehensive understanding of the complex communication between cells is achieved.
[0006] To achieve the above - mentioned goal, the technical solution adopted by the present invention is: an intelligent recognition method for cell - cell communication in spatial transcriptomic data, comprising the following steps:
[0007] I. Single - cell and spatial data pre - processing: Perform quality control and pre - processing on single - cell transcriptomic data and spatial transcriptomic data to reduce non - biological signal noise generated during experiments or sequencing, and eliminate technical variability caused by different experimental conditions or technologies, providing a more reliable data basis for subsequent model construction and analysis;
[0008] II. Integrating single - cell and spatial data: Use a graph attention neural network to mine information from the pre - processed single - cell transcriptomic data and spatial transcriptomic data: Use the self - attention mechanism to mine the gene expression information within the single - cell transcriptomic data and spatial transcriptomic data, and then use the cross - attention mechanism to mine the association information between single - cell omics and spatial omics, finally obtaining single - cell gene expression data with spatial position information;
[0009] III. Digitization of cell features: Using the single - cell gene expression data, construct a sub - graph for each cell, which is used to reveal the relationship between the corresponding cell and its surrounding cells, extract local features from the sub - graph, and then use a sub - graph encoder to integrate these local features to further generate global features;
[0010] IV. Construction of the cell - cell communication recognition model: Use a graph attention neural network to process single - cell gene expression features integrated with spatial positions, and utilize the attention mechanism to refine the relationships between cells; optimize the graph attention neural network by means of pre - training and fine - tuning: First, pre - train the model on a large - scale dataset containing thousands to tens of thousands of samples; then, fine - tune the model using known signal transduction pathways between specific cells to ensure that the model can accurately identify these biologically important interactions.
[0011] The above step one includes the following steps:
[0012] 1.1) First, eliminate the cells in the single - cell data that capture fewer than 200 genes.
[0013] 1.2) Subsequently, normalize the expression matrix of the single - cell data through formula (1):
[0014]
[0015] where C ij represents the original read of the i - th gene in the j - th cell, that is, the expression level of this gene directly obtained from the sequencing data; D ij represents the expression level of the i - th gene in the j - th cell after normalization, that is, the normalized read of this gene; is the median of the number of transcripts detected in each cell, used for normalization to eliminate the differences in sequencing depth between different cells; M represents the total number of cells, that is, the number of columns in the expression matrix;
[0016] 1.3) For the ST dataset with more than 1,000 detected genes, calculate the coefficient of variation CV of each gene using the following equation (2) i :
[0017]
[0018] where, σ i represents the standard deviation of the spatial distribution of the i - th gene at all detection points, u i represents the average expression level of the i - th gene at all detection points; according to the value of the coefficient of variation CV i select the top 1,000 genes with the highest variability as high - variability genes; then, compare these genes with the genes detected in the corresponding single - cell gene sequencing data to establish the ground truth for each dataset.
[0019] The above step two includes the following steps:
[0020] 2.1) Generate the matching descriptor f using the graph attention neural network i ∈R D ,where f i represents the feature descriptor of the i-th cell, which is a vector composed of D real numbers. Here, R represents the set of real numbers, and D is the dimension of the feature vector. Using a point encoder, the transcriptomic data of each cell and its spatial location information are combined to specifically represent each cell. A multi-layer perceptron (MLP) is used to embed the cell location into a high-dimensional vector, which is expressed as:
[0021]
[0022] where represents the initial feature vector of the i-th cell, which integrates gene expression data and cell location information; d i is the gene expression data obtained from the transcriptome, and p i represents the cell location information; MLP enc is a multi-layer perceptron encoder used to embed cell location information into a high-dimensional space;
[0023] 2.2) Create a graph to integrate single-cell transcriptomic data and spatial transcriptomic data. The nodes of the graph represent cells from transcriptomics. Self-connections connect each cell i to all other cells in the same omics, while cross-connections connect cell i to all other cells in different omics. Use the information propagation equation to enable information to propagate along self-connections and cross-connections;
[0024] 2.3) Represent the pairwise scores as a similarity matrix M ∈ R (m×n) to capture the similarity of the matching descriptors;
[0025]
[0026] Here, <·,·> represents the inner product; are the feature vectors of cells in the single-cell transcriptome and spatial transcriptome, respectively. A and B represent the sets of cells in the single-cell transcriptome and spatial transcriptome, respectively. m is the number of cells in the single-cell transcriptome, and n is the number of cells in the spatial transcriptome. To obtain the mapping matrix, minimize the following objective function for M:
[0027]
[0028] where cos sim is the cosine similarity function used to measure the similarity between two vectors; * represents matrix slicing; is the objective function used to optimize the similarity matrix M to obtain better cell matching; n genes and nspots represent the number of genes and the number of spatial positions respectively; (M T A *,k ) and B *,k represent the k-th columns of matrices M T A and B respectively; (M T A *,j ) and B *,j represent the j-th columns of matrices M T A and B respectively.
[0029] Step 3 above includes the following steps:
[0030] 3.1) Perform random walks on the graph created in Step 2. During the walks, mask some nodes and predict these nodes subsequently to capture the overall connection pattern of the graph; for each node and node pair in the graph, generate corresponding subgraphs g c , and these subgraphs g c constitute the set G c of pre-trained subgraphs; each subgraph g c is represented by a set of where v i is a node in subgraph g c , i is the index of the node, i = 1, 2... |V c |, and |V c | represents the number of nodes in g c ;
[0031] 3.2) For each node v c in subgraph g i , map its attributes and structure-based embeddings to a stacked vector through the function f attr (.), and then use the learnable embedding matrix W e to convert this vector into a low-dimensional representation h i . The embeddings of all nodes in subgraph g c are collectively represented as H c , and the initialization of these embeddings is achieved through the output embeddings obtained by the global feature generation method that captures the multi-relational graph structure. At the same time, when initializing the node features, the pre-trained node representation vectors are merged with the gene expression data obtained from the cells;
[0032] 3.3) For subgraph g c , where V c represents the set of nodes. For each node belonging to V c , there is a global input embedding, and these embeddings are represented by the matrix . Matrix H cIt not only contains the embeddings of all nodes, but also these embeddings are global, that is, they consider the structural information of the entire graph; further, through context learning, these global embeddings are transformed into new embeddings to reflect the most representative roles of nodes in g c ; meanwhile, in order to capture the high-order relationship dependencies between nodes, a semantic association matrix is introduced. This matrix serves as an asymmetric weight matrix, reflecting the different influences among units within the subgraph; in each transformation layer of the subgraph g c , and the global graph G, the weights of the matrix are iteratively learned to consider the local and global connection relationships between nodes.
[0033] Step 4 above includes the following steps:
[0034] 4.1) For each node in the graph constructed by fusing single-cell transcriptome data and spatial transcriptome data, create a node subgraph g c with a diameter equal to the maximum value of the shortest path distances between any two nodes in the graph. In the subgraph g c , randomly select a node v m for masking, and ensure that the structure of the graph remains unchanged during the process. Then, in the case of the given subgraph g c as the context, pre-train the model by maximizing the probability of correctly predicting the masked node v m . The formula for this probability is as follows:
[0035]
[0036] where θ represents the set of model parameters, G C is the set of all generated subgraphs, g c represents one of them, and p(v m |g c ,θ) represents the probability of correctly predicting the masked node v c under the condition of the given subgraph g m and the model parameters θ;
[0037] 4.2) For each pair of nodes ready for link prediction, generate multiple contexts. Under the condition of the given context g cp , the model is trained by maximizing the probability of observing the positive edge e p . Meanwhile, the model also performs negative sampling to learn to assign lower probabilities to the negatively sampled edge e n and its corresponding context g cn ; the training objective is to combine the positive edge dataset D p and the negative edge dataset D nConstructed by optimizing this training objective, the model can enhance its ability to accurately predict positive and negative edges. The formula of the training objective is as follows:
[0038]
[0039] In Equation (7), L represents the training objective function, e p represents the positive edge, e n represents the negative edge, g cp and g cn represent the contexts related to the positive and negative edges in sequence, θ represents the model parameters, and P(e p |g cp , θ) represents the probability of observing the positive edge e cp under the condition of the given context g p and the model parameters θ, and P(e n |g cn , θ) represents the probability of observing the negative edge e cn under the condition of the given context g n and the model parameters θ;
[0040] When calculating the probability of the edge between two nodes, expressed as e = (v i , v j ), the similarity score S(v i , v j ) is used for calculation, and its formula is: Here, and are the embedding vectors of nodes v i , v j respectively, and σ(·) is the sigmoid function.
[0041] Advantages of the present invention: The present invention utilizes cell-specific gene expression data and the spatial positions of cells, and uses graph neural networks and subgraph-based graph attention neural networks to systematically reveal the mechanisms of cell communication in normal and diseased tissues. By combining spatial transcriptomics data with single-cell transcriptomics data obtained from the same region, the problems of limited gene throughput and insufficient resolution of spatial omics data are effectively solved, thereby accurately revealing cell communication at the single-cell resolution in space. The advantages of the present invention specifically include the following:
[0042] I. The present invention combines single-cell transcriptomics data and spatial transcriptomics data, enhancing the ability to analyze cell-cell communication. For single-cell spatial transcriptomics datasets, the present invention adopts a similarity-based classification strategy, accurately classifying and analyzing single-cell spatial omics data by selecting the most similar and highest-ranked cell clusters. Meanwhile, using a graph neural network based on the attention mechanism, the relevant relationships and patterns between cells are identified, improving the accuracy and reliability of classification. For non-single-cell spatial transcriptomics datasets, by selecting and mapping the optimal combination of specific cells, the spatial transcriptomics distribution is reconstructed, and the relationships and spatial distribution patterns between cells are revealed, enhancing the interpretability and information accuracy of the data.
[0043] II. The present invention reveals the trends of these preferences in various spatial transcriptomics datasets by identifying the interaction preferences between different cell types. By accurately characterizing the spatial distribution of each cell type at the reconstructed single-cell resolution, information on the proximity relationship between different cell types is provided. The subgraph-based graph attention neural network considers the intercellular relationships at different levels, helping to construct a cell-cell communication network and revealing complex intercommunication relationships, which is helpful for more accurately determining the communication patterns and relationship characteristics between different cell types.
[0044] III. The present invention comprehensively and accurately infers cell-cell communication mediated by different ligand-receptor pairs by revealing ligand-receptor pairs between cells and their roles in cell-cell communication. By predicting and visualizing cell-cell communication at single-cell resolution and analyzing the relevant ligand-receptor pairs, a method for comprehensively and holistically understanding cell-cell communication is provided, which is helpful for understanding physiological and pathological processes and developing related treatment methods.
[0045] IV. The present invention provides a more comprehensive and clear perspective for understanding the cell-cell communication mechanism. By combining data from different omics and applying advanced neural network algorithms, the accuracy of identification is greatly improved. The present invention has independent innovation in the scheme combination and setting, providing new ideas for further studying the occurrence and development of diseases and providing an important basis for developing new treatment strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is the flowchart of intelligent identification of cell-cell communication for spatial transcriptome data of the present invention;
[0047] Figure 2 is the flowchart for constructing DeepTalk, the intelligent identification method of cell-cell communication of the present invention;
[0048] Figure 3 is the performance result graph of the intelligent identification method DeepTalk of the present invention for integrating single-cell data and spatial omics data;
[0049] Figure 4 It is a graph showing the prediction performance results of the intercellular communication intelligent recognition method DeepTalk of the present invention on single-cell spatial omics data.
[0050] Figure 5 It is a graph showing the prediction performance results of the intercellular communication intelligent recognition method DeepTalk of the present invention on non-single-cell spatial omics data. Detailed implementation manners
[0051] The following combines the accompanying drawings to describe in detail a specific implementation manner of the present invention. It should be understood that the protection scope of the present invention is not limited by the specific implementation manner.
[0052] See Figures 1 - 5 , the present invention discloses an intercellular communication intelligent recognition method for spatial transcriptome data.
[0053] As Figure 1 shown, the method of the present invention includes the following steps:
[0054] S1: Single-cell and spatial data preprocessing: This step aims to perform quality control and preprocessing on single-cell transcriptome data and spatial transcriptome data. This involves reducing non-biological signal noise generated during the experiment or sequencing, such as background noise and sequencing errors, and eliminating technical variability caused by different experimental conditions or techniques, such as batch effects and sequencing platform differences. Through this step, a more reliable data basis can be provided for subsequent model construction and analysis;
[0055] S1.1. Remove cells with poor quality; that is, remove cells in the single-cell data that capture fewer than 200 genes.
[0056] S1.2. Standardize the expression matrix of the single-cell data
[0057]
[0058] C ij represents the original reads of the i-th gene in the j-th cell, that is, the expression level of this gene directly obtained from the sequencing data; D ij represents the expression level of the i-th gene in the j-th cell after normalization processing, that is, the normalized reads of this gene; is the median of the number of transcripts detected in each cell, which is used for normalization processing to eliminate the differences in sequencing depth between different cells; M represents the total number of cells, that is, the number of columns in the expression matrix.
[0059] S1.3. For ST datasets with more than 1,000 detected genes, use the following equation to calculate the coefficient of variation (CV) of each genei )
[0060]
[0061] where σ i represents the standard deviation of the spatial distribution of the i-th gene across all detection points, and u i represents the average expression level of the i-th gene across all detection points; based on the value of the coefficient of variation (CV i ), the top 1,000 genes with the highest variability were selected as highly variable genes; subsequently, these genes were compared with the genes detected in the corresponding single-cell RNA sequencing data to establish the ground truth for each dataset.
[0062] S2: Integrating single-cell and spatial data: Using the preprocessed single-cell transcriptome data and spatial transcriptome data, in-depth information mining was performed through a graph attention neural network. Specifically, the self-attention mechanism was used to mine the gene expression information within single-cell omics, while the cross-attention mechanism was used to mine the association information between single-cell omics and spatial omics. Finally, single-cell gene expression data with clear spatial location information was obtained. As Figure 2 shown, based on the processed single-cell and spatial data, using a graph attention neural network, the self-attention mechanism was used to mine the information within omics, while the cross-attention mechanism was used to mine the data between omics, and single-cell data with spatial locations was obtained.
[0063] S2.1 Using a graph attention neural network to generate a matching descriptor f i ∈R D where f i represents the feature descriptor of the i-th cell, and D is the dimension of the feature vector. This step is achieved by enabling feature communication between the initial features. To ensure robust matching both within and outside omics, the aggregation of remote features is particularly important. Using a point encoder, the transcriptome data of each cell and its spatial location information are combined to specifically represent each cell. To enhance the ability of the graph attention neural network in subsequent inferences, we use a multi-layer perceptron (MLP) to embed the cell location into a high-dimensional vector, denoted as:
[0064]
[0065] where represents the initial feature vector of the i-th cell, which combines gene expression data and cell location information; d i is the gene expression data obtained from the transcriptome, and p i represents the cell location information; the point encoder enables the graph convolutional neural network to utilize d simultaneously in the subsequent inference stagei and p i ; MLP enc is a multi - layer perceptron encoder used to embed cell position information into a high - dimensional space.
[0066] S2.2 Create a graph to integrate single - cell transcriptomic data and spatial transcriptomic data. The nodes of the graph represent cells from transcriptomics; self - connections connect each cell i to all other cells in the same omics, while cross - connections connect cell i to all other cells in different omics; use the message - passing equation to enable information to propagate along self - connections and cross - connections;
[0067] S2.3 Represent the pairwise scores as a similarity matrix M ∈ R(m×n) to capture the similarity of matching descriptors.
[0068]
[0069] Here, <·,·> represents the inner product; are the feature vectors of cells in single - cell transcriptomics and spatial transcriptomics respectively. A and B represent the sets of cells in single - cell transcriptomics and spatial transcriptomics respectively; m is the number of cells in single - cell transcriptomics, and n is the number of cells in spatial transcriptomics. The matching descriptors are not normalized, and their magnitudes may vary for each feature during the training process to reflect the confidence level of the prediction; to obtain the mapping matrix, the following objective function is minimized for M:
[0070]
[0071] where cos sim is the cosine similarity function used to measure the similarity between two vectors; * represents matrix slicing; is the objective function used to optimize the similarity matrix M to obtain better cell matching; n genes and n spots represent the number of genes and the number of spatial positions respectively; (MTA *,k ) and B *,k represent the k - th columns of matrices M T A and B respectively; (M T A *,j ) and B *,j represent the j - th columns of matrices M T A and B respectively. The integration effect on multiple datasets is as Figure 3 shown.
[0072] S3: Digitalization of cell features. This step specifically processes the integrated single-cell data. By using the gene expression data of these single cells, a subgraph is constructed for each cell, which can fully reveal the relationship between the cell and its surrounding cells, and local features are extracted from it. Subsequently, a subgraph encoder is used to integrate these local features to further generate global features. As Figure 2 shown, using gene expression data, a subgraph is established for each cell to fully explore the relationship between the cell and its surrounding cells and establish local features. Through the subgraph encoder, the local features of the cells are integrated and global features are generated.
[0073] S3.1. Perform random walks on the graph. Here, the graph refers to the graph integrating two types of omics data constructed in step two, and its nodes represent cells of transcriptomics. During the walk, some nodes are masked and these nodes are attempted to be predicted later. The purpose of doing this is to capture the overall connection pattern of the graph. For each node and node pair in the graph, a corresponding subgraph g c is generated, and these subgraphs form a set G c of pre-trained subgraphs. Each subgraph g c is represented by a set where |V c | represents the number of nodes in g c ;
[0074] S3.2 For each node v c in the subgraph g i , where i is the index of the node, its attributes and structure-based embeddings are mapped to a stacked vector through the function f attr (.), and then a learnable embedding matrix W e is used to convert this vector into a low-dimensional representation h i . The embeddings of all nodes in the subgraph g c are collectively represented as H c . In addition, these embeddings can be initialized using the output embeddings obtained from other global feature generation methods that capture the multi-relational graph structure. For the initialization of node features, the pre-trained node representation vectors are combined with the gene expression data obtained from the cells;
[0075] S3.3 For the subgraph g c , where V c represents the set of nodes. For each node belonging to V c , there is a global input embedding, and these embeddings are represented by the matrix . The matrix H cIt not only contains the embeddings of all nodes, but also these embeddings are global, that is, they consider the structural information of the entire graph. Further, through context learning, these global embeddings are transformed into new embeddings to reflect the most representative roles of nodes in g c ; meanwhile, in order to capture the high-order relational dependencies between nodes, a semantic association matrix is introduced. This matrix serves as an asymmetric weight matrix, reflecting the different influences among units within the subgraph; in each transformation layer of the subgraph g c , and the global graph G, the weights of the matrix are iteratively learned to consider the local and global connection relationships between nodes;
[0076] S4: Construction of the cell-cell communication recognition model. This step uses a graph attention neural network to process and fuse single-cell gene expression features with spatial positions to extract the association information between cells. Subsequently, through the attention mechanism, the complex interaction network between cells is deeply analyzed and determined. To ensure the accuracy and generalization ability of the recognition model, a method combining pre-training and fine-tuning is used to optimize the graph attention neural network. First, pre-training is performed on a large-scale dataset containing thousands to tens of thousands of samples to help the recognition model master the general laws of cell-cell communication and thus improve its generalization performance. Then, fine-tuning is carried out for known specific cell signaling pathways to enable the recognition model to understand and adapt to the subtle features of these pathways.
[0077] S4.1. For each node in the graph constructed by fusing single-cell transcriptome data and spatial transcriptome data, create a node subgraph g c with a diameter equal to the maximum value of the shortest path distances between any two nodes in the graph. In the subgraph g c , randomly select a node v m for masking, and ensure that the structure of the graph remains unchanged during the process. Then, given the subgraph g c as the context, pre-train the model by maximizing the probability of correctly predicting the masked node v m , and the calculation formula for this probability is as follows:
[0078]
[0079] In the above formula, θ represents the set of model parameters, G C is the set of all generated subgraphs, g c represents one of the subgraphs, v m represents the masked node in the subgraph gc, and p(v m |g c ,θ) represents the probability of correctly predicting the masked node v given the subgraph g c and the model parameters θ.m The probability.
[0080] S4.2 generates multiple detailed contexts for each node pair for which link prediction is to be performed. In this stage, the model maximizes the probability of observing a positive edge e cp given the context g p At the same time, the model also learns to assign lower probabilities to the negatively sampled edges e n and their corresponding contexts g cn The overall training objective is constructed by combining the positive edge dataset D p and the negative edge dataset D n By optimizing this objective, the model can improve its ability to accurately predict positive and negative edges. The formula for the training objective is expressed as follows:
[0081]
[0082] In the above formula, L represents the training objective function, e p represents the positive edge, e n represents the negative edge, g cp and g cn represent the contexts related to the positive and negative edges respectively, θ represents the model parameters, P(e p |g cp ,θ) represents the probability of observing the positive edge e cp given the context g p and the model parameters θ, and P(e n |g cn ,θ) represents the probability of observing the negative edge e cn given the context g n and the model parameters θ. When calculating the probability of an edge between two nodes, denoted as e=(vi,vj), a similarity score S(v i ,v j ) is used for the calculation, and its formula is: Here, and are the embedding vectors of nodes v i ,v j respectively, and σ(·) is the sigmoid function. Therefore, the probability of an edge between two nodes is calculated based on the similarity of their embedding vectors. The performance of predicting cell-cell interactions on single-cell spatial transcriptomic data and non-single-cell spatial transcriptomic data is shown in Figure 4 and 5 respectively.
[0083] In summary, the present invention combines cell-specific gene expression data and the spatial location of cells to predict cell-cell communication. The present invention overcomes the limitations of single-cell transcriptome and spatial transcriptome data, solves the problems of accurate identification and analysis of cell-cell communication, helps to more comprehensively reveal the co-regulation between cells, and further deepens the understanding of complex biological processes. By combining spatial omics data with single-cell omics data obtained from the same region through a graph neural network, the disadvantages of limited gene flux and insufficient spatial resolution of spatial omics data are effectively solved, thus providing support for accurately revealing cell-cell communication at the single-cell resolution in space. Using a subgraph-based graph attention neural network to focus on the features of each cell, the mechanism of cell-cell communication existing in normal and diseased tissues is comprehensively revealed at the single-cell resolution, overcoming the problems of insufficient utilization of spatial information, high false positives, and low recognition accuracy in current identification methods.
[0084] The above are only several specific embodiments of the present invention disclosed. However, the embodiments of the present invention are not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. An intelligent recognition method for intercellular communication in spatial transcriptome data, characterized in that, It includes the following steps: Perform quality control and preprocessing on single-cell transcriptome data and spatial transcriptome data; Mine information from the preprocessed single-cell transcriptome data and spatial transcriptome data through a graph attention neural network: use the self-attention mechanism to mine the gene expression information within the single-cell transcriptome data and spatial transcriptome data, then use the cross-attention mechanism to mine the association information between single-cell omics and spatial omics, and finally obtain single-cell gene expression data with spatial location information; Utilize the single-cell gene expression data to construct a subgraph for each cell to reveal the relationship between the corresponding cell and its surrounding cells, extract local features from the subgraph, and then use a subgraph encoder to integrate these local features to further generate global features; Use a graph attention neural network to process the single-cell gene expression features fused with spatial positions, and use the attention mechanism to refine the intercellular relationships; optimize the graph attention neural network in a pre-training and fine-tuning manner: first, pre-train the model on a large-scale dataset containing thousands to tens of thousands of samples; then, fine-tune the model using known signal transduction pathways between specific cells to ensure that the model can accurately identify these biologically important interactions; Among them, mining information from the preprocessed single-cell transcriptome data and spatial transcriptome data through a graph attention neural network: using the self-attention mechanism to mine the gene expression information within the single-cell transcriptome data and spatial transcriptome data, then using the cross-attention mechanism to mine the association information between single-cell omics and spatial omics, and finally obtaining single-cell gene expression data with spatial location information, includes the following steps: 1) Use the graph attention neural network to generate the matching descriptor f i ∈R D ,f i represents the feature descriptor of the i-th cell, which is a vector composed of D real numbers. Here, R represents the set of real numbers, and D is the dimension of the feature vector. Using a point encoder, the transcriptome data of each cell and its spatial position information are combined to specifically represent each cell. A multi-layer perceptron MLP is used to embed the cell position into a high-dimensional vector, which is expressed as: (3) Among them, represents the initial feature vector of the i-th cell, which fuses gene expression data and cell location information; d i is the gene expression data obtained from the transcriptome, p i represents the cell location information; is a multi-layer perceptron encoder used to embed cell location information into a high-dimensional space; 2) Create a graph, integrate the single-cell transcriptome data and spatial transcriptome data, and the nodes of the graph represent cells from transcriptomics; self-connections connect each cell i to all other cells in the same omics, and cross-connections connect cell i to all other cells in different omics; use the information propagation equation to make information propagate along self-connections and cross-connections; 3) Represent the paired scores as a similarity matrix M ∈ R (m×n) , which is used to capture the similarity of the matching descriptors; (4) Here, <·,·> represents the inner product; are the feature vectors of cells in single-cell transcriptomics and spatial transcriptomics, respectively; A and B represent the sets of cells in single-cell transcriptomics and spatial transcriptomics, respectively; m is the number of cells in single-cell transcriptomics, and n is the number of cells in spatial transcriptomics; to obtain the mapping matrix, minimize the following objective function for M: (5) Among them is the cosine similarity function, which is used to measure the similarity between two vectors; * represents matrix slicing; is the objective function, which is used to optimize the similarity matrix M to obtain better cell matching; and represent the number of genes and the number of spatial positions respectively; respectively represent the k-th column of matrices and B respectively; the j-th column of matrices 2. The intercellular communication intelligent recognition method for spatial transcriptome data according to claim 1, characterized in that Perform quality control and preprocessing on single-cell transcriptome data and spatial transcriptome data, including the following steps: 1) First, eliminate cells in the single-cell data that capture fewer than 200 genes; 2) Subsequently, standardize the expression matrix of the single-cell data through formula (1): (1) Among them, C ij represents the original read count of the i-th gene in the j-th cell, that is, the expression level of this gene directly obtained from the sequencing data; D ij represents the expression level of the i-th gene in the j-th cell after normalization, that is, the normalized read count of this gene; is the median of the number of transcripts detected in each cell; M represents the total number of cells, that is, the number of columns in the expression matrix; 3) For the ST dataset with more than 1,000 detected genes, calculate the coefficient of variation CV for each gene using the following equation (2) i :[[]]END]] (2) Among them, σ i represents the standard deviation of the spatial distribution of the i-th gene at all detection points, and u i represents the average expression level of the i-th gene at all detection points; according to the coefficient of variation CV i value, the top 1,000 genes with the highest variability are selected as highly variable genes; subsequently, these genes are compared with the genes detected in the corresponding single-cell gene sequencing data to establish the ground truth for each dataset.
3. The intelligent cell - cell communication recognition method for spatial transcriptome data according to claim 1, wherein, Utilize the single-cell gene expression data to construct a subgraph for each cell to reveal the relationship between the corresponding cell and its surrounding cells, extract local features from the subgraph, and then use a subgraph encoder to integrate these local features to further generate global features, including the following steps: 1) Perform random walks on the created graph. During the walks, mask some nodes and predict these nodes subsequently to capture the overall connection pattern of the graph; for each node and node pair in the graph, generate the corresponding subgraph g c , and these subgraphs g c constitute the set of pre-trained subgraphs ; each subgraph is represented by a set of , where are the nodes in the subgraph , i is the index of the node, i = 1, 2... |V c |, and |V c | represents the number of nodes in g c ; 2) For subgraphs each node in , through the function map its attributes and structure-based embeddings to a stacked vector, and then use a learnable embedding matrix to convert this vector into a low-dimensional representation , the embeddings of all nodes in subgraph g c are collectively represented as , and the initialization of these embeddings is achieved through the output embeddings obtained by a global feature generation method that captures the multi-relational graph structure. Meanwhile, when initializing the node features, the pre-trained node representation vectors are merged with the gene expression data obtained from the cells; 3) For the subgraph , where represents the node set. For each node belonging to , there exists a global input embedding, and these embeddings are represented by the matrix ; furthermore, through context learning, these global embeddings are transformed into new embeddings to reflect the most representative roles of the nodes in ; meanwhile, in order to capture the high-order relational dependencies between nodes, a semantic association matrix is introduced; in the subgraph , and each transformation layer of the global graph G, the weights of the matrix are iteratively learned.
4. The intelligent cell-cell communication recognition method for spatial transcriptome data according to claim 1, wherein Use a graph attention neural network to process single-cell gene expression features that incorporate spatial positions, and utilize the attention mechanism to refine the intercellular relationships; optimize the graph attention neural network by means of pre-training and fine-tuning: First, pre-train the model on a large-scale dataset containing thousands to tens of thousands of samples; then, fine-tune the model using known signal transduction pathways between specific cells to ensure that the model can accurately identify these biologically important interactions, including the following steps: 1) For each node in the graph constructed by integrating single-cell transcriptome data and spatial transcriptome data, create a node subgraph g with a diameter equal to the maximum of the shortest path distances between any two nodes in the graph. c , in the subgraph g c , randomly select a node v m for masking, and ensure that the structure of the graph remains unchanged during the process. Then, with the given subgraph g c as the context, pre-train the model by maximizing the probability of correctly predicting the masked node v m . The formula for this probability is as follows: (6) Among them represents a set of model parameters, G C is a set of all generated subgraphs, g c represents one of the subgraphs, indicates that given the subgraph g c and the model parameters the probability of correctly predicting the masked node v m ; 2) For each node pair for which link prediction is to be performed, multiple contexts are generated. Given the context g cp , the model is trained by maximizing the probability of observing the positive edge e p . At the same time, the model also performs negative sampling to learn to assign a lower probability to the negatively sampled edge e n and its corresponding context g cn . The training objective is constructed by combining the positive edge dataset D p and the negative edge dataset D n . The formula for the training objective is as follows: (7) In formula (7), L represents the training objective function, e p represents the positive margin, and e n represents the negative margin, g cp and g cn successively represent the contexts related to the positive and negative margins, represents the model parameters, denotes the probability of observing the positive margin e cp given the context g and the model parameters p ; denotes the probability of observing the negative margin e cn given the context g and the model parameters n ; When calculating the probability of an edge between two nodes, denoted as e = , a similarity score is used for the calculation, and its formula is: ; here, and are the embedding vectors of nodes respectively, and σ(・) is the sigmoid function.
Citation Information
Patent Citations
Single cell transcriptome sequencing data interpolation method and system based on diffusion-noise reduction
CN114974421A
Spatial transcriptome biological tissue substructure analysis method fused with single cell transcriptome
CN115359845A