Batch correction and spatial domain division method and system for multi-slice spatial transcriptome data
By constructing a spatial adjacency matrix and batch correction, and using an embedded separating neural network to separate batch effects and biological effects, the accuracy and subjective bias problems of multi-slice spatial transcriptome data are solved, and efficient spatial domain partitioning is achieved.
Patent Information
- Application Number
- CN202511032697.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-12-05
AI Technical Summary
Existing technologies suffer from batch effects in the spatial domain partitioning of multi-slice spatial transcriptome data, resulting in insufficient accuracy and requiring significant human and material resources. Furthermore, traditional methods are subject to subjective bias.
By constructing a spatial adjacency matrix data of multiple slices, preliminary batch effect correction is performed. Gene expression data and spatial adjacency matrix information are fused, and batch effect and biological effect are separated by an embedded separation neural network. Finally, clustering is performed to output unified spatial domain partitioning data across slices.
It improves the accuracy of spatial domain division, avoids subjective bias and the need for human and material resources, and provides tools for in-depth analysis.
Smart Images

Figure CN121075412A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of spatial transcriptomics data analysis technology, and in particular to a batch correction and spatial domain partitioning method and system for multi-slice spatial transcriptomics data. Background Technology
[0002] Spatial domains are tissue regions with continuous spatial distribution and similar gene expression patterns. Spatial domain division of complex tissues is a fundamental step in understanding tissue spatial structure and biological function. Traditional methods rely on sophisticated biological instruments, requiring histologists to visually examine tissue sections and label different spatial regions; or to stain the tissue sections with biomarkers to roughly divide different spatial regions. These methods require significant human and material resources, and the division methods are inaccurate, potentially leading to subjective biases. In recent years, advancements in spatial transcriptomics (ST) technology have opened new avenues for a deeper understanding of tissue spatial structure and function. These advanced ST technologies can acquire transcriptomic features (gene expression data) and spatial location features (2D spatial coordinates) for each specific site in complex tissues. Therefore, spatial domain identification algorithms based on graph neural networks (GNNs) can be developed to extract latent features of spatial sites and cluster them to divide spatial regions. Due to current sequencing technology limitations, some large sections are divided into multiple subsections for separate, batch sequencing. Each subsection varies in size and sequencing environment, leading to batch effects in the sequencing results of different sections, affecting the accuracy of spatial domain division of multi-section spatial transcriptomic data.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this application is to provide a batch correction and spatial domain partitioning method and system for multi-slice spatial transcriptome data, aiming to improve the accuracy of spatial domain partitioning for multi-slice spatial transcriptome data.
[0005] To achieve the above objectives, this application proposes a batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data, the method comprising:
[0006] Based on the two-dimensional spatial coordinate data of multiple slices, construct the spatial adjacency matrix data of multiple slices;
[0007] Based on gene expression data and batch coding data from multiple slices, preliminary batch effect correction was performed to obtain preliminary corrected gene expression data.
[0008] The pre-corrected gene expression data is fused with the spatial adjacency matrix data to obtain gene expression data with fused spatial information;
[0009] Based on gene expression data and spatial adjacency matrix data that integrate spatial information, an embedding separation neural network is trained to generate node embedding data, and the node embedding data is separated into batch effect embedding data and biological effect embedding data.
[0010] Clustering is performed on the biological effect embedding data to output data with a unified spatial domain partitioning across slices.
[0011] In one embodiment, the step of constructing a spatial adjacency matrix data of multiple slices based on two-dimensional spatial coordinate data of multiple slices includes:
[0012] Based on the two-dimensional spatial coordinate data of each slice, calculate the Euclidean distance between nodes in each slice;
[0013] For each node, select a preset number of the nearest neighboring nodes from the corresponding multiple Euclidean distance data;
[0014] Based on the selected neighboring node data, construct the spatial adjacency matrix data for each slice; where each element represents the connection relationship between nodes.
[0015] In one embodiment, the step of performing preliminary batch effect correction processing on gene expression data and batch coding data based on multiple slices to obtain preliminary corrected gene expression data includes:
[0016] Stacked gene expression data from multiple slices are generated by stacking gene expression data.
[0017] Stacked gene expression data and batch encoded data are input into a conditional variational autoencoder for processing to reconstruct gene expression data;
[0018] By training the conditional variational autoencoder, the reconstruction is optimized using the mean squared error loss function, and the preliminarily corrected gene expression data is output.
[0019] In one embodiment, the step of inputting stacked gene expression data and batch-encoded data into a conditional variational autoencoder for processing to reconstruct gene expression data includes:
[0020] Stacked gene expression data and batch encoded data are input into a conditional variational autoencoder for processing, so as to generate latent representation data by processing the stacked gene expression data and batch encoded data through the full neural network of the conditional variational autoencoder.
[0021] The latent representation data is stacked with batch-encoded data and then input into the full neural network;
[0022] The reconstructed gene expression data is generated through the full neural network and used as the preliminary corrected gene expression data.
[0023] In one embodiment, the step of fusing the preliminarily corrected gene expression data with spatial adjacency matrix data to obtain gene expression data with fused spatial information includes:
[0024] A spatial map is constructed based on two-dimensional spatial coordinate data from multiple slices;
[0025] Spatial adjacency matrix data and pre-corrected gene expression data are used as inputs to a lightweight graph convolutional network;
[0026] A lightweight graph convolutional network is used to process spatial adjacency matrix data and pre-corrected gene expression data, and integrate spatial graph information to output gene expression data with fused spatial information.
[0027] In one embodiment, the step of processing spatial adjacency matrix data and pre-corrected gene expression data using a lightweight graph convolutional network, and integrating spatial graph information to output gene expression data fused with spatial information specifically includes integration according to the following formula:
[0028]
[0029] in, Gene expression data that incorporates spatial information () represents the graph convolution function. This represents the gene expression data after initial correction. Represents spatial adjacency matrix data; () represents the feature concatenation function. , , The angle matrix.
[0030] In one embodiment, the step of training an embedding separation neural network based on gene expression data with fused spatial information and global spatial adjacency matrix data to generate node embedding data, and separating the node embedding data into batch effect embedding data and biological effect embedding data includes:
[0031] Gene expression data that incorporates spatial information is stacked to generate global gene expression data;
[0032] Stack the spatial adjacency matrix data to generate global spatial adjacency matrix data;
[0033] Global gene expression data and global spatial adjacency matrix data are input into a graph attention neural network to generate node embedding data;
[0034] The node embedding data is separated into batch effect embedding data and biological effect embedding data. A classifier neural network is trained to classify the batch effect embedding data, and a discriminator neural network is trained to perform de-batch optimization on the biological effect embedding data.
[0035] In one embodiment, the method further includes training the discriminator neural network and the classifier neural network using the following loss function:
[0036] ;
[0037] Represents the loss function. Represents the true label vector. Represents the predicted probability vector. Indicates the total number of categories. Indicates the true label in the first place Values in each category This indicates that the model predicts the sample belongs to the first... The probability of each category.
[0038] In one embodiment, the step of clustering the biological effect embedding data to output data with a unified spatial domain partitioning across slices includes:
[0039] Embedded biological effects data as input;
[0040] The input data is processed using the K-means probabilistic clustering algorithm to calculate the similarity between nodes;
[0041] Based on similarity data, nodes are divided into multiple spatial domain groups to generate unified spatial domain partitioning data.
[0042] Furthermore, to achieve the above objectives, this application also proposes a batch correction and spatial domain partitioning system for multi-slice spatial transcriptome data. The system includes: a memory, a processor, and a batch correction and spatial domain partitioning program for multi-slice spatial transcriptome data stored in the memory and executable on the processor. The batch correction and spatial domain partitioning program for multi-slice spatial transcriptome data is configured to implement the steps of the batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data.
[0043] The batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data proposed in this application constructs a multi-slice spatial adjacency matrix based on two-dimensional spatial coordinate data of multiple slices. Based on gene expression data and batch coding data from multiple slices, preliminary batch effect correction is performed to obtain pre-corrected gene expression data. This pre-corrected gene expression data is then fused with the spatial adjacency matrix data to obtain gene expression data with fused spatial information. Next, based on the fused spatial information gene expression data and the spatial adjacency matrix data, an embedding separation neural network is trained to generate node embedding data. This node embedding data is then separated into batch effect embedding data and biological effect embedding data. Finally, the biological effect embedding data is clustered to output unified spatial domain partitioning data across slices. Thus, by constructing a spatial adjacency matrix data of multiple slices and preliminarily corrected gene expression data, spatial information and batch correction information are integrated, providing a reliable foundation for subsequent embedding separation and clustering processing. Furthermore, the embedding separation neural network is used to separate the node embedding data into batch effect embedding data and biological effect embedding data, achieving effective differentiation between batch effects and biological effects. Finally, by performing clustering processing on the biological effect embedding data, unified spatial domain partitioning data across slices is output. This not only improves the accuracy of spatial domain partitioning but also avoids the subjective bias and large investment of human and material resources required by traditional methods, providing a powerful tool for in-depth analysis of spatial transcriptomics data. Attached Figure Description
[0044] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0045] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a flowchart illustrating an embodiment of the batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data in this application.
[0047] Figure 2 For this application Figure 1 A detailed flowchart of step S100;
[0048] Figure 3 For this application Figure 1 A detailed flowchart of step S200;
[0049] Figure 4 For this application Figure 3 A detailed flowchart of step S220;
[0050] Figure 5 For this application Figure 1 Detailed flowchart of step S300;
[0051] Figure 6 For this application Figure 1 Detailed flowchart of step S400;
[0052] Figure 7 For this application Figure 1 Detailed flowchart of step S500;
[0053] Figure 8 This is a schematic diagram of a system for batch correction and spatial domain partitioning of multi-slice spatial transcriptome data according to an embodiment of this application.
[0054] Explanation of icon numbers:
[0055] 10. Memory; 20. Processor.
[0056] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0057] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0058] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0059] The main solution of this application embodiment is as follows: Based on the two-dimensional spatial coordinate data of multiple slices, construct the spatial adjacency matrix data of multiple slices, and perform preliminary batch effect correction processing based on the gene expression data and batch coding data of multiple slices to obtain the preliminary corrected gene expression data. Then, fuse the preliminary corrected gene expression data with the spatial adjacency matrix data to obtain gene expression data with fused spatial information. Then, based on the gene expression data with fused spatial information and the spatial adjacency matrix data, train the embedding separation neural network to generate node embedding data, and separate the node embedding data into batch effect embedding data and biological effect embedding data. Finally, perform clustering processing on the biological effect embedding data to output unified spatial domain partitioning data across slices.
[0060] In this embodiment, for ease of description, the following description focuses on identifying a batch correction and spatial domain partitioning system for multi-slice spatial transcriptome data.
[0061] Because existing technologies divide larger slices into multiple sub-slices for separate and batch sequencing, and each sub-slice has a different size and sequencing environment, the sequencing results of different slices have batch effects, which affect the accuracy of spatial domain division of multi-slice spatial transcriptome data.
[0062] The solution provided in this application integrates spatial information and batch correction information by constructing a spatial adjacency matrix data of multiple slices and preliminary corrected gene expression data, providing a reliable foundation for subsequent embedding separation and clustering processing. Furthermore, it utilizes an embedding separation neural network to separate node embedding data into batch effect embedding data and biological effect embedding data, achieving effective differentiation between batch effects and biological effects. Finally, by clustering the biological effect embedding data, it outputs unified spatial domain partitioning data across slices. This not only improves the accuracy of spatial domain partitioning but also avoids the subjective bias and significant human and material resource investment required by traditional methods, providing a powerful tool for in-depth analysis of spatial transcriptomics data.
[0063] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, a batch correction and spatial domain partitioning system for multi-slice spatial transcriptome data, etc. The following description uses a batch correction and spatial domain partitioning system for multi-slice spatial transcriptome data as an example to illustrate this embodiment and the subsequent embodiments.
[0064] Based on this, embodiments of this application provide a batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data, referring to... Figure 1 In this embodiment, the batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data includes steps S100-S500, wherein:
[0065] Step S100: Based on the two-dimensional spatial coordinate data of multiple slices, construct the spatial adjacency matrix data of multiple slices.
[0066] In this embodiment, the two-dimensional spatial coordinate data of multiple slices are obtained through spatial location barcodes or multi-round high-precision microscopy imaging techniques used in spatial transcriptome sequencing. To construct the spatial adjacency matrix data, the system first calculates the relative positional relationships between nodes (e.g., cells or feature points) within each slice, and then determines which nodes are connected based on preset spatial connection rules or thresholds. In the spatial adjacency matrix data, each element represents whether a connection exists between a pair of nodes, providing a foundation for subsequent spatial information fusion. The construction of the spatial adjacency matrix helps to consider the spatial proximity between cells in subsequent analysis, thereby more accurately reflecting the real situation within the organism.
[0067] In one feasible implementation, refer to Figure 2 Step S100 includes steps S110 to S130, wherein:
[0068] Step S110: Based on the two-dimensional spatial coordinate data of each slice, calculate the Euclidean distance data between nodes in each slice.
[0069] In this embodiment, the Euclidean distance is calculated as the square root of the sum of the squares of the coordinate differences between two points, reflecting the straight-line distance between two points in space. By calculating the Euclidean distance of each node to other nodes, the spatial structure within the slice can be quantified, providing data support for the subsequent construction of the spatial adjacency matrix.
[0070] Step S120: For each node, select a preset number of nearest neighboring node data from the corresponding multiple Euclidean distance data.
[0071] In this embodiment, to construct more accurate spatial connections, the system selects the 8 nearest neighbors from multiple Euclidean distance data for each node, based on a preset number of adjacent nodes (e.g., 8 nearest neighbors). This selection considers both the spatial proximity between nodes and limits the complexity of the connections, helping to maintain computational efficiency and accuracy in subsequent analysis.
[0072] Step S130: Based on the selected neighboring node data, construct the spatial adjacency matrix data for each slice; where each element represents the connection relationship between nodes.
[0073] In this embodiment, the spatial adjacency matrix is constructed by converting the selected neighboring node data into a matrix form. Each element in the matrix represents whether a connection exists between a pair of nodes, using binary representation: 1 indicates a connection, and 0 indicates no connection. This not only intuitively displays the spatial structure within the slice but also facilitates subsequent spatial information fusion. Through the spatial adjacency matrix, the system can effectively capture the spatial proximity between cells. In subsequent analysis steps, the spatial adjacency matrix will be combined with the pre-corrected gene expression data and input together into a lightweight graph convolutional network to achieve deep fusion of spatial information and gene expression data.
[0074] Step S200: Based on gene expression data and batch coding data from multiple slices, perform preliminary batch effect correction to obtain preliminary corrected gene expression data.
[0075] In this embodiment, each slice is batch-encoded using a one-hot encoding method according to its batch number, and the gene expression data of the multi-slice ST data and the batch codes of each slice are stacked and merged. Then, a conditional variational autoencoder is used to perform preliminary batch effect correction on the merged data. This process aims to reduce or eliminate data bias introduced by different batches, thereby obtaining preliminary corrected gene expression data.
[0076] In one feasible implementation, refer to Figure 3 Step S200 includes steps S210 to S230, wherein:
[0077] Step S210: Stack the gene expression data of multiple slices to generate stacked gene expression data.
[0078] In this embodiment, the construction of stacked gene expression data involves integrating gene expression data from different slices to form a comprehensive dataset containing gene expression information from all slices. This allows for the simultaneous consideration of gene expression from multiple slices in subsequent analyses, thereby improving the comprehensiveness and accuracy of the analysis.
[0079] Step S220: Input the stacked gene expression data and batch encoding data into the conditional variational autoencoder for processing to reconstruct the gene expression data.
[0080] In this embodiment, the Conditional Variational Autoencoder (CVA) can reconstruct input data by learning the latent representation of the data. In batch effect correction applications, the CVA can learn data features unrelated to batch effects, i.e., biological effect features, thereby generating preliminarily corrected gene expression data. This process aims to reduce data bias between batches, making data from different batches more comparable after correction. Through joint processing of stacked gene expression data and batch-encoded data, the CVA can capture the latent relationship between batch encoding and gene expression data, thus achieving batch effect correction. The corrected gene expression data will more accurately reflect the true gene expression situation in the organism, providing a reliable foundation for subsequent spatial information fusion and embedding separation.
[0081] In one feasible implementation, refer to Figure 4 Step S220 includes steps S221 to S222, wherein:
[0082] Step S221: The stacked gene expression data and batch coding data are input into the conditional variational autoencoder for processing, so as to generate latent representation data by processing the stacked gene expression data and batch coding data through the full neural network of the conditional variational autoencoder.
[0083] In this embodiment, the fully neural network structure of the conditional variational autoencoder can learn the complex distribution of data and extract biological effect features that are independent of batch effects. Through latent representation data, preliminary corrected gene expression data can be reconstructed, which, while removing batch effects, retains the true gene expression information within the organism.
[0084] Step S222: Stack the latent representation data and batch encoded data, and input them into the full neural network to generate reconstructed gene expression data through the full neural network.
[0085] In this embodiment, the latent representation data and batch-encoded data are stacked again and input into the full neural network to further utilize the batch-encoded information to refine the correction process of gene expression data. Through the complex learning capabilities of the full neural network, the system can generate more accurate reconstructed gene expression data based on latent representation and batch encoding.
[0086] Step S230: By training the conditional variational autoencoder, the reconstruction is optimized using the mean squared error loss function, and the preliminarily corrected gene expression data is output.
[0087] In this embodiment, by training a conditional variational autoencoder and optimizing the reconstruction process using a mean squared error loss function, it is possible to ensure that the output pre-corrected gene expression data retains biological effect characteristics as much as possible while effectively reducing batch-to-batch data bias. The mean squared error loss function can quantify the difference between the reconstructed gene expression data and the original data. By adjusting the parameters of the conditional variational autoencoder through the backpropagation algorithm, the reconstruction error is minimized, thereby obtaining high-quality pre-corrected gene expression data.
[0088] Step S300: The preliminarily corrected gene expression data is fused with the spatial adjacency matrix data to obtain gene expression data with fused spatial information.
[0089] In this embodiment, to combine spatial information with gene expression data, the system first aligns the pre-corrected gene expression data with the spatial adjacency matrix data, ensuring that the gene expression data of each node corresponds to its position in the spatial adjacency matrix. Then, using spatial information processing techniques such as graph convolutional networks, the spatial adjacency matrix is used as the edge information of the graph, and the pre-corrected gene expression data is used as the node information of the graph, for information fusion and propagation. This process aims to capture the influence of spatial structure on gene expression, generating fused data that includes both gene expression information and spatial location information, providing input for subsequent biological effect embedding data generation. By fusing spatial information, the system can more comprehensively understand the spatial distribution characteristics of gene expression, providing an information foundation for in-depth analysis of spatial transcriptomics data.
[0090] In one feasible implementation, refer to Figure 5 Step S300 includes steps S310 to S330, wherein:
[0091] Step S310: Construct a spatial map based on the two-dimensional spatial coordinate data of multiple slices.
[0092] In this embodiment, two-dimensional spatial coordinate data can be transformed into a spatial graph using techniques such as graph convolutional networks. Nodes in the spatial graph represent cells or feature points in a slice, while edges represent the connections between nodes based on spatial adjacency matrix data. This transformation allows spatial information to be processed by a computer in the form of a graph structure, providing a foundation for the subsequent fusion of spatial information with gene expression data.
[0093] Step S320: Use the spatial adjacency matrix data and the preliminarily corrected gene expression data as input to the lightweight graph convolutional network.
[0094] In this embodiment, the Lightweight Graph Convolutional Network (LTCNN) is a neural network model specifically designed for processing graph-structured data. It can achieve deep data fusion and propagation by learning the spatial relationships and feature information between nodes. In this embodiment, the LTCNN takes spatial adjacency matrix data and pre-corrected gene expression data as input. Through its internal learning mechanism, it effectively fuses spatial information with gene expression data. The fused data not only retains the feature information of gene expression but also incorporates the spatial proximity between cells, providing a more comprehensive and accurate information foundation for subsequent biological effect embedding data generation.
[0095] Step S330: The spatial adjacency matrix data and the preliminarily corrected gene expression data are processed by a lightweight graph convolutional network, and the spatial graph information is integrated to output gene expression data with fused spatial information.
[0096] In this embodiment, the lightweight graph convolutional network utilizes graph convolution operations to capture spatial relationships and feature information between nodes when processing spatial adjacency matrix data and pre-corrected gene expression data. Through the stacking of multiple graph convolutional layers, the network can progressively extract deep-level fusion features, which include both gene expression information and spatial proximity between cells. Ultimately, the network integrates spatial graph information, outputting fused data that includes both gene expression and spatial location information. This data provides a richer and more accurate information foundation for subsequent biological effect embedding data generation, helping to further improve the accuracy of spatial domain partitioning.
[0097] In one feasible implementation, step S330 includes integrating spatial map information according to the following formula to output gene expression data fused with spatial information:
[0098]
[0099] in, Gene expression data that incorporates spatial information () represents the graph convolution function. This represents the gene expression data after initial correction. Represents spatial adjacency matrix data; () represents the feature concatenation function. , , The angle matrix.
[0100] In this embodiment,
[0101] In the above formula, the graph convolution function () By stacking multiple graph convolutional layers, deep-level fusion features are gradually extracted. Feature concatenation function. () is used to integrate spatial map information with gene expression data, thereby outputting fused data that contains both gene expression information and spatial location information. Such fused data provides a more comprehensive and accurate information foundation for the subsequent generation of biological effect embedding data.
[0102] Step S400: Based on gene expression data and spatial adjacency matrix data with fused spatial information, train an embedding separation neural network to generate node embedding data, and separate the node embedding data into batch effect embedding data and biological effect embedding data.
[0103] In this embodiment, to extract useful biological information from gene expression data incorporating spatial information while removing batch effects, the system employs an embedding-separating neural network. This neural network learns low-dimensional embedding representations of the data, which include both biological effect features and batch effect influences. Through a specific training strategy, the system can separate node embedding data into batch effect embedding data and biological effect embedding data. Batch effect embedding data reflects data bias introduced by different batches, while biological effect embedding data retains the true gene expression characteristics and spatial structure information within the organism. This separation helps to focus on biological effects in subsequent analyses, avoiding interference from batch effects. When training the embedding-separating neural network, the system uses spatial adjacency matrix data as edge information of the graph and gene expression data incorporating spatial information as node information of the graph. Through the complex learning capabilities of the neural network, it gradually extracts deep-level feature representations and achieves the separation of batch effects and biological effects. Finally, the system outputs high-quality biological effect embedding data, providing a reliable information foundation for subsequent spatial domain partitioning steps.
[0104] In one feasible implementation, refer to Figure 6 Step S400 includes steps S410 to S420, wherein:
[0105] Step S410: Stack the gene expression data with fused spatial information to generate global gene expression data.
[0106] In this embodiment, the purpose of stacking global gene expression data is to integrate gene expression data from different slices, after preliminary correction and fusion of spatial information, to form a comprehensive global dataset. This dataset can more comprehensively reflect the gene expression status of the entire sample, providing richer and more accurate information for subsequent analysis. By stacking, the system can integrate gene expression data from multiple slices, ensuring that data from all slices can be considered simultaneously during analysis, thereby improving the accuracy and comprehensiveness of the analysis.
[0107] Step S420: Stack the spatial adjacency matrix data to generate global spatial adjacency matrix data.
[0108] In this embodiment, the purpose of stacking global spatial adjacency matrix data is to integrate spatial adjacency relationships from different slices to form a global spatial structure description. Through stacking, the system can integrate spatial adjacency matrix data from multiple slices, ensuring that spatial information from all slices is considered simultaneously during analysis, thereby more accurately capturing the spatial proximity and spatial distribution characteristics between cells. The global spatial adjacency matrix data and global gene expression data will jointly serve as input to the embedding segregating neural network, providing a comprehensive information foundation for the subsequent generation of biological effect embedding data.
[0109] Step S430: Input the global gene expression data and the global spatial adjacency matrix data into the graph attention neural network to generate node embedding data.
[0110] In this embodiment, the graph attention neural network is a neural network model specifically designed for processing graph-structured data. It can generate low-dimensional embedding representations of nodes by learning the relationships and feature information between nodes. In this embodiment, the graph attention neural network takes global gene expression data and global spatial adjacency matrix data as input. Through its internal learning mechanism, it captures the complex relationship between gene expression and spatial structure, and generates node embedding data that contains both gene expression information and spatial location information.
[0111] Understandably, this application employs a dual spatial information integration mechanism through a graph attention neural network and a lightweight graph convolutional network. First, the lightweight graph convolutional network achieves initial fusion of spatial information and gene expression data when processing the pre-corrected gene expression data and spatial adjacency matrix data. Subsequently, the graph attention neural network further utilizes global gene expression data and global spatial adjacency matrix data, leveraging its powerful learning capabilities to deeply explore the potential connections between gene expression and spatial structure. This dual spatial information integration mechanism can capture spatial features at different scales, improve the model's ability to distinguish complex organizational structures, reduce information loss, and contribute to improving the clarity of spatial domain boundary delineation. Furthermore, by dynamically learning the spatial association weights between nodes through the multi-head attention mechanism in the graph attention neural network, it avoids the over-smoothing caused by ordinary graph convolutional models, enhancing the model's robustness.
[0112] Step S440: Separate the node embedding data into batch effect embedding data and biological effect embedding data, and train a classifier neural network to classify the batch effect embedding data and train a discriminator neural network to perform de-batch optimization on the biological effect embedding data.
[0113] In this embodiment, to separate batch effects and biological effects from node embedding data, the system employs a specific decomposition strategy. This strategy leverages the learning capabilities of neural networks to identify and separate feature information related to batch effects and biological effects. By training a classifier neural network, the system can effectively classify batch effect embedding data, thereby further understanding data biases between different batches. Simultaneously, a discriminator neural network is trained to de-batch optimize the biological effect embedding data, aiming to reduce the impact of batch effects on biological effect embedding data and improve data comparability and accuracy.
[0114] In one feasible implementation, this application employs the cross-entropy loss function to train the discriminator neural network and the classifier neural network:
[0115] ;
[0116] Represents the loss function. Represents the true label vector. Represents the predicted probability vector. Indicates the total number of categories. Indicates the true label in the first place Values in each category This indicates that the model predicts the sample belongs to the first... The probability of each category.
[0117] In this embodiment, the cross-entropy loss function is used to quantify the difference between the predictions of the discriminator neural network and the classifier neural network and the true labels. By minimizing the loss function, the parameters of the neural network can be adjusted to make the predictions closer to the true labels. During training, the system iteratively updates the weights of the neural network until the loss function converges to a preset value. In this way, the system can generate high-quality biological effect embedding data. This data, while removing batch effects, retains the true gene expression characteristics and spatial structure information of the organism, providing a reliable information foundation for subsequent spatial domain partitioning steps.
[0118] It is understandable that steps S400 and S200 form a dual batch effect removal mechanism. Step S200 uses a conditional variational autoencoder to initially correct gene expression data, removing direct biases between batches; while step S400 further utilizes an embedding separation neural network to separate batch effect embeddings and biological effect embeddings from gene expression data that incorporates spatial information, achieving a deeper level of batch effect removal. This dual batch effect removal strategy helps ensure that the final biological effect embeddings accurately reflect the true gene expression characteristics and spatial structure information within the organism while minimizing batch interference. Furthermore, the interpretability separation mechanism for biological effect embeddings and batch effect embeddings in this application preserves the original node embeddings while achieving information separation: separating batch effects from the original node embeddings and extracting biological effects for spatial domain partitioning, applicable to multi-slice noise application scenarios caused by different sequencing technologies and conditions.
[0119] Step S500: Cluster the biological effect embedding data to output unified spatial domain partitioning data across slices.
[0120] In this embodiment, to transform the biological effect embedding data into biologically meaningful spatial domain partitioning results, the system performs clustering processing on this data, dividing the samples in the dataset into several groups or clusters. This results in high similarity among samples within the same cluster and low similarity between samples from different clusters. In this embodiment, the system employs appropriate clustering algorithms, such as K-means, Louvain, or mclust clustering, to process the biological effect embedding data. The choice of clustering algorithm depends on the characteristics of the data and the needs of the analysis. Through clustering, the system can group cells or feature points with similar gene expression characteristics and spatial location information into the same spatial domain, thereby achieving a unified spatial domain partitioning of the entire sample. The output unified spatial domain partitioning data across slices provides an important foundation for subsequent biological effect analysis and spatial transcriptomics research, helping to reveal the complex relationship between gene expression and spatial structure, and the role of these relationships in organismal function and disease development.
[0121] In one feasible implementation, step S500 includes steps S510 to S530, wherein:
[0122] Step S510: Use the biological effect embedding data as input.
[0123] In this embodiment, the biological effect embedding data is used as the input to the clustering algorithm in order to utilize the gene expression characteristics and spatial location information contained in these data to perform accurate spatial domain division.
[0124] Step S520: The input data is processed using the K-means probabilistic clustering algorithm to calculate the similarity between nodes.
[0125] In this embodiment, the K-means probabilistic clustering algorithm can iteratively adjust the position of cluster centers to maximize the similarity of samples within the same cluster and minimize the similarity of samples between different clusters. In this embodiment, the K-means probabilistic clustering algorithm calculates the similarity between each node and other nodes; this similarity data reflects the proximity of gene expression characteristics and spatial locations between nodes.
[0126] Step S530: Based on similarity data, the nodes are divided into multiple spatial domain groups to generate unified spatial domain partitioning data.
[0127] In this embodiment, based on similarity data, the algorithm divides nodes into multiple spatial domain groups, with nodes within each group sharing similar gene expression characteristics and spatial location information. Thus, through K-means probabilistic clustering, the system can generate unified spatial domain partitioning data across slices. This data provides crucial information for subsequent biological effect analysis and spatial transcriptomics research. These spatial domain partitions help reveal the complex relationship between gene expression and spatial structure, and the potential role of these relationships in organismal function and disease development.
[0128] The batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data proposed in this application constructs a multi-slice spatial adjacency matrix based on two-dimensional spatial coordinate data of multiple slices. Based on gene expression data and batch coding data from multiple slices, preliminary batch effect correction is performed to obtain pre-corrected gene expression data. This pre-corrected gene expression data is then fused with the spatial adjacency matrix data to obtain gene expression data with fused spatial information. Next, based on the fused spatial information gene expression data and the spatial adjacency matrix data, an embedding separation neural network is trained to generate node embedding data. This node embedding data is then separated into batch effect embedding data and biological effect embedding data. Finally, the biological effect embedding data is clustered to output unified spatial domain partitioning data across slices. Thus, by constructing a spatial adjacency matrix data of multiple slices and preliminarily corrected gene expression data, spatial information and batch correction information are integrated, providing a reliable foundation for subsequent embedding separation and clustering processing. Furthermore, the embedding separation neural network is used to separate the node embedding data into batch effect embedding data and biological effect embedding data, achieving effective differentiation between batch effects and biological effects. Finally, by performing clustering processing on the biological effect embedding data, unified spatial domain partitioning data across slices is output. This not only improves the accuracy of spatial domain partitioning but also avoids the subjective bias and large investment of human and material resources required by traditional methods, providing a powerful tool for in-depth analysis of spatial transcriptomics data.
[0129] This application also provides a batch correction and spatial domain partitioning system for multi-slice spatial transcriptome data, referencing... Figure 8 The system includes: a memory 10, a processor 20, and a batch correction and spatial domain partitioning program for multi-slice spatial transcriptome data stored in the memory 10 and executable on the processor 20. The batch correction and spatial domain partitioning program for multi-slice spatial transcriptome data is configured to implement the steps of the batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data.
[0130] The batch correction and spatial domain partitioning system for multi-slice spatial transcriptome data provided in this application, employing the batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data in the above embodiments, can improve the accuracy of spatial domain partitioning for multi-slice spatial transcriptome data. Compared with the prior art, the beneficial effects of the batch correction and spatial domain partitioning system for multi-slice spatial transcriptome data provided in this application are the same as those of the batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data provided in the above embodiments, and other technical features of the batch correction and spatial domain partitioning system for multi-slice spatial transcriptome data are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0131] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data, characterized in that, The method includes: Based on the two-dimensional spatial coordinate data of multiple slices, construct the spatial adjacency matrix data of multiple slices; Based on gene expression data and batch coding data from multiple slices, preliminary batch effect correction was performed to obtain preliminary corrected gene expression data. The pre-corrected gene expression data is fused with the spatial adjacency matrix data to obtain gene expression data with fused spatial information; Based on gene expression data and spatial adjacency matrix data that integrate spatial information, an embedding separation neural network is trained to generate node embedding data, and the node embedding data is separated into batch effect embedding data and biological effect embedding data. Clustering is performed on the biological effect embedding data to output data with a unified spatial domain partitioning across slices.
2. The batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data as described in claim 1, characterized in that, The steps for constructing a spatial adjacency matrix based on two-dimensional spatial coordinate data of multiple slices include: Based on the two-dimensional spatial coordinate data of each slice, calculate the Euclidean distance between nodes in each slice; For each node, select a preset number of the nearest neighboring nodes from the corresponding multiple Euclidean distance data; Based on the selected neighboring node data, construct the spatial adjacency matrix data for each slice; where each element represents the connection relationship between nodes.
3. The batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data as described in claim 1, characterized in that, The step of performing preliminary batch effect correction on gene expression data and batch coding data based on multiple slices to obtain preliminary corrected gene expression data includes: Stacked gene expression data from multiple slices are generated by stacking gene expression data. Stacked gene expression data and batch encoded data are input into a conditional variational autoencoder for processing to reconstruct gene expression data; By training the conditional variational autoencoder, the reconstruction is optimized using the mean squared error loss function, and the preliminarily corrected gene expression data is output.
4. The batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data as described in claim 3, characterized in that, The step of inputting stacked gene expression data and batch encoded data into a conditional variational autoencoder for processing to reconstruct gene expression data includes: Stacked gene expression data and batch encoded data are input into a conditional variational autoencoder for processing, so as to generate latent representation data by processing the stacked gene expression data and batch encoded data through the full neural network of the conditional variational autoencoder. Potential representation data and batch-encoded data are stacked and input into the full neural network to generate reconstructed gene expression data.
5. The batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data as described in claim 1, characterized in that, The step of fusing the pre-corrected gene expression data with the spatial adjacency matrix data to obtain gene expression data with fused spatial information includes: A spatial map is constructed based on two-dimensional spatial coordinate data from multiple slices; Spatial adjacency matrix data and pre-corrected gene expression data are used as inputs to a lightweight graph convolutional network; A lightweight graph convolutional network is used to process spatial adjacency matrix data and pre-corrected gene expression data, and integrate spatial graph information to output gene expression data with fused spatial information.
6. The batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data as described in claim 5, characterized in that, The steps of processing spatial adjacency matrix data and pre-corrected gene expression data using a lightweight graph convolutional network, and integrating spatial graph information to output gene expression data with fused spatial information, specifically include integration according to the following formula: ; in, Gene expression data that incorporates spatial information () represents the graph convolution function. This represents the gene expression data after initial correction. Represents spatial adjacency matrix data; () represents the feature concatenation function. , , The angle matrix.
7. The batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data as described in claim 1, characterized in that, The steps of training an embedding separation neural network based on gene expression data with fused spatial information and global spatial adjacency matrix data to generate node embedding data, and separating the node embedding data into batch effect embedding data and biological effect embedding data include: Gene expression data that incorporates spatial information is stacked to generate global gene expression data; Stack the spatial adjacency matrix data to generate global spatial adjacency matrix data; Global gene expression data and global spatial adjacency matrix data are input into a graph attention neural network to generate node embedding data; The node embedding data is separated into batch effect embedding data and biological effect embedding data. A classifier neural network is trained to classify the batch effect embedding data, and a discriminator neural network is trained to perform de-batch optimization on the biological effect embedding data.
8. The batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data as described in claim 7, characterized in that, The method also includes training the discriminator neural network and the classifier neural network using the following loss function: ; Represents the loss function. Represents the true label vector. Represents the predicted probability vector. Indicates the total number of categories. Indicates the true label in the first place Values in each category This indicates that the model predicts the sample belongs to the first... The probability of each category.
9. The batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data as described in claim 1, characterized in that, The step of clustering the biological effect embedding data to output a unified spatial domain partitioning data across slices includes: Embedded biological effects data as input; The input data is processed using the K-means probabilistic clustering algorithm to calculate the similarity between nodes; Based on similarity data, nodes are divided into multiple spatial domain groups to generate unified spatial domain partitioning data.
10. A batch correction and spatial domain partitioning system for multi-slice spatial transcriptome data, characterized in that, The system includes: a memory, a processor, and a batch correction and spatial domain partitioning program for multi-slice spatial transcriptome data stored in the memory and executable on the processor, the batch correction and spatial domain partitioning program for multi-slice spatial transcriptome data being configured to implement the steps of the batch correction and spatial domain partitioning method for multi-slice spatial transcriptome data as described in any one of claims 1 to 9.