Method for generating spatial transcriptomics data from histological image based on deep learning

By combining deep learning methods with convolutional neural networks and Transformer architecture, the accuracy and computational complexity problems in spatial transcriptomics data prediction were solved, efficient gene expression prediction and spatial region identification were achieved, and the application of precision medicine and tumor microenvironment research were promoted.

CN120673839APending Publication Date: 2025-09-19NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510770541.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies for spatial transcriptomics data prediction have problems such as insufficient prediction accuracy, excessive number of model parameters, and high computational complexity, and lack joint modeling of the cross-scale association between tissue macrostructure and cellular microcomponents.

Method used

A deep learning-based approach is adopted, combining convolutional neural networks, Transformer architecture, and graph neural networks. The model is optimized through the leave-one-out-cross validation method to build a global association model across image blocks. The Mamba module, weight parameter sharing, and self-attention mechanism are used to reduce the number of computational parameters and improve feature extraction capabilities.

Benefits of technology

It improves the accuracy of gene expression prediction and spatial region identification, reduces computational complexity and the number of model parameters, supports the adaptation of multiple technology platforms, promotes the popularization of precision medicine, and provides effective tumor microenvironment analysis and targeted treatment guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673839A_ABST
    Figure CN120673839A_ABST
Patent Text Reader

Abstract

The invention provides a histological image driving space transcriptome data generation method based on deep learning, and the method comprises the steps: constructing a training / testing set through a standardized preprocessing process, and screening a target data set; then, segmenting the full-view tissue slice into image blocks taking a space transcriptome sequencing point as a center, and inputting the image blocks into a convolutional neural network to extract local potential features; further, through a collaborative architecture of a state space model network VMama and Transform, a global association model of cross-image blocks is established; finally, a graph neural network is adopted to fuse a self-attention mechanism, and dynamic representation of the spatial locus and gene expression relation is achieved. The invention provides powerful guarantee for disclosing the molecular characteristics of the tissue, and specifically comprises the following steps: technically optimizing a network architecture and a standardized data preprocessing process; in application, accurate gene expression and spatial region prediction are carried out; the stability and robustness of the model are ensured through multi-stage cross validation in experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of multimodal data fusion analysis in bioinformatics, and specifically relates to a computational method that integrates the Transformer architecture with graph neural networks to achieve gene expression prediction of spatial transcriptome data and analysis of tissue spatial molecular features through tissue pathology images. Background Art

[0002] Bioinformatics, an interdisciplinary field at the intersection of life sciences and information sciences, is dedicated to analyzing genetic regulatory mechanisms and tumor microenvironmental characteristics through advanced computational models. It holds significant scientific value in revealing the laws governing life processes and the mechanisms of disease. Recent breakthroughs in spatial transcriptomics (ST) technology have enabled the simultaneous acquisition of RNA expression profiles and histological images at single-cell spatial resolution, providing a new research paradigm for analyzing cellular interaction networks and tissue microenvironments. However, ST technology faces bottlenecks such as high data acquisition costs and limited throughput, severely hindering its clinical translation and application.

[0003] To address these challenges, the academic community has conducted extensive research, proposing a new paradigm for predicting spatial gene expression based on whole-slide stained histological images (WSIs). Literature research indicates that several representative achievements have been made in this field. Schmauch et al. proposed HE2RNA (Schmauch B, Romagnoni A, Pronier E, Saillard C, Maillé P, Calderaro J, Kamoun A, Sefta M, Toldo S, Zaslavskiy M: A deep learning model to predict RNA-seq expression of tumors from whole slide images. Nat. Commun. 11, 3877. In.; 2020). Through deep feature extraction, they successfully established a mapping relationship between WSI and RNA-seq data, revealing characteristic patterns of tumor heterogeneity. ST-Net (He B, Bergenstråhle L, Stenbeck L, Abid A, AnderssonA, Borg Å, Maaskola J, Lundeberg J, Zou J: Integrating spatial geneexpression and breast tumour morphology via deep learning. Nature biomedicalengineering 2020, 4(8):827-834) pioneered a convolutional neural network architecture based on an image partitioning strategy, which achieved effective prediction of spatially variable gene expression.HistToGene (Pang M, Su K, Li M: Leveraging information in spatial transcriptomics to predict super-resolution gene expression from histology images in tumors. BioRxiv 2021:2021.2011.2028.470212) and Hist2ST (Pang M, Su K, Li M: Leveraging information in spatial transcriptomics to predict super-resolution gene expression from histology images in tumors. BioRxiv 2021:2021.2011. 2028.470212) further improved the prediction accuracy through the Transformer self-attention mechanism and graph neural network hierarchical feature fusion, respectively.

[0004] Despite significant progress in existing methods, three key scientific challenges remain. First, existing models often focus on extracting features at a single scale and lack the ability to jointly model the cross-scale associations between tissue macrostructure and cellular microstructures. Second, the complex nonlinear relationships between cell morphological parameters (such as nuclear-cytoplasmic ratio and density distribution) and gene expression have not yet been effectively encoded. Furthermore, the complexity of model architectures leads to an exponential increase in the number of parameters, resulting in severe computational redundancy and limiting the feasibility of clinical deployment. Summary of the Invention

[0005] The present invention aims to provide a spatial transcriptome prediction method based on graph neural network and feature embedding, which effectively solves the bottleneck problems of insufficient prediction accuracy, excessive number of model parameters and high computational complexity in the existing technology.

[0006] The technical solution to achieve the objectives of the present invention is: a method for generating spatial transcriptomics data from histological images based on deep learning, comprising the following steps: Step 1: Preprocess the spatial transcriptome dataset generated by the 10X Genomics platform (containing histological images, spatial gene expression data, and point coordinates). The top 1000 highly variable genes in each tissue section were selected, and genes with fewer than 1000 expression spots across all tissue sections were excluded. The dataset was then normalized and natural logarithm-transformed. Histological images were divided into N blocks of size (3 × W × H), where W = H = 112 pixels. The dataset was then divided into training and test sets.

[0007] Step 2: Build a model based on a deep learning algorithm and train it on the training set. During the training process, the optimal hyperparameters are determined using the leave-one-out cross validation method. Building a model based on a deep learning algorithm involves segmenting full-field tissue sections into image blocks centered on spatial transcriptome sequencing points, which are then fed into a convolutional neural network to extract local latent features. Furthermore, a global correlation model across image blocks is established through the collaborative architecture of the state-space model network VMamba and the Transformer. Finally, a graph neural network is integrated with a self-attention mechanism to dynamically represent the relationship between spatial sites and gene expression. Step 3: Input the test data set into the model with the updated optimal parameters to obtain a set of gene prediction results; Step 4: Select one sub-dataset from each of the disjoint sub-datasets as the validation dataset, and the remaining sub-datasets as the training dataset. Repeat steps 3 and 4 until all sub-datasets have been trained as both validation and training datasets, obtaining gene prediction results for all groups. Collect all prediction results to evaluate model performance.

[0008] Compared with the existing technology, the present invention has the following significant advantages: (1) it reduces the number of computational parameters, making the improved algorithm have obvious advantages in efficiency, deployment and generalization; (2) it improves the accuracy of gene expression prediction and spatial region identification.

[0009] This invention provides a powerful guarantee for revealing tissue molecular characteristics: technically through optimized network architecture and standardized data preprocessing; applied through precise gene expression and spatial region prediction; and experimentally through multi-stage cross-validation to ensure model stability and robustness. This multi-dimensional safeguard ensures that the model operates efficiently while also demonstrating significant advantages in actual clinical and scientific research. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 Schematic diagram of the method of the present invention. DETAILED DESCRIPTION

[0011] The present invention will be further described below with reference to the accompanying drawings.

[0012] The accompanying figure shows a schematic diagram of the method of the present invention, that is, generating spatial transcriptomics data from histological images. In conjunction with the accompanying figure, the method of the present invention is illustrated with a real case study. The specific steps are as follows: First, the spatial transcriptome dataset, which contains histological images, spatial gene expression data, and spot coordinates, is preprocessed. Training and validation datasets are then constructed. Next, the samples are divided into training and validation sets. A model based on a deep learning algorithm is then constructed and trained on the training set. During training, the optimal hyperparameters are determined using a leave-one-out-crossover validation method. The algorithm model specifically comprises the following modules: a convolution module, a custom lightweight convolutional layer with a Mamba module, a weight parameter sharing module, a Transformer module, and a GNN module. Subsequently, the histological images are divided into multiple fixed-size image blocks based on the spot coordinate information provided by the spatial transcriptome data. Each block is input into the model, which has been updated with the optimal parameters, to obtain a set of gene prediction results. Finally, the above steps are repeated, selecting one sub-dataset from each disjoint sub-dataset as the validation dataset and the remaining sub-datasets as the training dataset. This process continues until all sub-datasets have been trained as both validation and training datasets, resulting in gene prediction results for all groups. All prediction results are collected and displayed to the user in a table or graphical format for easy reference.

[0013] The above process will be described in more detail below with reference to the accompanying drawings.

[0014] Step 1: Dataset construction: Collect the spatial transcriptomics dataset generated by the 10X Genomics platform (including histological images, spatial gene expression data, and point coordinates) and preprocess the data. The specific steps are as follows: Download the dataset from the public database (https: / / www.synapse.org / ); normalize and transform the dataset into natural logarithm. The specific formula is:

[0015] in, Indicates the number of j The raw counts of genes, It is i The total number of spots, is the scale factor, Represents the processed data sample.

[0016] Step 2: Build a model based on a deep learning algorithm and train it on the training set. During the training process, the optimal hyperparameters are determined using a leave-one-out cross validation method. The constructed deep learning algorithm model includes: a convolution module, a custom lightweight convolution layer with a Mamba module, a weight parameter sharing module, a Transformer module, and a GNN module. Specifically: Step 2.1: The input data is used to capture the local feature relationship within the image block through convolution operation. The formula is:

[0017] in, Is the output feature map at position The value at Is the index of the input feature traversing the convolution kernel dimension (from 0 to K h -1 or K w -1); w is the weight matrix of the convolution kernel; K h and K w are the height and width of the convolution kernel respectively; b is the bias term in the convolution operation.

[0018] Step 2.2: Next, the convolutional features are input into a custom lightweight convolutional layer containing a Mamba module. The data is processed in multiple layers, including a dimensionality reduction layer and a dimensionality increase layer. The two layers are jump-connected using a weight parameter sharing module to achieve deeper feature extraction. The lightweight convolutional layer containing the Mamba module is specifically as follows:

[0019] in, (B is the batch size, C is the number of channels, and N is the spatial dimension after flattening); Indicates input Perform layer normalization; Indicates that Normalized It is divided into four parts along the channel dimension, recorded as ; Indicates the Apply the Mamba module; is the scaling factor of the jump connection (a learnable parameter); Indicates concatenating the four processed blocks in the channel dimension; is the weight matrix of the linear layer, which is used to project the concatenated results to obtain the final output .

[0020] The weight parameter sharing module is specifically:

[0021] in, is the input feature representation i Feature map of the branch input; represents the spatial attention part, where Express Perform global average pooling and maximum pooling in the spatial dimension to extract the global average and most significant information, and then perform splicing input convolution function In the Sigmoid activation function Get the spatial attention weight and use it to After element-by-element multiplication, add itself; Represents the channel attention part, performs global average pooling (GAP) on each feature processed by spatial attention, and then splices the GAP results of all scales into the global description , for each scale i , through the channel attention function Calculate channel attention and then get weight through Sigmoid activation Finally, the channel attention weight is multiplied by the feature after spatial attention fusion and the residual connection is added to get the output .

[0022] Step 2.3: Again, embed the spatial position code of each point into the extracted feature tensor, and input the fused features into the Transformer module to capture the global relationship. Specifically:

[0023] in, and is the 2D coordinate of each point encoded by the Pytorch function torch.nn.Embedding Fusion Features The multi-head attention layer in the Transformer module is used to capture the spatial connection between each patch. The multi-head attention layer is composed of multiple attention heads in a linear manner. Specifically:

[0024] in, is the weight matrix of the aggregated attention head, Q, K, and V represent Query, Key, and Value respectively. A single attention head is specifically:

[0025]

[0026] in, is the weight matrix; It is called Attention Map; Q=K=V. In the Transformer module, the final output dimension g maintains the fusion features of the input Same dimensions.

[0027] Step 2.4, finally use the GNN module to learn the domain relationship between each spot to build local spatial dependencies. We build the nearest neighbor graph Select the four nearest neighbors for each spot, where Represents the number of points, and E represents the edge connected to the nearest neighbor. The Euclidean distance between two points u and v is calculated as:

[0028] Then, GNN gathers the information of each point through the aggregation layer, specifically:

[0029]

[0030] in, represents the neighbor set of node v; AGG is the mean aggregator, is the historical feature of the aggregated neighbor node u, It is the domain aggregation feature, is a learnable weight matrix, is the activation function RELU.

[0031] Step 3: Based on the gene expression prediction values ​​obtained by the deep learning algorithm model, we use the Pearson correlation coefficient (PCC) and Rand index (ARI) to evaluate the accuracy of the model's predicted gene expression and clustering effect, respectively. Specifically:

[0032] in, and are the observed and predicted values ​​of gene expression, respectively; represents covariance; Represents variance.

[0033]

[0034] in, is the total number of data samples, The true label i Class and predicted label j The size of the intersection between classes; Indicates the true label i The total number of elements in the class; Indicates the predicted label j The total number of elements in the class.

[0035] Step 4: Input the test data set into the model after the optimal parameter update to obtain a set of gene prediction results.

[0036] Step 5: Select one set of sub-datasets from the disjoint sub-datasets as the validation dataset, and the remaining sub-datasets as the training dataset. Repeat steps 2, 3, and 4 until all sub-datasets are trained as validation datasets and training datasets, and obtain accurate gene prediction results for all groups.

[0037] In summary, the embodiments of the present invention provide a strong guarantee for revealing tissue molecular characteristics, specifically through: technically optimizing the network architecture and standardized data preprocessing process; applied through accurate gene expression and spatial region prediction; and experimentally ensuring model stability and robustness through multi-stage cross-validation. This multi-dimensional safeguard ensures that the model operates efficiently while also demonstrating significant advantages in actual clinical and scientific research.

[0038] On the other hand, the present invention uses the Mamba module: adopts the state space model (SSM) to model the time series with linear complexity, significantly reducing the amount of calculation; the weight sharing module: shares weight parameters in the channel attention and spatial attention calculations, reduces redundant parameters, and reduces the number of model parameters by 30%-50%, making the model possible for lightweight applications; by generating enhanced samples (random grayscale, rotation, flipping) and introducing learnable weights, the problem of small sample data is solved.

[0039] On the other hand, the present invention uses a convolution module to extract the texture and morphological features of image blocks at the local visual feature level, retaining the microscopic details of the histological image; the self-attention mechanism of the Transformer module models the association of long-distance spots, solving the problem of traditional methods ignoring the global structure; the GNN module explicitly aggregates the gene expression patterns of adjacent spots, capturing the spatial heterogeneity of the tissue microenvironment (such as tumor infiltration boundaries), and on the HER2+ breast cancer dataset, the PCC is improved by about 3%.

[0040] Applications have proven that the present invention can effectively generate spatial transcriptome data from conventional H&E-stained sections (WSIs), with costs significantly lower than those of ST technology, thereby promoting the popularization of precision medicine. It also provides multi-technology platform adaptation, supporting mainstream spatial transcriptome platforms such as 10xVisium, Space-TREX, and Stereo-seq. At the same time, the pre-trained model can be directly migrated to new tissue types (such as mouse brain and human dorsolateral prefrontal cortex) with only a small amount of annotated data fine-tuning. Analysis at the molecular mechanism level has also been effectively applied: the predicted gene expression profiles are significantly enriched in tumor-related pathways (such as HMGB2 and TFF3), two genes closely related to breast cancer. The associations between spots captured by GNN can identify the tumor microenvironment, which can provide effective guidance for targeted therapy. In terms of interpretability and visualization support: the prediction results are superimposed on the H&E image as a pseudo-color heat map, intuitively displaying the spatial heterogeneity of gene expression.

[0041] This invention addresses the high cost and data scarcity of spatial transcriptomics through a lightweight architecture, multi-scale feature fusion, and cross-platform generalization. It also achieves leading-edge prediction accuracy and biological interpretability. Its core advantages include: clinical applicability, adapting to routine pathology slides and lowering the threshold for scientific research; scientific rigor, preserving spatial neighborhood relationships and biological pathway information; and technological foresight, integrating cutting-edge technologies such as Mamba and self-distillation to advance the industrialization of large-scale model-based pathology. This technology provides an efficient and cost-effective solution for tumor microenvironment research, drug development, and precision medicine.

Claims

1. A method for generating spatial transcriptomics data from histological images based on deep learning, characterized in that: The specific steps are: Step 1: A standardized preprocessing process was performed on the spatial transcriptomics dataset generated by the 10X Genomics platform to construct a training / test set and select the target dataset, which contained histological images, spatial gene expression data, and point coordinates. Step 2: Build a model based on a deep learning algorithm and train it on the training set. During the training process, the optimal hyperparameters are determined through the leave-one-out cross validation method. Building a model based on a deep learning algorithm involves segmenting full-field tissue sections into image blocks centered on spatial transcriptome sequencing points, which are then fed into a convolutional neural network to extract local latent features. Furthermore, a global correlation model across image blocks is established through the collaborative architecture of the state-space model network VMamba and the Transformer. Finally, a graph neural network is integrated with a self-attention mechanism to dynamically represent the relationship between spatial sites and gene expression. Step 3: Input the test data set into the model with the updated optimal parameters to obtain a set of gene prediction results; Step 4: Select one sub-dataset from each of the disjoint sub-datasets as the validation dataset, and the remaining sub-datasets as the training datasets. Repeat steps 3 and 4 until all sub-datasets are trained as validation datasets and training datasets. Gene prediction results for all groups are obtained, and all prediction results are collected to evaluate model performance.

2. The method for generating spatial transcriptomics data from histological images based on deep learning according to claim 1, characterized in that Step 1 included selecting the top 1000 highly variable genes in each tissue section, excluding genes with less than 1000 expression spots in all tissue sections, followed by normalization and natural logarithm transformation; For histological images, they are divided into N blocks of size 3×W×H, where W=H=112 pixels, and the dataset is divided into training and test sets.

3. The method for generating spatial transcriptomics data from histological images based on deep learning according to claim 2, characterized in that The formula for normalizing and natural logarithm transforming the sample data set in step 1 is: (1) in, Indicates the number of j The raw counts of genes, It is i The total number of spots, is the scale factor, Represents the processed data sample.

4. The method for generating spatial transcriptomics data from histological images based on deep learning according to claim 1, characterized in that The specific steps for building a model based on deep learning algorithm in step 2 are: First, the input data is convolved to capture local feature relationships within the image block, and the activation function GELU is added after each convolution layer. Second, the convolved features are input into a custom lightweight convolution layer containing a Mamba module, and the data is processed in multiple layers, including dimensionality reduction and dimensionality increase layers, which are jump-connected through a weight-sharing module to achieve the extraction of deeper features. Third, the spatial position encoding of each point is embedded in the extracted feature tensor, and the fused features are input into the Transformer module to capture global relationships. The graph neural network module is then used to learn the domain relationships between each spot. Finally, the learned features are used to predict gene expression and the prediction results are scored using evaluation indicators. The model based on the deep learning algorithm includes the following modules: convolution module, customized lightweight convolution layer with Mamba module, weight parameter sharing module, Transformer module and GNN module.

5. The method for generating spatial transcriptomics data from histological images based on deep learning according to claim 4, characterized in that The input data captures the local feature relationship within the image block through convolution operation, the formula is as follows: (2) in, Is the output feature map at position The value at Is the index of the input feature traversing the convolution kernel dimension, from 0 to K h -1 or Kw-1; w is the weight matrix of the convolution kernel; K h and K w are the height and width of the convolution kernel respectively; b is the bias term in the convolution operation.

6. The method for generating spatial transcriptomics data from histological images based on deep learning according to claim 4, characterized in that The specific formula of the customized lightweight convolutional layer containing the Mamba module is as follows: (3) in, , B is the batch size, C is the number of channels, and N is the spatial dimension after flattening; Indicates input Perform layer normalization; Indicates that the normalized It is divided into four parts along the channel dimension, recorded as ; Indicates the Apply the Mamba module; is the scaling factor of the jump connection, which is a learnable parameter; Indicates concatenating the four processed blocks in the channel dimension; is the weight matrix of the linear layer, which is used to project the concatenated results to obtain the final output .

7. The method for generating spatial transcriptomics data from histological images based on deep learning according to claim 4, characterized in that: The specific formula of the weight parameter sharing module is as follows: (4) in, is the input feature representation i Feature map of the branch input; represents the spatial attention part, where right Perform global average pooling and maximum pooling in the spatial dimension to extract the global average and most significant information, and then input the convolution function after splicing. , after Sigmoid activation function Get the spatial attention weight, which is After element-wise multiplication, add itself; Represents the channel attention part, performs global average pooling (GAP) on the features processed by spatial attention, and splices the GAP of all scales into the global description , for each scale , through the channel attention function Calculate channel attention and then activate it with Sigmoid to get the weight Finally, the channel attention weight is multiplied by the feature of the fused spatial attention, and the residual connection is added to get the output .

8. The method for generating spatial transcriptomics data from histological images based on deep learning according to claim 4, characterized in that: The spatial position encoding of each point is embedded into the extracted feature tensor, and the fused features are input into the Transformer module to capture the global relationship, as follows: (5) in, and is the 2D coordinate of each point encoded by the function torch.nn.Embedding in Pytorch; Fusion Features The multi-head attention layer in the Transformer module is used to capture the spatial connection between each patch. The multi-head attention layer is composed of multiple attention heads in a linear manner. Specifically: (6) in, is the weight matrix of the aggregated attention head, Q, K, and V represent Query, Key, and Value respectively; a single attention head is specifically: (7) (8) in, is the weight matrix; It is called Attention Map; Q=K=V; in the Transformer module, the final output dimension g maintains the fusion features of the input Same dimensions.

9. The method for generating spatial transcriptomics data from histological images based on deep learning according to claim 4, characterized in that: Finally, the GNN module is used to learn the domain relationship between each spot to build local spatial dependencies, including building the nearest neighbor graph , select the four nearest neighboring points for each spot, where Represents the number of points, E represents the edge connected to the nearest neighbor; the Euclidean distance between points u and v is calculated as: (9) Then, GNN gathers the information of each point through the aggregation layer, specifically: (10) (11) in, represents the neighbor set of node v; AGG is the mean aggregator, is a learnable weight matrix, is the activation function RELU.

10. The method for generating spatial transcriptomics data from histological images based on deep learning according to claim 4, characterized in that: It also includes algorithm model evaluation indicators based on deep learning, using the Pearson correlation coefficient PCC and the Rand index ARI to evaluate the accuracy of the model's predicted gene expression and clustering effect, respectively. Specifically: (12) in, and are the observed and predicted values ​​of gene expression, respectively; represents covariance; represents variance; (13) in, is the total number of data samples, The true label Class and predicted label j The size of the intersection between classes; Indicates the true label i The total number of elements in the class; Indicates the predicted label j The total number of elements in the class.