A method, device and equipment for predicting super-resolution gene expression maps

By constructing a hierarchical feature extraction model and a GCN-based super-resolution gene expression prediction model, using histological images and spatial transcriptome data, the problem of low gene expression accuracy in the prior art is solved, and a more accurate and comprehensive prediction of super-resolution gene expression map is achieved.

CN118486375BActive Publication Date: 2025-05-16YUNNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410709969.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-03
Publication Date
2025-05-16
Estimated Expiration
2044-06-03

AI Technical Summary

Technical Problem

The prior art fails to fully utilize the correlation between gene expression patterns and histological image features, resulting in a low gene expression accuracy of super-resolution tissue structures.

Method used

By acquiring histological images and spatial transcriptome data, a hierarchical feature extraction model is constructed, including local ViT modules, global ViT modules and feature stacking modules, and multi-scale histological features are extracted. Then, weakly supersupervised learning of superresolution gene expression prediction model based on GCN graph neural network is used to generate a more accurate superresolution gene expression map.

Benefits of technology

It achieves a more accurate and comprehensive prediction of super-resolution gene expression map, covering the entire tissue area, and improves gene expression resolution and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118486375B_ABST
    Figure CN118486375B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of neural network prediction technology, and discloses a prediction method, device and equipment for super-resolution gene expression maps. The method comprises: after dividing a histological image into large image blocks and small image blocks according to pixel size, the large image blocks, the small image blocks and the original histological image are input into a pre-trained hierarchical feature extraction model, and a hierarchical histological feature map is output; each super-pixel point in the hierarchical histological feature map is used as a node, divided into a series of sub-graphs, and the series of sub-graphs are input into a trained super-resolution gene expression prediction model, and a predicted super-resolution gene expression map is output. The present invention uses the ViT module to perform multi-scale extraction of histological image features, and captures the complex correlation information between adjacent super-pixel points through a graph neural network, so as to obtain a super-resolution gene expression map that covers the entire tissue area more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of neural network prediction technology, and in particular to a method, device and equipment for predicting super-resolution gene expression maps. Background Art

[0002] With the continuous progress of research, spatial transcriptomics technology has been applied to various biological fields, such as tumor microenvironment, embryonic development, and neurology. However, all platforms of spatial transcriptomics technology do not provide single-cell resolution and cannot cover the entire transcriptome.

[0003] In existing spatial transcriptomics experimental methods, there is usually a trade-off between resolution and gene throughput. Imaging-based methods provide higher resolution and sensitivity, but lower gene throughput. Next-generation sequencing-based methods can in principle sample the entire transcriptome, but have lower resolution and sensitivity, which limits their ability to study detailed gene expression patterns. For example, in the ST platform, the diameter of the spot (spatial transcriptome sequencing point) is 100μm, the center distance between two spots is 200μm, and the size of a single cell is about 8*8μm. 2 , which is far from single-cell resolution, greatly affects researchers' understanding of tissue or cellular heterogeneity and makes it impossible to more finely resolve the differences between different cell types and subtypes in tissues.

[0004] At present, researchers have proposed a variety of methods to improve the gene expression resolution of transcriptomics. However, all existing methods have their shortcomings. DIST (spatial transcriptomics enhancement using deep learning) uses convolutional neural networks to predict the uncovered areas between spots under self-supervised learning, but it does not reduce the coverage area of ​​the spots and cannot reach the single-cell level. BayesSpace (Spatial transcriptomics at subspot resolution with BayesSpace) divides the spots into a series of sub-spots, but does not predict the unmeasured areas. iStar (Inferring super-resolution tissue architecture by integrating spatial transcriptomics with histology) can improve the gene expression of transcriptomics to the single-cell level and cover the entire transcriptome, but it ignores the connection between adjacent cells, resulting in low accuracy.

[0005] Therefore, the prior art still needs to be improved and developed. Summary of the invention

[0006] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to provide a method, device and equipment for predicting super-resolution gene expression maps, aiming to solve the problem that the prior art does not fully utilize the characteristics of gene expression patterns related to histological image features, resulting in low accuracy of gene expression in predicted super-resolution tissue structures.

[0007] The first aspect of the present invention provides a method for predicting a super-resolution gene expression map, comprising: obtaining a histological image and corresponding spatial transcriptome data; dividing the histological image into a large image block and a small image block according to a size of 4096*4096 pixels and a size of 256*256 pixels, respectively; constructing a hierarchical feature extraction model, the hierarchical feature extraction model comprising a local ViT module, a global ViT module and a feature stacking module; inputting the large image block, the small image block and the histological image into the hierarchical feature extraction model, and outputting a hierarchical histological feature map; using the hierarchical histological feature map as input data and the corresponding spatial transcriptome data as pseudo labels, training a super-resolution gene expression prediction model based on a GCN graph neural network through weak supervised learning to obtain a trained super-resolution gene expression prediction model; extracting features of a histological image to be predicted based on the hierarchical feature extraction model to obtain a hierarchical histological feature map to be predicted; using each superpixel point in the hierarchical histological feature map to be predicted as a node, dividing it into a series of sub-maps, inputting the series of sub-maps into the trained super-resolution gene expression prediction model, and outputting a predicted super-resolution gene expression map.

[0008] Optionally, in a first implementation of the first aspect of the present invention, the large image block, the small image block and the histological image are input into the hierarchical feature extraction model to output a hierarchical histological feature map, including the steps of: assuming that the height and width of the histological image X are M and N respectively, and the large image block after division is X mn , then the histological image is represented by a large image block as The large image block X mn Divide into small image blocks X mnpq , then the large image block is represented as The histological image is represented by small image blocks: The small image block is divided into 256 sub-image blocks of size 16*16. After the sub-image blocks are encoded and embedded to form a two-dimensional matrix, a random vector Token is added as the input of the local ViT module, and the Token corresponding to each sub-image block is obtained as a low-level local feature vector, and a CLS-Token representing the high-level local feature vector of the small image block; the CLS-Token of all small image blocks contained in the large image block is merged into a matrix as the input of the global ViT module, and a high-level global feature vector integrating the information of all small image blocks is obtained; based on the low-level local feature vector and the high-level global feature vector, a low-level local feature map and a high-level global feature map are formed respectively; after the low-level local feature map, the high-level global feature map and the original histological image are size-aligned, they are input into the feature stacking module for stacking in the channel dimension, and a hierarchical histological feature map is output.

[0009] Optionally, in a second implementation of the first aspect of the present invention, the formula for obtaining the low-level local feature vector is: Among them, f0 represents the component of feature information extracted in the local ViT module, The dimension of the feature vector is c0; the formula for obtaining the high-level local feature vector is: Among them, f1 represents the part of the local ViT module that generates CLS-Token by fusing feature information. The dimension of the feature vector is c1; the expression formula of the high-level global feature vector is: Among them, f2 represents the global ViT module.

[0010] Optionally, in a third implementation of the first aspect of the present invention, after the low-level local feature map, the high-level global feature map and the original histological image are size-aligned, they are input into the feature stacking module for stacking in the channel dimension, and the step of outputting a hierarchical histological feature map includes: assuming that the width and height of the super-resolution gene expression map to be predicted are M′×N′, if the size of the low-level local feature map, the high-level global feature map and the original histological image is less than M′×N′, the image is reduced to M′×N′ by averaging; if the size is greater than M′×N′, the image is expanded to M′×N′ by using a broadcasting mechanism; the aligned low-level local feature map, the high-level global feature map and the original histological image are input into the feature stacking module for stacking in the channel dimension to obtain a hierarchical histological feature map. Among them, h mn is the hierarchical histological feature vector at the superpixel point (m,n).

[0011] Optionally, in a fourth implementation of the first aspect of the present invention, the stratified histological feature map is used as input data and the corresponding spatial transcriptome data is used as a pseudo-label, and the super-resolution gene expression prediction model based on the GCN graph neural network is trained through weakly supervised learning to obtain the trained super-resolution gene expression prediction model. The steps include: constructing a super-resolution gene expression prediction model based on the GCN graph neural network, the super-resolution gene expression prediction model comprising two layers of GCN modules and an output module consisting of a linear layer and an exponential linear unit, and the weakly supervised loss function of the super-resolution gene expression prediction model is Among them, S is the total number of spots in the entire histological image, K is the number of predicted genes, and f k is the gene expression prediction model for predicting k genes, g ks is the expression level of gene k at spot s, is the mask of s in spot, that is, the set of superpixels covered by s in spot, hmn is the hierarchical histological feature vector at the superpixel point (m,n), and θ is the parameter of the super-resolution gene expression prediction model; the hierarchical histological feature map is used as input, and its corresponding spatial transcriptomics data is used as the label to train the super-resolution gene expression prediction model. During the training process, the gradient of the weak supervision loss function is calculated by the back propagation algorithm, and the parameters of the super-resolution gene expression prediction model are updated to minimize the weak supervision loss function, thereby obtaining the trained super-resolution gene expression prediction model.

[0012] Optionally, in a fifth implementation of the first aspect of the present invention, each superpixel point in the stratified histological feature map to be predicted is taken as a node and divided into a series of sub-maps, and the series of sub-maps are input into the trained super-resolution gene expression prediction model, and the predicted super-resolution gene expression map is output, comprising the steps of: taking each superpixel point in the stratified histological feature map H to be predicted as a node and dividing it into n series of sub-maps, then H = {H1, H2, H i K,H n}, where H i represents the ith series subgraph; the n series subgraphs are input into the trained super-resolution gene expression prediction model, the formula of the super-resolution gene expression prediction model is as follows: H′ i =Flatten(H i ), Y i 0 =σ(AH′ i W (0) ), Y i 1 =σ(AY i W (1) ), Z′ i =ELU(Y i 1 W (2) +b) and Z i =Flatten -1 (Z′ i ), where Flatten means converting the shape of the series subgraph from three-dimensional to two-dimensional, and H′ i represents a two-dimensional series of subgraphs; σ represents the nonlinear activation function ReLU; W (0) , W (1) are the weight matrices of the first and second layers of GCN respectively; W (2) and b are the weight matrix and bias term of the output layer respectively; Y i 0 represents the output of the first layer of GCN, Y i 1 represents the output of the second layer GCN, A represents the adjacency matrix of the superpixel point, Z′i Represents the output of the output module, Flatten -1 Indicates converting a two-dimensional matrix into a three-dimensional matrix, Z i is a gene expression subgraph, whose number of channels is the number of genes K; n gene expression subgraphs are merged into a predicted super-resolution gene expression map Z = {Z1, Z2, K, Z n}.

[0013] The second aspect of the present invention provides a prediction device for super-resolution gene expression maps, including: an acquisition module for acquiring a histological image and corresponding spatial transcriptome data; a division module for dividing the histological image into large image blocks and small image blocks according to the sizes of 4096*4096 pixels and 256*256 pixels, respectively; a model construction module for constructing a hierarchical feature extraction model, wherein the hierarchical feature extraction model includes a local ViT module, a global ViT module, and a feature stacking module; a sample image feature extraction module for inputting the large image block, the small image block, and the histological image into the hierarchical feature extraction model and outputting a hierarchical histological feature map; a training module for The method comprises the following steps: using the stratified histological feature map as input data and the corresponding spatial transcriptome data as pseudo-labels, training a super-resolution gene expression prediction model based on a GCN graph neural network through weakly supervised learning to obtain a trained super-resolution gene expression prediction model; a feature extraction module for an image to be tested, used for extracting features of a histological image to be predicted based on the stratified feature extraction model to obtain a stratified histological feature map to be predicted; and a prediction module, used for taking each superpixel point in the stratified histological feature map to be predicted as a node, dividing the result into a series of sub-maps, inputting the series of sub-maps into the trained super-resolution gene expression prediction model, and outputting a predicted super-resolution gene expression map.

[0014] A third aspect of the present invention provides a prediction device for a super-resolution gene expression map, comprising: a memory and at least one processor, wherein the memory stores computer-readable instructions, and the memory and the at least one processor are interconnected via lines; the at least one processor calls the computer-readable instructions in the memory so that the prediction device for the super-resolution gene expression map performs each step of the prediction method for the super-resolution gene expression map as described above.

[0015] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-readable instructions, which, when executed on a computer, enable the computer to execute the various steps of the method for predicting super-resolution gene expression maps as described above.

[0016] Beneficial effects: The present invention provides a method for predicting super-resolution gene expression maps. Based on histological images and spatial transcriptome data, the ViT module is used to perform multi-scale extraction of histological image features, and the complex correlation information between adjacent superpixels is captured through a graph neural network, thereby obtaining a more accurate super-resolution gene expression map that can cover the entire tissue area. The present invention fully utilizes the characteristics of gene expression patterns related to histological image features, and further utilizes the strong correlation between adjacent cells through a graph neural network, making the predicted gene expression more accurate and comprehensive. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flow chart of a method for predicting a super-resolution gene expression map provided by an embodiment of the present invention;

[0018] Figure 2 This is a training flow chart of the hierarchical feature extraction model and the super-resolution gene expression prediction model of the present invention.

[0019] Figure 3 It is a box plot formed by the super-resolution gene expression predicted by different methods in Example 1 and the RSME index of the real data.

[0020] Figure 4 It is a box plot formed by the SSIM index of the super-resolution gene expression predicted by different methods in Example 1 and the real data.

[0021] Figure 5 This is a visualization heat map of the super-resolution gene expressions predicted by different methods based on the HBC_S1R1 dataset in Example 1.

[0022] Figure 6 A schematic diagram of the structure of a prediction device for super-resolution gene expression maps provided by an embodiment of the present invention;

[0023] Figure 7 A schematic diagram of the structure of a super-resolution gene expression map prediction device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0024] Embodiments of the present invention provide a method, apparatus, device and storage medium for predicting super-resolution gene expression maps. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] The rapid development of spatial transcriptomics technology has made it possible to measure gene expression in the original tissue context and retain spatial information. Understanding the relative position of transcripts is crucial for understanding biological, physiological and pathological processes. Therefore, spatial transcriptomics technology has been used in various biological fields. An ideal spatial transcriptomics platform should provide single-cell resolution and cover the entire transcriptome. However, existing technologies still have great difficulties in generating such data.

[0026] Based on this, the present invention provides a method for predicting super-resolution gene expression profiles, such as Figure 1 As shown, it includes the steps of:

[0027] S10, obtaining histological images and corresponding spatial transcriptome data;

[0028] Specifically, spatial transcriptomics is a technology and method for studying the spatial distribution and tissue specificity of gene expression at the tissue or cell level. It provides an in-depth understanding of the expression location and spatial relationship of genes in tissues by combining transcriptomics data with spatial information. Spatial transcriptome data is obtained through spatial transcriptome sequencing technology, which can measure the total mRNA of intact tissues on tissue sections, and combine the spatial information of total mRNA with morphological content to map the location where all gene expressions occur. In this process, histological images play an important role. The histological images are pictures of biological tissue sections or cell structures observed and photographed under a microscope. First, histological images are an important source of data for spatial transcriptomics research. Histological images observed and photographed under a microscope can show the microstructure and cell type of biological tissues, providing rich information for spatial transcriptomics. Based on these histological images, researchers can combine high-throughput sequencing technology and bioinformatics analysis methods to quantitatively analyze the gene expression in tissue sections, and then study the expression pattern of genes in space. Secondly, by comparing the histological images and gene expression data of different tissue sections or different cell types, researchers can verify the accuracy and reliability of spatial transcriptomics analysis results. There is a close correlation between histological images and spatial transcriptomics. They are interdependent and mutually reinforcing, and jointly promote the development of biology and medicine. Therefore, in order to accurately improve the gene expression resolution in spatial transcriptomics to a level close to that of a single cell and cover the entire tissue area, the present invention needs to first obtain histological images and their corresponding spatial transcriptome data as training samples.

[0029] S20, dividing the histological image into large image blocks and small image blocks according to the sizes of 4096*4096 pixels and 256*256 pixels respectively;

[0030] Specifically, in order to facilitate feature extraction of histological images with different resolutions, the present embodiment may first scale the histological images so that the size of each pixel is 0.5×0.5 μm. 2 ; This scaling ensures that a 16×16 pixel image block has a size of 8×8μm 2 , which is about the size of a single cell. In order to simplify the subsequent division process, the rescaled image is cropped so that its height and width are both divisible by 256. Then, this embodiment divides the scaled histological image into image blocks hierarchically, where the large image block (4096*4096 pixel size) reflects the global structure of the tissue, and the small image blocks (256*256 pixel size) divided from the large image block reflect the local fine-grained cell structure of the tissue. Suppose the height and width of the histological image X are M and N respectively, and the large image block after division is X mn, then the histological image is represented by a large image block as The large image block X mn Divide into small image blocks X mnpq , then the large image block is represented as The histological image is represented by small image blocks:

[0031] S30, constructing a hierarchical feature extraction model, wherein the hierarchical feature extraction model includes a local ViT module, a global ViT module, and a feature stacking module;

[0032] In this embodiment, the local ViT module and the global ViT module both refer to different variants or application methods based on the Visual Transformer (ViT) model, which are used to process image tasks. Generally speaking, the ViT model regards images as sequence data, captures spatial and temporal information in images through a self-attention mechanism, and captures position information in images through position encoding; this global processing method enables the ViT model to have a strong global perception ability and better understand the overall structure and content of the image. The local ViT module of this embodiment can be represented by f0*f1, and the global ViT module can be represented by f2, both of which are used to extract histological image features at different levels.

[0033] Due to the benefits of transfer learning, the hierarchical feature extraction model constructed in this embodiment uses only histological images, without gene expression or other labels, and can be pre-trained by a self-supervised learning method. As an example, DINO (a label-free self-distillation form, a self-supervised learning method) can be used to train the hierarchical feature extraction model on the histological image. The pre-trained hierarchical feature extraction model can already capture the histological features well without fine-tuning. Therefore, the large image block and the small image block and the original histological image are input into the pre-trained hierarchical feature extraction model, and the hierarchical histological feature map can be output.

[0034] S40, inputting the large image block, the small image block and the histological image into the hierarchical feature extraction model, and outputting a hierarchical histological feature map;

[0035] In this embodiment, the small image block is divided into 256 sub-image blocks of size 16*16. After the sub-image blocks are encoded and embedded to form a two-dimensional matrix, a random vector Token is added as the input of the local ViT module, and the Token corresponding to each sub-image block is obtained as a low-level local feature vector, and a CLS-Token representing a high-level local feature vector of the small image block is obtained; the CLS-Tokens of all small image blocks contained in the large image block are merged into a matrix as the input of the global ViT module, and a high-level global feature vector integrating information of all small image blocks is obtained; a low-level local feature map and a high-level global feature map are formed based on the low-level local feature vector and the high-level global feature vector, respectively; after the low-level local feature map, the high-level global feature map and the original histological image are size-aligned, they are input into the feature stacking module for stacking in the channel dimension, and a hierarchical histological feature map is output.

[0036] Specifically, if Figure 2 As shown in the figure, the shape of the small image block is (256, 256, 3). The small image block is first divided into sub-image blocks (Patch), each sub-image block is 16*16 in size, corresponding to a Token, and a small image block has 256 sub-image blocks in total; after encoding and embedding, a small image block becomes a two-dimensional matrix of (256, C); then a random vector Token is added to the two-dimensional matrix, and its purpose is to fuse all sub-image information as the representative of this small image block; then the shape of the input small image block becomes (257, C), as the input of the local ViT module, the output shape becomes (257, 384), the number of Tokens remains unchanged, still 257 (including 256 sub-image blocks and 1 representative of the small image block, that is, CLS-Token), only the dimension of each Token is changed. That is to say, the first 16*16 tokens correspond to 256 sub-image blocks divided in a small image block, and the last CLS-Token integrates all the information of the 256 sub-image blocks divided in a small image block, which serves as the representative vector of this small image block. This representative vector CLS-Token only integrates the information of the small image block, so this embodiment merges the CLS-Tokens of all small image blocks contained in a large image block into a matrix as the input of the global ViT module. The output shape remains unchanged, but the CLS-Tokens of all small image blocks contained in a large image block have been integrated to obtain a high-level global feature vector.

[0037] In this embodiment, the formula for obtaining the low-level local feature vector is: Among them, f0 represents the component of feature information extracted in the local ViT module, The dimension of the feature vector is c0; the formula for obtaining the high-level local feature vector is: Among them, f1 represents the part of the local ViT module that generates CLS-Token by fusing feature information. The dimension of the feature vector is c1; the expression formula of the high-level global feature vector is: Among them, f2 represents the global ViT module.

[0038] In this embodiment, a low-level local feature map and a high-level global feature map are formed based on the low-level local feature vector and the high-level global feature vector, wherein the high-level global feature map Its size is (M / 4096)×(N / 4096)×C1; low-level local feature map Its size is (M / 256)×(N / 256)×C0, and the size of the histology image is M×N×3.

[0039] In this embodiment, the sizes of the low-level local feature map, the high-level global feature map, and the original histological image are different. In order to obtain super-resolution gene expression data with a height and width of M′×N′, this embodiment uses a simple method to align the above three feature image data. If the size of the low-level local feature map, the high-level global feature map, and the original histological image is less than M′×N′, the image is reduced to M′×N′ by averaging; if the size is greater than M′×N′, the image is expanded to M′×N′ using a broadcasting mechanism. The aligned low-level local feature map, high-level global feature map, and original histological image are input into the feature stacking module and stacked in the channel dimension to obtain a layered histological feature map. Among them, h mn is the hierarchical histological feature vector at the superpixel point (m, n), and the size of the hierarchical histological feature map is (M′)×(N′)×(C0+C1+3).

[0040] S50, using the stratified histological feature map as input data and the corresponding spatial transcriptome data as pseudo-labels, training a super-resolution gene expression prediction model based on the GCN graph neural network through weakly supervised learning to obtain a trained super-resolution gene expression prediction model;

[0041] In this embodiment, a super-resolution gene expression prediction model based on a GCN graph neural network is first constructed. The super-resolution gene expression prediction model includes a two-layer GCN module and an output module composed of a linear layer and an exponential linear unit. The weak supervision loss function of the super-resolution gene expression prediction model is Among them, S is the total number of spots in the entire histological image, K is the number of predicted genes, and fk is the gene expression prediction model for predicting k genes, g ks is the expression level of gene k at spot s, is the mask of s in spot, that is, the set of superpixels covered by s in spot, h mn is the hierarchical histological feature vector at the superpixel point (m, n), θ is the parameter of the super-resolution gene expression prediction model. In this embodiment, the area of ​​a sequencing point (spot) usually includes dozens of superpixels. The size of a superpixel point in real space is approximately the size of a single cell, and the sequencing point contains dozens of single cells. Therefore, in this embodiment, the gene expression of each spot is modeled as the sum of the gene expression of all superpixels in the area covered by it.

[0042] The super-resolution gene expression prediction model is trained using the hierarchical histological feature map as input and its corresponding spatial transcriptomics data as labels. During the training process, a superpixel point is used as one of several nodes. After passing through two layers of GCN modules, it interacts with the surrounding superpixels and the dimension is mapped to 256. Then it is input into the output module, the dimension is mapped from 256 to the number of genes, and then activated by an exponential linear unit. Finally, the gradient of the weak supervision loss function is calculated through the back-propagation algorithm, and the parameters of the super-resolution gene expression prediction model are updated to minimize the weak supervision loss function to obtain the trained super-resolution gene expression prediction model. During the model training process, superpixels outside the spot coverage range do not participate in the training. After the model training process, the predicted gene expression value of gene k at the superpixel (m,n) is z kmn =f k (h mn ), in this way, we can obtain the expression map Z of all genes in the tissue, whose size is (M′)×(N′)×K.

[0043] To evaluate the accuracy of the super-resolution gene expression predicted by the model, for each gene, we consider both the true value and the predicted gene expression as images, where the image intensity is normalized to the range of [0.0, 1.0]. The prediction accuracy is then measured by the root mean square error (RMSE) and structural similarity (SSIM).

[0044] RMSE calculates the average squared error between the predicted value and the true value and squares it to keep the error dimension consistent with the original data. The smaller the RMSE, the more accurate the model's prediction. RMSE is often used to evaluate regression models, for example, to measure the prediction error of the model in a prediction task. To calculate RMSE, the true value and predicted gene expression images are flattened into vectors y i , RMSE is equal to the Euclidean distance between two vectors,

[0045] RMSE is a direct and fast metric to evaluate the prediction accuracy of any vectorizable result, but for image data, RMSE ignores the spatial context in the image. Therefore, in addition to RMSE, we also calculated SSIM to evaluate the similarity between the spatial structures of the true gene expression image and the predicted gene expression image. SSIM is an image similarity metric that is widely used in computer vision and medical imaging. It not only considers the brightness error, but also the structural and perceptual properties of the image, and its value ranges from -1 to 1. The closer to 1, the higher the similarity between images. In our context, SSIM can capture both the global features and the fine-grained spatial structure in super-resolution gene expression images. The larger the RMSE, the more accurate the model's prediction. Here, x and y represent the two images to be compared, and u x and u y Represent the mean brightness information of images x and y respectively, and Respectively represent the contrast information variance of images x and y, σ xy represents the structural information covariance of images x and y, c1 and c2 are two constants used to stabilize the calculation.

[0046] S60, performing feature extraction on the histological image to be predicted based on the hierarchical feature extraction model to obtain a hierarchical histological feature map to be predicted;

[0047] S70, taking each superpixel point in the layered histological feature map to be predicted as a node, dividing it into a series of sub-graphs, inputting the series of sub-graphs into the trained super-resolution gene expression prediction model, and outputting a predicted super-resolution gene expression map.

[0048] Specifically, after the histological image to be predicted is input into the hierarchical feature extraction model, the hierarchical histological feature map H to be predicted is obtained, and this embodiment uses it to predict the super-resolution gene expression map. The histological feature vector of each superpixel point in the hierarchical histological feature map contains not only its local cell feature information, but also global relationship information with other regions in the entire tissue. The high-level global feature map establishes long-distance spatial dependencies and expresses global relationship information, even if the superpixels are physically far away from each other. However, in the tissue structure, the gene expression of cells is highly correlated with their surrounding neighboring cells, far more than with tissue regions that are physically far away from each other. Therefore, when predicting gene expression at superpixels, this embodiment further uses the GCN module to capture the complex correlation information between adjacent superpixel points and generate a more accurate representation. Since the number of superpixel points is generally hundreds of thousands or even millions, in order to reduce the computational complexity, this embodiment divides the hierarchical histological feature map H into a series of sub-graphs. Assuming that it is divided into n series of sub-graphs, then H={H1,H2,H i K,H n}, where H i represents the ith series subgraph, H i The size is

[0049] Each superpixel is recorded as a node, and the four nearest nodes are selected as its neighbor nodes according to the physical distance to generate the adjacency matrix A. After the information fusion between adjacent cells of the two-layer GCN module, a 512-dimensional output representation is obtained at each superpixel; the output module contains a linear layer with 512 input nodes and K output nodes, where K represents the number of genes; the output module also contains an exponential linear unit ELU activation to ensure that the predicted gene expression is non-negative.

[0050] The n series of subgraphs are input into the trained super-resolution gene expression prediction model. The formula of the super-resolution gene expression prediction model is as follows: i =Flatten(H i ), Y i 0 =σ(AH′ i W (0) ), Y i 1 =σ(AY i W (1) ), Z′ i =ELU(Y i 1 W (2) +b) and Z i =Flatten -1 (Z′ i), where H i represents the i-th series sub-graph;

[0051] The n series of subgraphs are input into the trained super-resolution gene expression prediction model. The formula of the super-resolution gene expression prediction model is as follows: i =Flatten(H i ), Y i 0 =σ(AH′ i W (0) ), Y i 1 =σ(AY i W (1) ), Z′ i =ELU(Y i 1 W (2) +b) and Z i =Flatten -1 (Z′ i ), where Flatten means converting the shape of the series subgraph from three-dimensional to two-dimensional, and H′ i represents a two-dimensional series of subgraphs; σ represents the nonlinear activation function ReLU; W (0) , W (1) are the weight matrices of the first and second layers of GCN respectively; W (2) and b are the weight matrix and bias term of the output layer respectively; Y i 0 represents the output of the first layer of GCN, Y i 1 represents the output of the second layer GCN, A represents the adjacency matrix of the superpixel point, Z′ i Represents the output of the output module, Flatten -1 Indicates converting a two-dimensional matrix into a three-dimensional matrix, Z i is a gene expression subgraph, whose number of channels is the number of genes K; n gene expression subgraphs are merged into a predicted super-resolution gene expression map Z = {Z1, Z2, K, Z n}, whose size is (M′)×(N′)×K.

[0052] Based on histological images and spatial transcriptome data, the present invention uses ViTs to perform multi-scale extraction of histological image features, and captures the complex correlation information between adjacent superpixels through graph neural networks, thereby obtaining more accurate super-resolution gene expression of tissue structures. The present invention makes full use of the characteristics of gene expression patterns related to histological image features, and further utilizes the strong correlation between adjacent cells through graph neural networks, so that the predicted gene expression is smoother. The effectiveness of the present invention in combining two modal data to predict super-resolution gene expression maps is mainly reflected in: a) It can be used to improve the resolution of gene expression in spatial transcriptomics data, reconstruct subtle changes and structural information in tissue space, so that researchers can more accurately identify and analyze the differences and correlations between different cell types and states; b) Gene expression prediction is performed on uncovered areas (gaps between spots and background images, etc.), which promotes a comprehensive understanding of the entire cell population and can fully express the true state of the tissue, especially when studying cell heterogeneity and local differences is crucial.

[0053] The following is a further explanation of a method for predicting a super-resolution gene expression map of the present invention through specific examples:

[0054] Example 1

[0055] Since spatial transcriptomics data do not have super-resolution gene expression, we cannot evaluate the accuracy of the super-resolution gene expression predicted by the model, so we apply it to four 10X Xenium datasets (a type of data that reaches pixel-level gene expression). The abbreviations of the above four datasets are:

[0056] 1. HBC_S1R1:10x Xenium human breast cancer sample 1-replicate1data;

[0057] 2. HBC_S1R2:10x Xenium human breast cancer sample 1-replicate2data;

[0058] 3. HBC_S2R1:10x Xenium human breast cancer sample 1data;

[0059] 4. HP:10x Xenium human pancreas section data.

[0060] First, the gene expression of the Xenium dataset is divided into a rectangular network of superpixels, and the size of the superpixel is set to 0.5*0.5um 2 , the superpixel rectangular network is used as the real data (as a label for indicator evaluation, not involved in training). Then, according to the size, shape and layout of the sequencing spots of Visium (a type of spatial transcriptomics data), the gene expression of the Xenium dataset is divided into a series of sequencing sites: the diameter is 55um, and the distance from the center of adjacent sequencing spots is 100um. The generated data is used as pseudo Visium (idle) data.

[0061] The method of the present invention uses histological images as input and pseudo-Visium data as pseudo-labels, and obtains predicted super-resolution gene expression through a graph neural network trained by weakly supervised learning.

[0062] Existing benchmark methods iStar, XFuse, and TESLA can also obtain super-resolution gene expression through histological images and idle data; the benchmark method STAGE uses the relationship between spatial coordinates and gene expression to obtain high-density gene expression through autoencoders.

[0063] The method of the present invention and the other four benchmark methods are applied to the above four data sets to predict super-resolution gene expression, among which iStar infers super-resolution gene expression through histological image features, but it does not further consider the close relationship between adjacent cells and does not effectively model it; XFuse uses a deep generative model to characterize the transcriptome of micrometer-scale anatomical features, but it is overly dependent on histological images and performs well only on a few highly variable genes closely related to histological images, and cannot predict the super-resolution gene expression of other genes well, and the running speed is too slow, and the super-resolution gene expression inference of a data set takes several days or even dozens of days. TESLA believes that the gene expression at the superpixel point is expected to be similar to the gene expression of its adjacent sequencing sites, so TESLA detects the sequencing sites of its first 10 nearest neighbors based on the Euclidean distance, and then calculates the weighted average. This method does not take into account the global dependency, nor does it perform deep feature extraction on histological images; STAGE models the relationship between the two-dimensional spatial coordinates of the sequencing site and gene expression through an autoencoder, and does not take into account the mutual dependency between histological images and gene expression at all. The predicted super-resolution gene expression and the real data are compared with two indicators (RSME and SSIM). The box plots based on these two indicators are as follows: Figure 3 and Figure 4As shown in the figure, it can be seen that the RSME value of the method provided by the present invention is lower than that of all the benchmark methods, and its SSIM value is greater than that of all the benchmark methods. Specifically, on the first data set, our upper quartile is close to the lower quartile of the best benchmark method iStar, and on the last three data sets, our median is also close to the lower quartile; the same is true for the SSIM indicator.

[0064] Furthermore, based on the six genes in the HBC_S1R1 dataset, the super-resolution gene expression profiles predicted by the existing benchmark method and the pseudo-Visium data of the present invention were visualized as heat maps. Figure 4 As shown in the figure, each column is a visualized heat map of a gene, and the gene name is marked above each column. The first row in the figure is pseudo-Visium data; the second row is the true value of the super-resolution gene expression data; the third row is the super-resolution gene expression data predicted by the method of the present invention, and the HBC_S1R1 dataset is used for training and prediction; the fourth row is the super-resolution gene expression data predicted by the method of the present invention, and the HBC_S1R2 dataset is used for training, and the HBC_S1R1 dataset is used for prediction; the fifth row is the super-resolution gene expression data predicted by the iStar method, and the HBC_S1R1 dataset is used for training and prediction; the sixth row is the super-resolution gene expression data predicted by the iStar method, and the HBC_S1R2 dataset is used for training, and the HBC_S1R1 dataset is used for prediction; the seventh row is the super-resolution gene expression data predicted by XFuse, and the eighth row is the super-resolution gene expression data predicted by TESLA. From Figure 5 As can be seen from the results, the images in the third and fourth rows are closest to the true values ​​in the second row, that is, the super-resolution gene expression data predicted by the method of the present invention are closest to the true values.

[0065] The above describes the prediction method of the super-resolution gene expression map in the embodiment of the present invention. The following describes the prediction device of the super-resolution gene expression map in the embodiment of the present invention. Figure 6 In one embodiment of the present invention, a prediction device for super-resolution gene expression profiles includes:

[0066] An acquisition module 10, for acquiring histological images and corresponding spatial transcriptome data;

[0067] A division module 20, for dividing the histological image into large image blocks and small image blocks according to the sizes of 4096*4096 pixels and 256*256 pixels respectively;

[0068] A model building module 30, used to build a hierarchical feature extraction model, wherein the hierarchical feature extraction model includes a local ViT module, a global ViT module and a feature stacking module;

[0069] A sample image feature extraction module 40, used for inputting the large image block, the small image block and the histological image into the hierarchical feature extraction model, and outputting a hierarchical histological feature map;

[0070] A training module 50 is used to train a super-resolution gene expression prediction model based on a GCN graph neural network by using the stratified histological feature map as input data and the corresponding spatial transcriptome data as pseudo labels through weakly supervised learning to obtain a trained super-resolution gene expression prediction model;

[0071] The image feature extraction module 60 is used to extract features of the histological image to be predicted based on the hierarchical feature extraction model to obtain a hierarchical histological feature map to be predicted;

[0072] The prediction module 70 is used to divide each superpixel point in the layered histological feature map to be predicted into a series of sub-maps as a node, input the series of sub-maps into the trained super-resolution gene expression prediction model, and output a predicted super-resolution gene expression map.

[0073] above Figure 6 The prediction device for super-resolution gene expression maps in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The prediction device for super-resolution gene expression maps in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0074] Figure 7 1 is a schematic diagram of the structure of a prediction device for a super-resolution gene expression map provided by an embodiment of the present invention. The prediction device 100 for the super-resolution gene expression map may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 11 (for example, one or more processors) and a memory 12, and one or more storage media 13 (for example, one or more massive storage devices) storing application programs 133 or data 132. Among them, the memory 12 and the storage medium 13 can be short-term storage or permanent storage. The program stored in the storage medium 13 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the prediction device 100 for the super-resolution gene expression map. Furthermore, the processor 11 can be configured to communicate with the storage medium 13 to execute a series of instruction operations in the storage medium 13 on the prediction device 100 for the super-resolution gene expression map.

[0075] The super-resolution gene expression map prediction device 100 may also include one or more power supplies 14, one or more wired or wireless network interfaces 15, one or more input and output interfaces 16, and / or one or more operating systems 131, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will appreciate that Figure 7 The device structure shown does not constitute a limitation on the super-resolution gene expression map prediction device 100 , and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0076] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of a method for predicting a super-resolution gene expression map.

[0077] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device, or unit can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0078] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.

[0079] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting super-resolution gene expression profiles, characterized in that: Includes steps: Obtain histological images and corresponding spatial transcriptome data; Divide the histological image into large image blocks according to the size of 4096*4096 pixels, and then divide the large image blocks into small image blocks according to the size of 256*256 pixels; Constructing a hierarchical feature extraction model, wherein the hierarchical feature extraction model includes a local ViT module, a global ViT module, and a feature stacking module; Assume that the height and width of the histological image X are M and N respectively, and the large image block after division is X mn , then the histological image is represented by a large image block as The large image block X mn Divide into small image blocks X mnpq , then the large image block is represented as The histological image is represented by small image blocks: The small image block is divided into 256 sub-image blocks of size 16*16. After the sub-image blocks are encoded and embedded into a two-dimensional matrix, a random vector Token is added as the input of the local ViT module. The Token corresponding to each sub-image block is obtained as a low-level local feature vector, and the CLS-Token representing the high-level local feature vector of the small image block is obtained. Merge the CLS-Tokens of all small image blocks contained in the large image block into a matrix as the input of the global ViT module, and obtain a high-level global feature vector that integrates the information of all small image blocks; Forming a low-level local feature map and a high-level global feature map based on the low-level local feature vector and the high-level global feature vector respectively; After the low-level local feature map, the high-level global feature map and the original histological image are aligned in size, they are input into the feature stacking module for stacking in the channel dimension, and a hierarchical histological feature map is output; Using the stratified histological feature map as input data and the corresponding spatial transcriptome data as pseudo-labels, a super-resolution gene expression prediction model based on the GCN graph neural network is trained through weakly supervised learning to obtain a trained super-resolution gene expression prediction model; Extracting features from the histological image to be predicted based on the hierarchical feature extraction model to obtain a hierarchical histological feature map to be predicted; Each superpixel point in the layered histological feature map to be predicted is used as a node to divide it into a series of sub-maps, and the series of sub-maps are input into the trained super-resolution gene expression prediction model to output a predicted super-resolution gene expression map.

2. The method for predicting super-resolution gene expression profiles according to claim 1, characterized in that: The formula for obtaining the low-level local feature vector is: Among them, f0 represents the component of feature information extracted in the local ViT module, The dimension of the feature vector is c0; the formula for obtaining the high-level local feature vector is: Among them, f1 represents the part of the local ViT module that generates CLS-Token by fusing feature information. The dimension of the feature vector is c1; the expression formula of the high-level global feature vector is: Among them, f2 represents the global ViT module.

3. The method for predicting super-resolution gene expression profiles according to claim 1, characterized in that: After the low-level local feature map, the high-level global feature map and the original histological image are aligned in size, the steps of inputting them into the feature stacking module for stacking in the channel dimension and outputting the hierarchical histological feature map include: Assuming that the width and height of the super-resolution gene expression map to be predicted are M′×N′, if the size of the low-level local feature map, the high-level global feature map and the original histological image is smaller than M′×N′, the image is reduced to M′×N′ by averaging; if the size is larger than M′×N′, the image is expanded to M′×N′ by broadcasting mechanism; The aligned low-level local feature map, high-level global feature map and original histological image are input into the feature stacking module for stacking in the channel dimension to obtain a hierarchical histological feature map. Among them, h mn is the hierarchical histological feature vector at the superpixel point (m,n).

4. The method for predicting super-resolution gene expression profiles according to claim 1, characterized in that: The steps of using the stratified histological feature map as input data and the corresponding spatial transcriptome data as pseudo labels to train the super-resolution gene expression prediction model based on the GCN graph neural network through weakly supervised learning to obtain the trained super-resolution gene expression prediction model include: A super-resolution gene expression prediction model based on a GCN graph neural network is constructed. The super-resolution gene expression prediction model includes a two-layer GCN module and an output module composed of a linear layer and an exponential linear unit. The weakly supervised loss function of the super-resolution gene expression prediction model is: Among them, S is the total number of spots in the entire histological image, K is the number of predicted genes, and f k is the gene expression prediction model for predicting k genes, g ks is the expression level of gene k at spot s, is the mask of s in spot, that is, the set of superpixels covered by s in spot, h mn is the hierarchical histological feature vector at the superpixel point (m, n), and θ is the parameter of the super-resolution gene expression prediction model; The super-resolution gene expression prediction model is trained with the hierarchical histological feature map as input and its corresponding spatial transcriptomics data as labels. During the training process, the gradient of the weakly supervised loss function is calculated by the back-propagation algorithm, and the parameters of the super-resolution gene expression prediction model are updated to minimize the weakly supervised loss function to obtain the trained super-resolution gene expression prediction model.

5. The method for predicting super-resolution gene expression profiles according to claim 4, characterized in that: Each superpixel point in the layered histological feature map to be predicted is used as a node to divide it into a series of sub-maps, the series of sub-maps are input into the trained super-resolution gene expression prediction model, and the predicted super-resolution gene expression map is output, including the steps of: Each superpixel point in the layered histological feature map H to be predicted is used as a node and divided into n series of sub-maps, then H = {H1, H2, H i K,H n }, where H i represents the i-th series sub-graph; The n series of subgraphs are input into the trained super-resolution gene expression prediction model. The formula of the super-resolution gene expression prediction model is as follows: i ′=Flatten(H i ), Y i 0 =σ(AH i 'W (0) ), Y i 1 =σ(AY i W (1) ), Z i ′=ELU(Y i 1 W (2) +b) and Z i =Flatten -1 (Z i ′); Flatten means converting the shape of the series subgraph from three-dimensional to two-dimensional, H i ′ represents a two-dimensional series of subgraphs; σ represents the nonlinear activation function ReLU; W (0) , W (1) are the weight matrices of the first and second layers of GCN respectively; W (2) and b are the weight matrix and bias term of the output layer respectively; Y i 0 represents the output of the first layer of GCN, Y i 1 represents the output of the second layer GCN, A represents the adjacency matrix of superpixel points, Z i ′ represents the output of the output module, Flatten -1 Indicates converting a two-dimensional matrix into a three-dimensional matrix, Z i is the gene expression subgraph, whose number of channels is the number of genes K; The n gene expression sub-graphs are merged into a predicted super-resolution gene expression map Z = {Z1, Z2, K, Z n }.

6. A super-resolution gene expression map prediction device, characterized in that: include: An acquisition module, used to acquire histological images and corresponding spatial transcriptome data; A division module, used for dividing the histological image into large image blocks according to the size of 4096*4096 pixels, and then dividing the large image blocks into small image blocks according to the size of 256*256 pixels; A model building module, used to build a hierarchical feature extraction model, wherein the hierarchical feature extraction model includes a local ViT module, a global ViT module and a feature stacking module; The sample image feature extraction module is used to set the height and width of the histological image X to M and N respectively, and the large image block after division is X mn , then the histological image is represented by a large image block as The large image block X mn Divide into small image blocks X mnpq , then the large image block is represented as The histological image is represented by small image blocks: The small image block is divided into 256 sub-image blocks of size 16*16. After the sub-image blocks are encoded and embedded to form a two-dimensional matrix, a random vector Token is added as the input of the local ViT module, and the Token corresponding to each sub-image block is obtained as a low-level local feature vector, and a CLS-Token representing the high-level local feature vector of the small image block; the CLS-Token of all small image blocks contained in the large image block is merged into a matrix as the input of the global ViT module to obtain a high-level global feature vector that integrates the information of all small image blocks; based on the low-level local feature vector and the high-level global feature vector, a low-level local feature map and a high-level global feature map are respectively formed; after the low-level local feature map, the high-level global feature map and the original histological image are size-aligned, they are input into the feature stacking module for stacking in the channel dimension, and a hierarchical histological feature map is output; A training module, for training a super-resolution gene expression prediction model based on a GCN graph neural network by using the stratified histological feature map as input data and the corresponding spatial transcriptome data as pseudo-labels through weakly supervised learning to obtain a trained super-resolution gene expression prediction model; A feature extraction module for the image to be tested, used for extracting features of the histological image to be predicted based on the hierarchical feature extraction model to obtain a hierarchical histological feature map to be predicted; The prediction module is used to divide each superpixel point in the layered histological feature map to be predicted into a series of sub-maps as a node, input the series of sub-maps into the trained super-resolution gene expression prediction model, and output a predicted super-resolution gene expression map.

7. A super-resolution gene expression map prediction device, characterized in that: comprising a memory and at least one processor, wherein the memory has computer-readable instructions stored therein; The at least one processor calls the computer-readable instructions in the memory to execute each step of the method for predicting a super-resolution gene expression map according to any one of claims 1 to 5.

8. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the method for predicting a super-resolution gene expression map according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Domain self-adaptive equipment operation inspection system and method

    CN112183788A

  • System and method for image preprocessing

    CN114787876A