Spatial transcriptomic gene expression prediction method
By segmenting the stained pathological image data of the target tissue into images and extracting cross-modal features, and combining multimodal features and spatial neighborhood aggregation networks, the problems of high cost and poor prediction accuracy in existing technologies are solved, and efficient spatial transcriptome gene expression prediction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2026-04-17
- Publication Date
- 2026-06-26
AI Technical Summary
Existing spatial transcriptome sequencing technologies are costly, have complex experimental procedures, limited throughput, and poor accuracy in gene expression profiling prediction methods based on local image patches.
By acquiring stained pathological image data of the target tissue, the images are segmented, and the cross-modal image features of the image blocks and the second cross-modal image features of adjacent image blocks are determined. Feature extraction and alignment are performed using a multimodal feature extraction network and a dual alignment contrast learning network. Gene expression and cell type prediction are then performed by combining a spatial neighborhood aggregation network to generate a spatial transcriptome gene expression prediction map.
It improves the prediction accuracy of gene expression profiles, enhances the richness of feature information based on stained pathological image data, and improves the prediction accuracy of spatial gene expression by leveraging the dependence of neighboring cell interactions.
Smart Images

Figure CN122290706A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a spatial transcriptome gene expression prediction method. Background Technology
[0002] Spatial gene expression refers to the technique of measuring gene transcriptional activity at specific spatial locations in tissue sections. It is of great significance for understanding cell behavior, intercellular interactions, and the molecular mechanisms of disease development in the tissue microenvironment.
[0003] Spatial transcriptome sequencing technology faces practical challenges such as high cost, complex experimental procedures, and limited throughput, hindering its widespread application in large-scale clinical studies and routine pathological diagnosis. Currently, the main approach employs prediction methods based on local image patches, utilizing convolutional neural networks or visual transformers to extract image features from histopathological images, and then using regression models to predict gene expression profiles at corresponding locations. However, this method relies solely on image feature information, resulting in poor accuracy in predicting gene expression profiles. Summary of the Invention
[0004] Therefore, it is necessary to provide a spatial transcriptome gene expression prediction method that can improve the prediction accuracy of gene expression profiles, addressing the aforementioned technical problems.
[0005] Firstly, this application provides a spatial transcriptome gene expression prediction method, including:
[0006] Acquire stained pathological image data of the target tissue;
[0007] The stained pathological image data is divided into image blocks to obtain at least two image blocks to be identified.
[0008] For each image block, determine the first cross-modal image features of the image block and the second cross-modal image features of the adjacent image blocks of the image block;
[0009] Based on the aggregation results of the first cross-modal image features and the second cross-modal image features, image patches are predicted to determine the gene expression prediction results and cell type prediction results corresponding to the image patches;
[0010] Based on the gene expression prediction results and cell type prediction results of each image patch, a spatial transcriptome gene expression prediction map of the target tissue is generated.
[0011] In one embodiment, for each image block, determining a first cross-modal image feature of the image block and a second cross-modal image feature of adjacent image blocks includes:
[0012] The multimodal feature extraction network of the target prediction model is used to extract features from each image block to be identified, thereby obtaining the image block features of each image block to be identified.
[0013] By using a dual-alignment contrastive learning network of the target prediction model, image patch features are processed to obtain cross-modal image features;
[0014] From the cross-modal image features, determine the first cross-modal image features of the image patch and the second cross-modal image features of the adjacent image patches of the image patch.
[0015] In one embodiment, image patch features are processed to obtain cross-modal image features, including:
[0016] Align the image patch features to obtain aligned image patch features;
[0017] Cross-attention processing is applied to the aligned image patch features to obtain attention features;
[0018] Attention features and aligned image patch features are fused to obtain cross-modal image features of image patch features.
[0019] In one embodiment, based on the aggregation result of the first cross-modal image features and the second cross-modal image features, image patches are predicted to determine the gene expression prediction result and cell type prediction result corresponding to the image patches, including:
[0020] The spatial neighborhood aggregation network of the target prediction model is used to aggregate the first cross-modal image features and the second cross-modal image features to obtain the aggregated features of the image patch.
[0021] The aggregated features of the image patch are input into the prediction network of the target prediction model to obtain the gene expression prediction results and cell type prediction results corresponding to the image patch.
[0022] In one embodiment, the method further includes:
[0023] Acquire pathological image data, spatial transcriptome gene data and cell type abundance data of sample tissues. The pathological image data includes at least two pathological image blocks obtained by dividing the stained pathological images of sample tissues into blocks according to specified specifications.
[0024] Each pathological image patch, spatial transcriptome gene data, and cell type abundance data are input into the prediction model to be trained, which includes a multimodal feature extraction network and a dual alignment contrast learning network.
[0025] By using a multimodal feature extraction network, features are extracted from each pathological image block, spatial transcriptome gene data, and cell type abundance data to obtain the pathological image features, gene expression features, and cell type abundance features of each pathological image block.
[0026] By using a dual-alignment contrastive learning network, pathological image features, gene expression features, and cell type abundance features are mapped to a shared mapping space to obtain alignment loss. Based on the alignment loss, the parameters of the multimodal feature extraction network and the dual-alignment contrastive learning network are updated until the alignment loss converges, thus obtaining cross-modal features.
[0027] In one embodiment, pathological image data, spatial transcriptome gene data, and cell type annotation data of the sample tissue are acquired, including:
[0028] Acquire raw spatial transcriptome data, stained pathological images, and cell type description data of the sample tissue;
[0029] The raw spatial transcriptome data was screened to obtain cell type-specific marker genes, and the gene expression profiles of the cell type-specific marker genes were identified as spatial transcriptome gene data.
[0030] The stained pathological image is divided into blocks according to a specified size to obtain at least two pathological image blocks, which are then identified as pathological image data.
[0031] The cell type description data was converted into cell type description text and identified as cell type abundance data.
[0032] In one embodiment, raw spatial transcriptome data is screened to obtain cell type-specific marker genes, including:
[0033] Genes related to intercellular communication were screened from raw spatial transcriptome data;
[0034] Determine the degree of expression variation in genes related to intercellular communication;
[0035] Based on the degree of expression variation, identify hypervariable genes among genes related to intercellular communication;
[0036] Cell type-specific marker genes were screened from hypervariable genes.
[0037] In one embodiment, the cross-modal image features are obtained until the alignment loss converges, including:
[0038] When the alignment loss converges, cross-attention processing is performed using pathological image features as query indicators, gene expression features as keys, and cell type abundance features as values to obtain attention output features.
[0039] The attention output features and pathological image features are fused to obtain cross-modal image features of pathological image blocks.
[0040] In one embodiment, the prediction model to be trained further includes a spatial neighborhood aggregation network and a prediction network, and after the alignment loss converges and cross-modal features are obtained, it further includes:
[0041] The spatial neighborhood aggregation network is used to aggregate the cross-modal features of each pathological image block and its neighboring image blocks to obtain the neighborhood aggregation features of each pathological image block.
[0042] The neighborhood aggregation features of each pathological image patch are input into the prediction network to obtain gene expression prediction samples and cell type prediction samples corresponding to each pathological image patch.
[0043] The total target loss is determined based on the differences between spatial transcriptome gene data and gene expression prediction samples, the differences between cell type abundance data and cell type prediction samples, and alignment loss.
[0044] Based on the total target loss, the parameters of the spatial neighborhood aggregation module and the prediction network are updated until the total target loss converges. The multimodal feature extraction network, the dual alignment contrast learning network, the spatial neighborhood aggregation network, and the prediction network after the total target loss converges are determined as the target prediction model.
[0045] In one embodiment, the target total loss is determined based on the differences between spatial transcriptome gene data and gene expression prediction samples, the differences between cell type abundance data and cell type prediction samples, and alignment loss, including:
[0046] Gene expression loss was determined based on the difference between spatial transcriptome gene data and gene expression prediction samples;
[0047] The cell type loss for each cell type is determined based on the difference between cell type abundance data and cell type prediction samples.
[0048] Determine the rarity weight of each cell type; the rarity weight is positively correlated with the scarcity of the cell type.
[0049] The cell type loss is calculated by weighting the cell type loss according to the rarity weight of each cell type, and the overall cell type loss is obtained.
[0050] The total target loss is determined based on gene expression loss, cell type integrated loss, and alignment loss.
[0051] Secondly, this application also provides a spatial transcriptome gene expression prediction device, comprising:
[0052] The data acquisition module is used to acquire stained pathological image data of the target tissue;
[0053] The image segmentation module is used to segment stained pathological image data into images to obtain at least two image blocks to be identified.
[0054] The feature determination module is used to determine, for each image block, the first cross-modal image features of the image block and the second cross-modal image features of the adjacent image blocks of the image block;
[0055] The gene prediction module is used to predict image patches based on the aggregation results of the first cross-modal image features and the second cross-modal image features, and to determine the gene expression prediction results and cell type prediction results corresponding to the image patches.
[0056] The map combination module is used to generate a spatial transcriptome gene expression prediction map of the target tissue based on the gene expression prediction results and cell type prediction results of each image patch.
[0057] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0058] Acquire stained pathological image data of the target tissue;
[0059] The stained pathological image data is divided into image blocks to obtain at least two image blocks to be identified.
[0060] For each image block, determine the first cross-modal image features of the image block and the second cross-modal image features of the adjacent image blocks of the image block;
[0061] Based on the aggregation results of the first cross-modal image features and the second cross-modal image features, image patches are predicted to determine the gene expression prediction results and cell type prediction results corresponding to the image patches;
[0062] Based on the gene expression prediction results and cell type prediction results of each image patch, a spatial transcriptome gene expression prediction map of the target tissue is generated.
[0063] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0064] Acquire stained pathological image data of the target tissue;
[0065] The stained pathological image data is divided into image blocks to obtain at least two image blocks to be identified.
[0066] For each image block, determine the first cross-modal image features of the image block and the second cross-modal image features of the adjacent image blocks of the image block;
[0067] Based on the aggregation results of the first cross-modal image features and the second cross-modal image features, image patches are predicted to determine the gene expression prediction results and cell type prediction results corresponding to the image patches;
[0068] Based on the gene expression prediction results and cell type prediction results of each image patch, a spatial transcriptome gene expression prediction map of the target tissue is generated.
[0069] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0070] Acquire stained pathological image data of the target tissue;
[0071] The stained pathological image data is divided into image blocks to obtain at least two image blocks to be identified.
[0072] For each image block, determine the first cross-modal image features of the image block and the second cross-modal image features of the adjacent image blocks of the image block;
[0073] Based on the aggregation results of the first cross-modal image features and the second cross-modal image features, image patches are predicted to determine the gene expression prediction results and cell type prediction results corresponding to the image patches;
[0074] Based on the gene expression prediction results and cell type prediction results of each image patch, a spatial transcriptome gene expression prediction map of the target tissue is generated.
[0075] The aforementioned spatial transcriptome gene expression prediction method acquires stained pathological image data of the target tissue; divides the stained pathological image data into image blocks to obtain at least two image blocks to be identified; for each image block, it determines the first cross-modal image features of the image block and the second cross-modal image features of its neighboring image blocks. Therefore, this application constructs cross-modal image features for each image block based on the intrinsic correlation between images and other modalities, and determines the first cross-modal image features of the image block and the second cross-modal image features of its neighboring image blocks, which helps to improve the information richness of features obtained from stained pathological image data. Since neighboring cell interactions affect gene expression, this application can predict image blocks based on the aggregation results of the first and second cross-modal image features, determining the corresponding gene expression prediction results and cell type prediction results for each image block. Furthermore, based on the gene expression prediction results and cell type prediction results of each image block, a spatial transcriptome gene expression prediction map of the target tissue is generated. Therefore, this application, based on the highly information-rich first and second cross-modal image features, aggregates them using the dependencies between neighboring tissues before making predictions, which helps improve the accuracy of spatial gene expression prediction. The derivation of the technical features in the proprietary patent can achieve the beneficial effects addressing the technical problems in the background art. Attached Figure Description
[0076] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0077] Figure 1 This is a diagram illustrating the application environment of a spatial transcriptome gene expression prediction method in one embodiment;
[0078] Figure 2 This is a flowchart illustrating a spatial transcriptome gene expression prediction method in one embodiment;
[0079] Figure 3 This is a flowchart illustrating the spatial transcriptome gene expression prediction method in yet another embodiment;
[0080] Figure 4 This is a flowchart illustrating a spatial transcriptome gene expression prediction method in another embodiment;
[0081] Figure 5 This is a schematic diagram illustrating an application scenario of a specific embodiment of this application;
[0082] Figure 6 This is a structural block diagram of a spatial transcriptome gene expression prediction device in one embodiment;
[0083] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0084] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0085] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0086] The spatial transcriptome gene expression prediction method provided in this application can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on a cloud or other network server. The spatial transcriptome gene expression prediction method can be executed by terminal 102 or server 104. For example, terminal 102 acquires stained pathological image data of the target tissue and sends the data to server 104. Server 104 segments the stained pathological image data into image blocks, obtaining at least two image blocks to be identified. For each image block, it determines the first cross-modal image features of the image block and the second cross-modal image features of adjacent image blocks. Based on the aggregation result of the first and second cross-modal image features, it predicts the image block, determining the gene expression prediction result and cell type prediction result corresponding to the image block. Based on the gene expression prediction result and cell type prediction result of each image block, it generates a spatial transcriptome gene expression prediction map of the target tissue. Server 104 feeds back the spatial transcriptome gene expression prediction map of the target tissue to terminal 102. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, etc. IoT devices can include smart TVs, smart in-vehicle devices, projection devices, etc. Server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0087] In one exemplary embodiment, such as Figure 2 As shown, a spatial transcriptome gene expression prediction method is provided, which can be applied to... Figure 1 Taking the server in the example, the explanation includes the following steps S110 to S150. Wherein:
[0088] Step S110: Obtain stained pathological image data of the target tissue;
[0089] The target tissue is the tissue sample for which prediction is expected, and the staining pathological image data is an image of the target tissue that has been stained and whose digital pathological information has been preserved. For example, the staining pathological image data may include a digital pathological image of the target tissue after H&E (Hematoxylin-Eosin staining).
[0090] For example, this embodiment can receive stained pathological images of the target tissue transmitted from other devices or stored in a predetermined storage space as stained pathological image data of the target tissue.
[0091] Step S120: Divide the stained pathological image data into image blocks to obtain at least two image blocks to be identified;
[0092] For example, this embodiment can divide the stained pathological image data into image blocks according to a predetermined specification to obtain at least two image blocks to be identified. Furthermore, this embodiment can also perform standardization processing on each image block. For instance, this embodiment can divide the stained pathological image data into several image blocks, uniformly adjust the image blocks to a predetermined specification (such as 224×224×3 RGB format) and normalize them to achieve image standardization processing.
[0093] Step S130: For each image block, determine the first cross-modal image features of the image block and the second cross-modal image features of the adjacent image blocks of the image block;
[0094] The first cross-modal image feature is the cross-modal image feature corresponding to the image patch, and the second cross-modal image feature is the cross-modal image feature of the adjacent image patches of the image patch. The cross-modal image feature is an image feature that incorporates other modal features (such as gene expression features, cell type abundance features, etc.) corresponding to the image patch.
[0095] This embodiment can extract features from each image block to be identified, obtaining image block features for each image block. Then, alignment and cross-attention processing are performed on these image block features to obtain cross-modal image features. Alignment processing is an operation used to shorten the distance between image block features of the same image block and other modal features within a shared mapping space. Cross-attention processing is an operation that aggregates information related to image block features from other modal features under the guidance of the image block features. In another example, this embodiment can also use knowledge distillation to "distill" the semantic information of other modal features into the image feature encoder. After encoding each image block, the resulting image features (i.e., cross-modal image features) possess the representational ability of other modal features. Therefore, the cross-modal image features of the image block are determined as the first cross-modal image features, and the cross-modal image features of adjacent image blocks are determined as the second cross-modal image features.
[0096] In some embodiments, for each image block, determining a first cross-modal image feature of the image block and a second cross-modal image feature of adjacent image blocks includes:
[0097] Step S131: Through the multimodal feature extraction network of the target prediction model, feature extraction is performed on each image block to be identified to obtain the image block features of each image block to be identified;
[0098] Step S132: The image patch features are processed through the dual-alignment contrastive learning network of the target prediction model to obtain cross-modal image features;
[0099] Step S133: From the cross-modal image features, determine the first cross-modal image features of the image block and the second cross-modal image features of the adjacent image blocks of the image block.
[0100] The target prediction model is a model used to predict spatial gene expression maps. The multimodal feature extraction network can include feature encoders for each modality, such as image feature encoders, text feature encoders, and gene expression feature encoders. The dual-alignment contrastive learning network incorporates information from other modalities (text modality and expression modality) into image patch features during the training phase, so that it can output cross-modal image features with cross-modal information using only image patch features during the inference phase.
[0101] This embodiment can extract features from each image patch to be identified using a multimodal feature extraction network of the target prediction model, obtaining image patch features for each image patch. Then, through a dual-alignment contrastive learning network of the target prediction model, alignment and cross-attention processing are performed to obtain cross-modal image features. Alignment processing is an operation used to shorten the distance between image patch features and other modal features within the shared mapping space. Cross-attention processing is an operation that aggregates information related to image patch features from other modal features, guided by the image patch features. Therefore, this embodiment can determine the first cross-modal image features of an image patch and the second cross-modal image features of its neighboring image patches from the cross-modal image features. Neighboring image patches are other image patches that are connected to the edge of the image patch. For example, if the image patch has six image patches connected to its edge, then the first cross-modal image features include the cross-modal image features of the image patch, and the second cross-modal image features include the cross-modal image features of these six neighboring image patches.
[0102] In some embodiments, image patch features are processed to obtain cross-modal image features, including:
[0103] Step S210: Align the image patch features to obtain aligned image patch features;
[0104] Step S220: Perform cross-attention processing on the aligned image patch features to obtain attention features;
[0105] Step S230: The attention features and aligned image patch features are fused to obtain cross-modal image features of the image patch features.
[0106] The dual-alignment contrastive learning network includes a mapping layer, an attention mechanism layer, and a fusion layer. The mapping layer is used to map the features extracted by the multimodal feature extraction network to a shared mapping space. The attention mechanism layer is used to perform attention cross-processing on the aligned features of each modality in the shared mapping space. The fusion layer is used to fuse the aligned features (including at least image features) with the attention features output by the attention mechanism layer.
[0107] This embodiment uses the mapping layer of a dual-alignment contrastive learning network to align image patch features, mapping them to aligned image patch features. Then, a trained attention mechanism layer performs attention cross-processing on these aligned image patch features, outputting attention features related to the image patch features and containing other modal information. A fusion layer is used to fuse the aligned image patch features with the attention features output by the attention mechanism layer, obtaining cross-modal image features of the image patch features. Exemplarily, the fusion method can include stitching, tensor fusion, projection fusion, element fusion, etc.
[0108] This embodiment combines alignment and attention mechanisms to generate attention features with other modal information under the guidance of image patch features. The attention features are then fused with the aligned image patch features, enabling this embodiment to obtain cross-modal image features even with only an input image. This helps to improve the information richness of features obtained from stained pathological image data, thereby improving the accuracy of subsequent predictions of spatial gene expression.
[0109] Step S140: Based on the aggregation results of the first cross-modal image features and the second cross-modal image features, the image patch is predicted to determine the gene expression prediction result and cell type prediction result corresponding to the image patch;
[0110] Since gene expression in the real tissue microenvironment is influenced by the interactions of neighboring cells, this embodiment can aggregate the first and second cross-modal image features. The resulting aggregation not only characterizes the features of the image patch itself but also its features in long-range space. For example, this embodiment can sequentially concatenate the first and second cross-modal image features into a feature sequence according to the spatial positions of the image patch and its neighboring patches. Then, this feature sequence is used for sequence modeling (such as a Mamba model based on a selective state-space mechanism), enabling the feature representation of each image patch in the feature sequence to incorporate contextual information from neighboring image patches, thereby capturing spatial dependencies and obtaining a feature vector modeling sequence. Each vector in this feature vector modeling sequence already contains information resulting from spatial interactions. Global average pooling is then performed on each vector in the feature vector modeling sequence to obtain an aggregation result containing spatial dependencies. Furthermore, this embodiment can predict the image patch, determining the predicted gene expression and cell type corresponding to the image patch.
[0111] In some embodiments, based on the aggregation results of the first cross-modal image features and the second cross-modal image features, image patches are predicted to determine the gene expression prediction results and cell type prediction results corresponding to the image patches, including:
[0112] Step S141: Through the spatial neighborhood aggregation network of the target prediction model, the first cross-modal image features and the second cross-modal image features are aggregated to obtain the aggregated features of the image patch.
[0113] Step S142: Input the aggregated features of the image patch into the prediction network of the target prediction model to obtain the gene expression prediction results and cell type prediction results corresponding to the image patch.
[0114] The spatial neighborhood aggregation network is used to capture long-range spatial dependencies between cross-modal features of an image patch and its neighboring image patches, and to aggregate cross-modal features containing long-range spatial dependencies. For example, the spatial neighborhood aggregation network may include a sequence modeling model for capturing long-range spatial dependencies and a pooling layer for aggregation. This sequence modeling model may include models such as the Mamba model (a sequence modeling model based on a selective state space mechanism) or the Transformer model with an attention mechanism, which can be used to capture long-range spatial dependencies. The prediction network can be a multilayer perceptron (MLP) regression head (composed of several fully connected layers, nonlinear activation functions, normalization layers, and an output head), mapping the aggregated features to the target dimension to obtain the output result. This embodiment can also use a residual multilayer perceptron to improve the stability of deep training, or a gated multilayer perceptron to enhance nonlinear representation capabilities. It is understood that when multiple tasks need to be predicted simultaneously, the output head can also adopt a shared backbone + multi-branch output head structure to achieve joint prediction of multiple tasks such as gene expression and cell type.
[0115] This embodiment can aggregate the first and second cross-modal image features using a spatial neighborhood aggregation network of the target prediction model to obtain aggregated features of the image patch. For example, this embodiment can sequentially concatenate the first and second cross-modal image features in spatial neighborhood order (e.g., "top → bottom → left → right → top left → bottom right → center", or clockwise / counterclockwise order) to form a feature sequence. This feature sequence is then input into a sequence modeling model (e.g., a Mamba model based on a selective state-space mechanism) for sequence modeling, enabling the feature vectors corresponding to each image patch in the feature sequence to incorporate contextual information from neighboring patches, thereby capturing internal long-range spatial dependencies. The output of the sequence modeling model is still a feature vector modeling sequence with the same dimension as the feature sequence, but each feature vector in this sequence already contains information after spatial interaction (i.e., long-range spatial dependencies). Then, each feature vector in the feature vector modeling sequence is pooled (e.g., global average pooling) to obtain the aggregation result containing spatial dependencies, i.e., the aggregated features of the image patch. The aggregated features of the image patch are then input into the prediction network of the target prediction model to obtain the gene expression prediction results and cell type prediction results corresponding to the image patch.
[0116] Step S150: Based on the gene expression prediction results and cell type prediction results of each image patch, a spatial transcriptome gene expression prediction map of the target tissue is generated.
[0117] This embodiment can sequentially combine the gene expression prediction results and cell type prediction results of each image patch according to their spatial location in the target tissue to obtain a spatial transcriptome gene expression prediction map of the target tissue. For example, the spatial transcriptome gene expression prediction map includes at least a gene expression matrix, a spatial coordinate matrix corresponding to the gene expression matrix, and the spatial distribution of cell types. The gene expression matrix is a "location × gene" matrix, where each row represents a captured spatial location (i.e., an image patch), each column represents a gene, and the values in the matrix represent the expression level of that gene at that spatial location (e.g., mRNA count). The spatial coordinate matrix corresponds one-to-one with the expression matrix, recording the two-dimensional coordinates of each spatial location. The spatial distribution of cell types represents the precise location and abundance of different cell types (e.g., T cells, B cells, fibroblasts, tumor cell subtypes) in the target tissue.
[0118] In the aforementioned spatial transcriptome gene expression prediction method, stained pathological image data of the target tissue is acquired; the stained pathological image data is divided into image blocks to obtain at least two image blocks to be identified; for each image block, the first cross-modal image features of the image block and the second cross-modal image features of its adjacent image blocks are determined. Therefore, this application constructs cross-modal image features for each image block based on the intrinsic correlation between images and other modalities, and determines the first cross-modal image features of the image block and the second cross-modal image features of its adjacent image blocks, which helps to improve the information richness of features obtained from stained pathological image data. Since neighboring cell interactions affect gene expression, this application can predict image blocks based on the aggregation results of the first and second cross-modal image features, determining the corresponding gene expression prediction results and cell type prediction results for each image block. Furthermore, based on the gene expression prediction results and cell type prediction results of each image block, a spatial transcriptome gene expression prediction map of the target tissue is generated. Therefore, this application, based on the first cross-modal image features and the second cross-modal image features which have a high degree of information richness, aggregates them by leveraging the dependencies between neighboring tissues before making predictions, which helps to improve the prediction accuracy of spatial gene expression.
[0119] In one exemplary embodiment, such as Figure 3 As shown, the method further includes steps S310 to S340. Wherein:
[0120] Step S310: Obtain pathological image data, spatial transcriptome gene data and cell type abundance data of the sample tissue. The pathological image data includes at least two pathological image blocks obtained by dividing the stained pathological image of the sample tissue into blocks according to a specified specification.
[0121] Step S320: Input each pathological image block, spatial transcriptome gene data, and cell type abundance data into the prediction model to be trained, wherein the prediction model to be trained includes a multimodal feature extraction network and a dual alignment contrast learning network.
[0122] Step S330: Through a multimodal feature extraction network, feature extraction is performed on each pathological image block, spatial transcriptome gene data, and cell type abundance data to obtain the pathological image features, gene expression features, and cell type abundance features of each pathological image block.
[0123] Step S340: The pathological image features, gene expression features, and cell type abundance features are mapped to a shared mapping space through a dual alignment contrast learning network to obtain the alignment loss. The parameters of the multimodal feature extraction network and the dual alignment contrast learning network are updated based on the alignment loss until the alignment loss converges to obtain cross-modal features.
[0124] The multimodal feature extraction network can include feature encoders for various modalities such as images, gene expression, and text. For example, a multimodal feature extraction network may include an image feature encoder, a text feature encoder, and a gene expression feature encoder. The shared mapping space is a semantic space shared by features from different modalities.
[0125] This embodiment can acquire a sample dataset of tissue samples, which may include pathological image data, spatial transcriptome gene data, and cell type abundance data. The pathological image data includes at least two pathological image blocks obtained by dividing stained pathological images of the tissue samples into blocks according to specified specifications. Each pathological image block, spatial transcriptome gene data, and cell type abundance data are input into a prediction model to be trained. The prediction model to be trained includes a multimodal feature extraction network and a dual-alignment contrastive learning network. Through the multimodal feature extraction network, for each pathological image block, feature extraction is performed using an image feature encoder, feature extraction using a gene expression encoder, and feature extraction using a text feature encoder, resulting in pathological image features, gene expression features, and cell type abundance features for each pathological image block.
[0126] Furthermore, this embodiment uses a dual-alignment contrastive learning network to map pathological image features, gene expression features, and cell type abundance features to a shared mapping space. Then, based on the first distance between the pathological image features of the same image patch and their corresponding gene expression features and cell type abundance features within the shared mapping space, and the second distance between the pathological image features of an image patch and their gene expression features and cell type abundance features of different image patches, the alignment loss is determined. For example, this embodiment can calculate the first contrastive loss by using the first expression distance between the pathological image features of an image patch and their corresponding gene expression features of the same image patch, and the second expression distance between the pathological image features of an image patch and their gene expression features of different image patches. The second contrastive loss is calculated by using the first abundance distance between the pathological image features of an image patch and their corresponding cell type abundance features of the same image patch, and the second abundance distance between the pathological image features of an image patch and their cell type abundance features of different image patches. Furthermore, this embodiment can weight the first contrastive loss and the second contrastive loss to obtain the alignment loss. Furthermore, this embodiment updates the parameters of the multimodal feature extraction network and the dual-alignment contrast learning network based on alignment loss. This brings the pathological image features of an image patch closer to the corresponding gene expression features and cell type abundance features of the same image patch, while widening the pathological image features of an image patch to the gene expression features and cell type abundance features of different image patches, until the alignment loss converges. For example, the first and second contrast losses can be represented using the InfoNCE (Information Noise-Contrastive Estimation) loss function.
[0127] Then, based on the double-alignment contrastive learning network under the alignment loss convergence state, pathological image features, gene expression features, and cell type abundance features are mapped to pathological image features, gene expression features, and cell type abundance features under the alignment state. These features are then fused into cross-modal features. For example, in this embodiment, under the alignment loss convergence state, the pathological image features under the alignment state can be used as the query index, the gene expression features under the alignment state as the key, and the cell type abundance features under the alignment state as the value for cross-attention processing to obtain attention output features. These attention output features are then fused with the pathological image features under the alignment state to obtain the cross-modal image features of the pathological image patch. Further, this embodiment can also fuse the attention output features with the pathological image features, gene expression features, and cell type abundance features under the alignment state to obtain the cross-modal image features of the pathological image patch.
[0128] This embodiment proposes a dual-alignment contrastive learning network, which enables the establishment of intrinsic relationships between images, gene expression, and cell type composition while aligning multimodal features. This allows the network parameters of the trained dual-alignment contrastive learning network to output image features that also have the ability to represent other modal features such as gene expression and cell type abundance (i.e., cross-modal image features). This improves the feature richness of the image patch to be identified in subsequent use, thereby improving prediction accuracy.
[0129] In some embodiments, acquiring pathological image data, spatial transcriptome gene data, and cell type annotation data of sample tissue includes:
[0130] Step S410: Obtain raw spatial transcriptome data, stained pathological images, and cell type description data of the sample tissue;
[0131] Step S420: Screen the raw spatial transcriptome data to obtain cell type-specific marker genes, and determine the gene expression profile of the cell type-specific marker genes as spatial transcriptome gene data;
[0132] Step S430: Divide the stained pathological image into blocks according to the specified specifications to obtain at least two pathological image blocks, and determine them as pathological image data;
[0133] Step S440: Convert cell type description data into cell type description text and identify it as cell type abundance data.
[0134] The raw spatial transcriptome data includes raw gene expression information obtained from spatial transcriptome sequencing of the sample tissue. The stained pathological images are images of the sample tissue stained and preserving complete digital pathological information. For example, these stained pathological images may include digital pathological images of the sample tissue treated with H&E (Hematoxylin-Eosin staining). Cell type-specific marker genes are specific genes that can be used to identify cell types. Cell type description data is the cell type annotation data for the stained pathological images.
[0135] To improve data quality, this embodiment performs data preprocessing before model training using pathological image data, spatial transcriptome gene data, and cell type abundance data from sample tissues. This embodiment acquires raw spatial transcriptome data, stained pathological images, and cell type description data from sample tissues. The raw spatial transcriptome data is then screened to obtain cell type-specific marker genes that can be used to identify cell types. The gene expression profiles of these cell type-specific marker genes are then identified as spatial transcriptome gene data. The stained pathological images are then divided into at least two pathological image blocks according to a specified specification, which are identified as pathological image data. The specified specification is consistent with the model input requirements. The cell type description data is converted into cell type description text, which is identified as cell type abundance data. For example, this embodiment can use the CellChatDB database (intercellular communication database) to screen for intercellular communication-related genes in the raw spatial transcriptome data. Hypervariable genes are then screened from these intercellular communication-related genes. Cell type-specific marker genes are then screened from the hypervariable genes using the Cell2location method. The screened cell type-specific marker genes are then standardized to obtain spatial transcriptome gene data. After dividing the stained pathological image into several image blocks, each image block is uniformly adjusted to a specified size and normalized to obtain individual pathological image blocks. All pathological image blocks can be standardized using statistics (mean and standard deviation) from the ImageNet dataset to adapt to the image feature encoder. In this embodiment, cell type abundance information for each pathological image block is extracted from cell type description data containing cell type annotations, and corresponding cell type description text is generated based on the cell type abundance information. Furthermore, this embodiment can use pathological image data of sample tissues, spatial transcriptome gene data, and cell type abundance data as a sample dataset for subsequent model training and testing.
[0136] In some embodiments, raw spatial transcriptome data is screened to obtain cell type-specific marker genes, including:
[0137] Step S421: Screen for intercellular communication-related genes from the raw spatial transcriptome data;
[0138] Step S422: Determine the degree of expression variation of genes related to intercellular communication;
[0139] Step S423: Based on the degree of expression variation, identify hypervariable genes among the genes related to intercellular communication;
[0140] Step S424: Screen for cell type-specific marker genes from the hypervariable genes.
[0141] For example, this embodiment can screen cell-to-cell communication (CPC) related genes that match the CellChatDB database from the raw spatial transcriptome data, thereby determining the expression variation of these CPC related genes. Based on the expression variation, highly variable genes among the CPC related genes are identified. For example, the top N CPC related genes in terms of expression variation are identified as highly variable genes, or CPC related genes with expression variation exceeding a specified variation threshold are identified as highly variable genes. Then, based on the Cell2location method, cell type-specific marker genes are screened from the highly variable genes. The screened cell type-specific marker genes are standardized to obtain spatial transcriptome gene data. For example, log1p standardization is performed, i.e., log1p normalization is applied to gene expression values (UMI counts). 10 The logarithmic transformation of (1 + x) or ln(1 + x) can avoid infinitesimal values when taking the logarithm of genes with an expression level of 0, and at the same time can compress the variance of highly expressed genes, making the data distribution closer to a normal distribution.
[0142] In some embodiments, the cross-modal image features are obtained until the alignment loss converges, including:
[0143] Step S341: Under the condition that the alignment loss converges, cross-attention processing is performed using pathological image features as query indicators, gene expression features as keys, and cell type abundance features as values to obtain attention output features.
[0144] Step S342: The attention output features and pathological image features are fused to obtain the cross-modal image features of the pathological image block.
[0145] This embodiment, upon convergence of the alignment loss, performs cross-attention processing using aligned pathological image features as query indicators, aligned gene expression features as keys, and aligned cell type abundance features as values to obtain attention output features. These attention output features are then fused with the pathological image features to obtain cross-modal image features of the pathological image patch. This embodiment achieves fine-grained cross-modal interaction through a cross-attention mechanism, thereby establishing an intrinsic correlation between image, gene expression, and cell type, enabling the attention features to characterize information related to pathological image features in gene expression features and cell type abundance features.
[0146] In one exemplary embodiment, such as Figure 4 As shown, the prediction model to be trained also includes a spatial neighborhood aggregation network and a prediction network. After the alignment loss converges and cross-modal features are obtained, steps S510 to S540 are also included. Wherein:
[0147] Step S510: Through the spatial neighborhood aggregation network, the cross-modal features of each pathological image block and the adjacent image blocks of each pathological image block are aggregated to obtain the neighborhood aggregation features of each pathological image block.
[0148] Step S520: Input the neighborhood aggregation features of each pathological image block into the prediction network to obtain gene expression prediction samples and cell type prediction samples corresponding to each pathological image block.
[0149] Step S530: Determine the total target loss based on the differences between spatial transcriptome gene data and gene expression prediction samples, the differences between cell type abundance data and cell type prediction samples, and alignment loss.
[0150] Step S540: Based on the total target loss, update the parameters of the spatial neighborhood aggregation module and the prediction network until the total target loss converges. Then, determine the multimodal feature extraction network, dual alignment contrast learning network, spatial neighborhood aggregation network and prediction network after the total target loss converges as the target prediction model.
[0151] The training prediction model includes a spatial neighborhood aggregation network and a prediction network. The spatial neighborhood aggregation network is used to capture long-range spatial dependencies between cross-modal features of an image patch and its neighboring image patches, and to aggregate cross-modal features containing long-range spatial dependencies. For example, the spatial neighborhood aggregation network may include a sequence modeling model for capturing long-range spatial dependencies and a pooling layer for aggregation. The sequence modeling model may include models such as the Mamba model (a sequence modeling model based on a selective state-space mechanism) or a Transformer model with an attention mechanism, which can be used to capture long-range spatial dependencies. The prediction network may be a Multilayer Perceptron (MLP) regression head (composed of several fully connected layers, nonlinear activation functions, normalization layers, and an output head), which maps the aggregated features to the target dimension to obtain the output result. This embodiment can also use a residual MLP to improve the stability of deep training, or a gated MLP to enhance nonlinear representation capabilities. Understandably, when multiple tasks need to be predicted simultaneously, the output head can also adopt a structure of shared trunk + multi-branch output head to achieve joint prediction of multiple tasks such as gene expression and cell type.
[0152] This embodiment utilizes a spatial neighborhood aggregation network to aggregate the cross-modal features of each pathological image block and its neighboring image blocks, obtaining the neighborhood aggregation features of each pathological image block. For example, this embodiment can sequentially concatenate the cross-modal features of each pathological image block and its neighboring image blocks in a spatial neighborhood order (e.g., "top → bottom → left → right → top left → bottom right → center", or clockwise / counterclockwise order) to form a feature sequence. This feature sequence is then input into a sequence modeling model (e.g., a Mamba model based on a selective state-space mechanism) for sequence modeling. This allows the feature vector corresponding to each pathological image block in the feature sequence to incorporate contextual information from neighboring blocks, thereby capturing internal long-range spatial dependencies. The output of the sequence modeling model remains a feature vector modeling sequence with the same dimension as the feature sequence, but each feature vector in this sequence already contains information from spatial interactions (i.e., long-range spatial dependencies). Next, pooling (e.g., global average pooling) is performed on each feature vector in the feature vector modeling sequence to obtain neighborhood aggregation features containing spatial dependencies in the pathological image patches. Then, the neighborhood aggregation features of each pathological image patch are input into the prediction network to obtain gene expression prediction samples and cell type prediction samples corresponding to each pathological image patch. Based on the differences between spatial transcriptome gene data and gene expression prediction samples, the differences between cell type abundance data and cell type prediction samples, and the alignment loss, the target total loss is determined. Since the multimodal feature extraction network and the dual alignment contrast learning network have been trained to convergence of the alignment loss at this point, this embodiment can update the parameters of the spatial neighborhood aggregation module and the prediction network based on the target total loss until the target total loss converges. The multimodal feature extraction network, dual alignment contrast learning network, spatial neighborhood aggregation network, and prediction network after the target total loss converges are then determined as the target prediction model.
[0153] In some embodiments, the target total loss is determined based on the differences between spatial transcriptome gene data and gene expression prediction samples, the differences between cell type abundance data and cell type prediction samples, and alignment loss, including:
[0154] Step S531: Determine gene expression loss based on the difference between spatial transcriptome gene data and gene expression prediction samples;
[0155] Step S532: Determine the cell type loss for each cell type based on the difference between the cell type abundance data and the cell type prediction samples.
[0156] Step S533: Determine the rarity weight of each cell type. The rarity weight is positively correlated with the rarity of the cell type.
[0157] Step S534: The cell type loss is calculated by weighting the cell type loss according to the rarity weight of each cell type to obtain the comprehensive cell type loss;
[0158] Step S535: Determine the total target loss based on gene expression loss, cell type integrated loss, and alignment loss.
[0159] In real tissues, the majority of major cell types (such as epithelial cells) far outnumber rarer but functionally critical cell types (such as immune cells and stromal cells). However, since local gene expression is determined by the mixed cell composition, this cell population imbalance directly propagates to the prediction results. Therefore, this embodiment determines gene expression loss based on the difference between spatial transcriptome gene data and gene expression prediction samples, determines cell type loss for each cell type based on the difference between cell type abundance data and cell type prediction samples, and determines the rarity weight of each cell type. The rarity weight is positively correlated with the scarcity of the cell type. The cell type loss is weighted according to the rarity weight of each cell type to obtain the comprehensive cell type loss. Thus, this embodiment balances the imbalance in prediction results caused by excessive differences in the number of cell populations by weighting the cell type loss for each cell type. Furthermore, the target total loss is determined based on gene expression loss, comprehensive cell type loss, and alignment loss. Therefore, this embodiment further focuses on rare but functionally critical cell type-related genes, thereby balancing the impact of different cell types on the prediction results, effectively reducing prediction bias caused by cell type imbalance, and improving the accuracy of gene expression prediction.
[0160] like Figure 5 As shown in the illustration, a specific embodiment of this application proposes a spatial transcriptome gene expression prediction method, which includes a model training phase and a model application phase:
[0161] (1) Training phase.
[0162] The raw spatial transcriptome data of the sample tissues undergoes quality control and screening to exclude poor-quality loci and abnormally expressed genes. This embodiment can use the CellChatDB database (intercellular communication database) to screen for intercellular communication-related genes matching the CellChatDB database from the raw spatial transcriptome data. This allows for the determination of the expression variation levels of these genes. Based on the expression variation levels, hypervariable genes are identified among the intercellular communication-related genes. For example, the top N intercellular communication-related genes in terms of expression variation levels are designated as hypervariable genes, or genes with expression variation levels exceeding a specified variation threshold are also designated as hypervariable genes. Furthermore, based on the Cell2location method, cell type-specific marker genes were screened from hypervariable genes. The screened cell type-specific marker genes were then subjected to log1p normalization to obtain spatial transcriptome gene data. Log1p normalization is a logarithmic transformation of gene expression values (UMI counts) by performing log10(1 + x) or ln(1 + x). This avoids infinitesimal values when taking the logarithm of genes with 0 expression levels, and can also compress the variance of highly expressed genes, making the data distribution closer to a normal distribution.
[0163] The H&E stained pathological images of the sample tissue were divided into several image blocks. The image blocks were uniformly adjusted to 224×224×3 RGB format and normalized. All image blocks were then standardized using ImageNet statistics to obtain each pathological image block, which was then adapted to the pre-trained image encoder.
[0164] From cell type description data containing cell type annotations, cell type abundance information for each pathological image patch is extracted, a corresponding cell type abundance vector is constructed, and corresponding cell type description text is generated based on the cell type abundance vector. Therefore, this embodiment can use pathological image data, spatial transcriptome gene data, and cell type abundance data of sample tissues as sample datasets for subsequent model training and testing. For scenarios where cell type description data is unavailable, this embodiment can use only the pathological image data and spatial transcriptome gene data of sample tissues as sample datasets for subsequent model training and testing. For example, for each training dataset, one sample is reserved for testing, and the remaining samples are divided into training and validation datasets according to a specified ratio (e.g., 8:2) to ensure no sample overlap between the test and training data.
[0165] After obtaining the training dataset, it can be used to train the prediction model to be trained. The prediction model to be trained includes a multimodal feature extraction network, a dual-alignment contrastive learning network, a spatial neighborhood aggregation network, and a prediction network. In this embodiment, the spatial neighborhood aggregation network and the prediction network can be frozen first. Then, pathological image patches, gene expression, and cell type abundance text from the training set can be input into the image encoder, gene expression encoder, and text encoder in the multimodal feature extraction network, respectively, to extract the pathological image features f of each pathological image patch. HE Gene expression f exp and cell type abundance characteristics f cell Then, a dual-alignment contrastive learning network is used to apply the InfoNCE contrastive loss L between the image and representation. clip-exp InfoNCE contrast loss between image and text L clip-cell Determined alignment loss L clip-total Alignment loss L clip-total The network parameters of the multimodal feature extraction network and the dual alignment contrast learning network are updated until the alignment loss L is reached. clip-total Convergence is achieved to realize dual alignment of image-expression and image-text, combined with a cross-attention mechanism to realize cross-modal interaction. Then, the multimodal feature extraction network and the dual alignment contrast learning network are frozen. Each pathological image patch is input into the frozen multimodal feature extraction network and the dual alignment contrast learning network, and the aligned pathological image features of each pathological image patch are output. Then, the spatial neighborhood aggregation network based on the Mamba model captures spatial dependencies through the selective state space mechanism of the Mamba model, processes the feature sequences of 7 image patches (1 image patch and 6 neighboring image patches), and obtains neighborhood aggregation features through global average pooling. The neighborhood aggregation features are input into the gene expression prediction head for prediction to obtain gene expression prediction samples. By inputting the gene expression prediction samples into the cell type prediction head, cell type prediction samples are obtained. Based on the difference between the real spatial transcriptome gene data and the gene expression prediction samples, L... exp The difference between actual cell type abundance data and cell type prediction samples L cell and alignment loss L clip-total The total target loss is determined. Based on the total target loss, the parameters of the spatial neighborhood aggregation module and the prediction network are updated until the total target loss converges. The multimodal feature extraction network, dual alignment contrast learning network, spatial neighborhood aggregation network and prediction network after the total target loss converges are determined as the target prediction model.
[0166] Among them, the total cell loss L corresponds to the difference between the actual cell type abundance data and the predicted cell type samples. cellThe long-tailed distribution problem can be solved by combining mean squared error and equalization term. Specifically, gene expression loss is determined based on the difference between spatial transcriptome gene data and gene expression prediction samples. Cell type loss is determined based on the difference between cell type abundance data and cell type prediction samples. The rarity weight of each cell type is determined, and the rarity weight is positively correlated with the scarcity of cell types. The cell type loss is weighted according to the rarity weight of each cell type to obtain the comprehensive cell type loss. The weighted sum of gene expression loss, comprehensive cell type loss and alignment loss is determined as the target total loss.
[0167] This embodiment can also train multiple target prediction models, and then evaluate the model performance of each target prediction model through a validation dataset, calculate indicators such as average PCC (Pearson Correlation Coefficient) and average SSIM (Structural Similarity Index Measure), save the prediction results and model weights, and select the target prediction model with the best model performance.
[0168] (2) Model application stage:
[0169] The H&E stained pathological images to be tested are acquired, preprocessed, and adjacent image patches corresponding to each image patch are extracted. Each image patch and its corresponding adjacent image patches are input into the trained target prediction model to obtain gene expression prediction results and cell type abundance prediction results. The gene expression prediction results and cell type abundance prediction results of each image patch are integrated to generate a complete spatial transcriptome gene expression prediction map.
[0170] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0171] Based on the same inventive concept, this application also provides a spatial transcriptome gene expression prediction device for implementing the spatial transcriptome gene expression prediction method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more spatial transcriptome gene expression prediction device embodiments provided below can be found in the limitations of the spatial transcriptome gene expression prediction method described above, and will not be repeated here.
[0172] In one exemplary embodiment, such as Figure 6 As shown, a spatial transcriptome gene expression prediction device 600 is provided, comprising: a data acquisition module 610, an image segmentation module 620, a feature determination module 630, a result prediction module 640, and a map generation module 650, wherein:
[0173] Data acquisition module 610 is used to acquire stained pathological image data of the target tissue;
[0174] The image segmentation module 620 is used to segment the stained pathological image data into images to obtain at least two image blocks to be identified.
[0175] The feature determination module 630 is used to determine, for each image block, a first cross-modal image feature of the image block and a second cross-modal image feature of the adjacent image blocks of the image block;
[0176] The result prediction module 640 is used to predict image patches based on the aggregation results of the first cross-modal image features and the second cross-modal image features, and to determine the gene expression prediction results and cell type prediction results corresponding to the image patches.
[0177] Map generation module 650 is used to generate a spatial transcriptome gene expression prediction map of the target tissue based on the gene expression prediction results and cell type prediction results of each image patch.
[0178] In some embodiments, the feature determination module 630 is further configured to:
[0179] The multimodal feature extraction network of the target prediction model is used to extract features from each image block to be identified, and the image block features of each image block to be identified are obtained. The image block features are then processed by the dual alignment contrast learning network of the target prediction model to obtain cross-modal image features. From the cross-modal image features, the first cross-modal image features of the image block and the second cross-modal image features of the adjacent image blocks of the image block are determined.
[0180] In some embodiments, the feature determination module 630 is further configured to:
[0181] Align the image patch features to obtain aligned image patch features. Perform cross-attention processing on the aligned image patch features to obtain attention features. Fuse the attention features and aligned image patch features to obtain cross-modal image features of the image patch features.
[0182] In some embodiments, the result prediction module 640 is further configured to:
[0183] The spatial neighborhood aggregation network of the target prediction model is used to aggregate the first cross-modal image features and the second cross-modal image features to obtain the aggregated features of the image patch. The aggregated features of the image patch are then input into the prediction network of the target prediction model to obtain the gene expression prediction results and cell type prediction results corresponding to the image patch.
[0184] In some embodiments, the spatial transcriptome gene expression prediction device 600 further includes a model training module for:
[0185] Pathological image data, spatial transcriptome gene data, and cell type abundance data of sample tissues are acquired. The pathological image data includes at least two pathological image blocks obtained by dividing the stained pathological images of the sample tissues into blocks according to specified specifications. Each pathological image block, spatial transcriptome gene data, and cell type abundance data are input into a prediction model to be trained. The prediction model to be trained includes a multimodal feature extraction network and a dual alignment contrast learning network. Through the multimodal feature extraction network, features are extracted from each pathological image block, spatial transcriptome gene data, and cell type abundance data to obtain pathological image features, gene expression features, and cell type abundance features of each pathological image block. Through the dual alignment contrast learning network, the pathological image features, gene expression features, and cell type abundance features are mapped to a shared mapping space to obtain an alignment loss. The parameters of the multimodal feature extraction network and the dual alignment contrast learning network are updated based on the alignment loss until the alignment loss converges to obtain cross-modal features.
[0186] In some embodiments, the model training module is further configured to:
[0187] The raw spatial transcriptome data, stained pathological images, and cell type description data of the sample tissue were obtained. The raw spatial transcriptome data were screened to obtain cell type-specific marker genes. The gene expression profiles of the cell type-specific marker genes were identified as spatial transcriptome gene data. The stained pathological images were divided into blocks according to specified specifications to obtain at least two pathological image blocks, which were identified as pathological image data. The cell type description data were converted into cell type description text and identified as cell type abundance data.
[0188] In some embodiments, the model training module is further configured to:
[0189] From the raw spatial transcriptome data, we screened out genes related to intercellular communication, determined the degree of expression variation of these genes, identified hypervariable genes among them, and screened out cell type-specific marker genes from these hypervariable genes.
[0190] In some embodiments, the model training module is further configured to:
[0191] When the alignment loss converges, cross-attention processing is performed using pathological image features as query indicators, gene expression features as keys, and cell type abundance features as values to obtain attention output features. The attention output features and pathological image features are then fused to obtain cross-modal image features of the pathological image patch.
[0192] In some embodiments, the prediction model to be trained further includes a spatial neighborhood aggregation network and a prediction network, and the model training module is further configured to:
[0193] The spatial neighborhood aggregation network aggregates the cross-modal features of each pathological image patch and its neighboring image patches to obtain the neighborhood aggregation features of each pathological image patch. These neighborhood aggregation features are then input into the prediction network to obtain gene expression prediction samples and cell type prediction samples corresponding to each pathological image patch. Based on the differences between spatial transcriptome gene data and gene expression prediction samples, the differences between cell type abundance data and cell type prediction samples, and the alignment loss, the target total loss is determined. Based on the target total loss, the parameters of the spatial neighborhood aggregation module and the prediction network are updated until the target total loss converges. The multimodal feature extraction network, dual alignment contrast learning network, spatial neighborhood aggregation network, and prediction network after the target total loss converges are determined as the target prediction model.
[0194] In some embodiments, the model training module is further configured to:
[0195] Gene expression loss is determined based on the difference between spatial transcriptome gene data and gene expression prediction samples. Cell type loss is determined based on the difference between cell type abundance data and cell type prediction samples. The rarity weight of each cell type is determined, and the rarity weight is positively correlated with the scarcity of the cell type. The cell type loss is weighted according to the rarity weight of each cell type to obtain the comprehensive cell type loss. Based on gene expression loss, comprehensive cell type loss, and alignment loss, the target total loss is determined.
[0196] Each module in the aforementioned spatial transcriptome gene expression prediction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0197] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a spatial transcriptome gene expression prediction method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0198] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0199] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0200] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0201] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0202] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0203] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0204] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A spatial transcriptome gene expression prediction method, characterized by, The method comprises: acquiring staining pathological image data of a target tissue; image blocking on the staining pathological image data to obtain at least two image blocks to be identified; for each image block, determining a first cross-modal image feature of the image block and a second cross-modal image feature of an adjacent image block of the image block; based on an aggregation result of the first cross-modal image feature and the second cross-modal image feature, predicting the image block to determine a gene expression prediction result and a cell type prediction result corresponding to the image block; based on the gene expression prediction result and the cell type prediction result of each image block, generating a spatial transcriptome gene expression prediction map of the target tissue.
2. The method of claim 1, wherein, The determination of the first cross-modal image feature of the image block and the second cross-modal image feature of the adjacent image block of the image block for each image block comprises: extracting features of each image block to be identified through a multi-modal feature extraction network of a target prediction model to obtain image block features of each image block to be identified; processing the image block features through a dual alignment contrast learning network of the target prediction model to obtain cross-modal image features; determining the first cross-modal image feature of the image block and the second cross-modal image feature of the adjacent image block of the image block from the cross-modal image features.
3. The method of claim 2, wherein, The processing of the image block features to obtain cross-modal image features comprises: aligning the image block features to obtain aligned image block features; performing cross-attention processing on the aligned image block features to obtain attention features; fusing the attention features and the aligned image block features to obtain cross-modal image features of the image block features.
4. The method of claim 2, wherein, The prediction of the image block based on the aggregation result of the first cross-modal image feature and the second cross-modal image feature to determine the gene expression prediction result and the cell type prediction result corresponding to the image block comprises: aggregating and processing the first cross-modal image feature and the second cross-modal image feature through a spatial neighborhood aggregation network of the target prediction model to obtain aggregated features of the image block; inputting the aggregated features of the image block into a prediction network of the target prediction model to obtain the gene expression prediction result and the cell type prediction result corresponding to the image block.
5. The method of claim 2, wherein, The method further comprises: acquiring pathological image data, spatial transcriptome gene data and cell type abundance data of a sample tissue, the pathological image data comprising at least two pathological image blocks obtained by blocking staining pathological images of the sample tissue according to a specified specification; inputting each pathological image block, the spatial transcriptome gene data and the cell type abundance data into a prediction model to be trained, wherein the prediction model to be trained comprises a multi-modal feature extraction network and a dual alignment contrast learning network; The multi-modal feature extraction network is used for feature extraction on each of the pathological image blocks, the spatial transcriptome gene data, and the cell type abundance data, to obtain pathological image features, gene expression features, and cell type abundance features of each of the pathological image blocks; The dual alignment contrast learning network is used for mapping the pathological image features, the gene expression features, and the cell type abundance features to a shared mapping space to obtain an alignment loss, and the multi-modal feature extraction network and the dual alignment contrast learning network are updated based on the alignment loss until the alignment loss converges, to obtain cross-modal features.
6. The method of claim 5, wherein, The pathological image data, the spatial transcriptome gene data, and the cell type annotation data of the sample tissue are obtained, including: The original spatial transcriptome data, the stained pathological image, and the cell type description data of the sample tissue are obtained; The original spatial transcriptome data is filtered to obtain cell type-specific marker genes, and a gene expression profile of the cell type-specific marker genes is determined as the spatial transcriptome gene data; The stained pathological image is divided into at least two pathological image blocks according to a specified specification to obtain the pathological image data; The cell type description data is converted into cell type description text to obtain the cell type abundance data.
7. The method of claim 6, wherein, The original spatial transcriptome data is filtered to obtain cell type-specific marker genes, including: Cell-to-cell communication related genes are filtered from the original spatial transcriptome data; The expression variation degree of the cell-to-cell communication related genes is determined; High-variable genes are determined from the cell-to-cell communication related genes according to the expression variation degree; Cell type-specific marker genes are filtered from the high-variable genes.
8. The method of claim 5, wherein, The alignment loss converges to obtain cross-modal image features, including: In the case where the alignment loss converges, the pathological image features are taken as query indicators, the gene expression features are taken as keys, and the cell type abundance features are taken as values for cross-attention processing to obtain attention output features; The attention output features and the pathological image features are fused to obtain cross-modal image features of the pathological image blocks.
9. The method of claim 5, wherein, The to-be-trained prediction model further includes a spatial neighborhood aggregation network and a prediction network, and after the alignment loss converges to obtain cross-modal features, the following steps are further included: The spatial neighborhood aggregation network is used for aggregating the cross-modal features of each of the pathological image blocks and neighboring image blocks of each of the pathological image blocks, respectively, to obtain neighborhood aggregation features of each of the pathological image blocks; The neighborhood aggregation features of each of the pathological image blocks are input into the prediction network, respectively, to obtain gene expression prediction samples and cell type prediction samples corresponding to each of the pathological image blocks; A target total loss is determined according to differences between the spatial transcriptome gene data and the gene expression prediction samples, differences between the cell type abundance data and the cell type prediction samples, and the alignment loss. Based on the target total loss, the parameters of the spatial neighborhood aggregation module and the prediction network are updated until the target total loss converges. The multimodal feature extraction network, dual alignment contrast learning network, spatial neighborhood aggregation network and prediction network after the target total loss converges are determined as the target prediction model.
10. The method of claim 9, wherein, The step of determining the target total loss based on the differences between the spatial transcriptome gene data and the gene expression prediction samples, the differences between the cell type abundance data and the cell type prediction samples, and the alignment loss includes: Gene expression loss is determined based on the difference between the spatial transcriptome gene data and the gene expression prediction samples; Based on the difference between the cell type abundance data and the cell type prediction samples, the cell type loss for each cell type is determined; Determine the rarity weight for each of the cell types, wherein the rarity weight is positively correlated with the scarcity of the cell type; The cell type loss is calculated by weighting the cell type loss according to the rarity weight of each cell type, and the overall cell type loss is obtained. The target total loss is determined based on the gene expression loss, the cell type combined loss, and the alignment loss.