Method and system for tissue region segmentation based on pathological images and spatial transcriptome

By fusing pathological images and spatial transcriptome information through a multi-stage, multi-modal deep learning framework, the problem of insufficient accuracy and stability in tissue region segmentation in existing technologies is solved, achieving high-precision tissue region segmentation and disease progression modeling, and improving the robustness and biological interpretability of the model.

CN122134744APending Publication Date: 2026-06-02SICHUAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN UNIV
Filing Date
2026-02-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies struggle to deeply integrate morphological information from high-resolution pathological images with molecular expression information from spatial transcriptomics within the same analytical framework. This results in insufficient accuracy, robustness, and biological interpretability in tissue region segmentation, making it difficult to meet the needs of clinical pathological diagnosis and refined tissue region analysis.

Method used

By constructing a multi-stage, multi-modal deep learning framework that integrates pathological images and spatial transcriptome data, molecular expression information is introduced as a constraint during the training phase. By utilizing spatial adjacency graphs and morpho-molecular consistency learning mechanisms, the model's segmentation accuracy and stability during the inference phase are improved, supporting uncertainty assessment of tissue regions and disease progression modeling.

Benefits of technology

It significantly improves the accuracy and boundary localization capabilities of complex tissue region segmentation, enhances the ability to identify morphologically similar but molecularly functionally different regions, improves the training stability and cross-sample generalization ability of the model, and supports the assessment of tissue region uncertainty and disease progression analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134744A_ABST
    Figure CN122134744A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for tissue region segmentation based on pathological images and spatial transcriptomics. The method includes: acquiring the H&E pathological image to be analyzed; preprocessing the H&E pathological image; and inputting the preprocessed H&E pathological image into a pre-trained tissue region segmentation model to obtain the tissue region segmentation result corresponding to the H&E pathological image. This invention focuses on the automated and fine segmentation of tissue regions in complex tissue sections, solving the key technical bottlenecks of existing technologies such as limited spatial resolution, insufficient utilization of molecular information, discontinuous region boundaries, and insufficient biological interpretability and reliability of segmentation results. It has good practicality and promotional value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image processing technology, and in particular relates to a method and system for tissue region segmentation based on pathological images and spatial transcriptomics. Background Technology

[0002] Precise segmentation of tissue regions is a fundamental and crucial step in digital pathology analysis, playing a vital role in tumor diagnosis, disease staging and prognostic assessment, biomarker discovery, and personalized treatment planning. In actual pathological analysis, tissues typically consist of various regions with different biological functions and clinical significance, including the tumor core, tumor periphery, areas enriched by immune cells, necrotic areas, normal parenchyma, and stroma. Accurately distinguishing these regions is an important prerequisite for achieving refined interpretation and intelligent analysis of pathological information.

[0003] Existing tissue region segmentation methods primarily rely on pathologists' manual interpretation based on hematoxylin and eosin (H&E) stained whole-section images, or on the analysis of pathological images using automated image segmentation models based on deep learning. These methods mainly use tissue morphological features as the basis for discrimination, such as cell morphology, tissue structure, and staining patterns. However, in practical applications, different tissue regions may be highly similar in morphology, but their molecular functional states and biological behaviors can differ significantly. For example, there is functional heterogeneity between different tumor subregions, differences in the invasiveness of tumor margins and core regions, and differences in immune cell infiltration patterns. Relying solely on morphological information can easily lead to inaccurate tissue region segmentation, limiting the biological interpretability and stability of the segmentation results.

[0004] Spatial transcriptome analysis tools such as SpaGCN and BayesSpace employ tissue region segmentation techniques. These techniques acquire spatial transcriptome sequencing data from tissue slices, assigning each spatial sequencing site a set of high-dimensional gene expression characteristics and its two-dimensional spatial coordinates within the tissue. Based on this, a site adjacency graph structure is constructed according to spatial distance relationships. Gene expression similarity and spatial proximity are used as joint modeling inputs, and graph convolutional networks, Markov random fields, or spatial Bayesian statistical models are used to perform cluster analysis on the spatial sites. This divides the tissue slices into several regions with similar characteristics at the molecular expression level, characterizing the molecular functional heterogeneity and potential biological partitions within the tissue. In practical applications, the tissue region segmentation results of technologies such as SpaGCN and BayesSpace are significantly limited by the sampling density and physical resolution of the spatial transcriptome raster. They can usually only perform region labeling at a relatively coarse point scale, making it difficult to form continuous, smooth region boundaries that conform to the real tissue structure. At the same time, these methods do not incorporate cell morphology, tissue structure, and boundary information from high-resolution pathological H&E images during the modeling process. This results in insufficient region localization accuracy in tissue regions with obvious morphological differences and spatial continuity, such as the tumor core, tumor edge, and immune cell-rich areas. As a result, they cannot meet the requirements of spatial accuracy and interpretability for clinical pathological diagnosis and fine tissue region analysis.

[0005] U-Net, Hover-Net, nnU-Net, and other tissue segmentation models based purely on H&E pathological slides employ techniques that involve H&E staining of tissue slides to obtain high-resolution digital pathological images. These methods utilize deep convolutional neural networks to automatically learn color distribution, texture features, cell morphology, and local tissue structure information from the images. By performing pixel-level or region-level classification, the pathological images are predicted pixel by pixel, thereby dividing the tissue into tumor areas, stroma areas, necrotic areas, or other predefined tissue regions. The training process typically relies on manually labeled morphological tags, and the model's discrimination criteria mainly come from the image's appearance features. The tissue segmentation models based on H&E pathological images rely entirely on tissue morphology features for region discrimination, lacking molecular-level information support such as spatial transcriptomics. In tissue regions with similar morphology but significantly different molecular functional states, region confusion or missegmentation is prone to occur. Furthermore, the segmentation results are difficult to explain from the perspective of gene expression or molecular mechanisms, and cannot effectively distinguish tumor subregions or immune microenvironment regions with different biological behaviors and prognostic significance. This limits the application value of such methods in fine tumor partitioning, biomarker discovery, and mechanism research.

[0006] Spatial transcriptomics technology can measure gene expression profiles at different spatial locations while preserving tissue spatial location information, providing a new means to reveal molecular heterogeneity within tissues. Spatial transcriptomics data can identify cell populations with different functional characteristics and their spatial distribution patterns at the molecular level, thus providing important molecular evidence for tissue region division. However, this technology is usually limited by the spatial resolution of sequencing arrays or captured regions, and its precision cannot match that of high-resolution pathological images. Furthermore, spatial transcriptomics data is high-dimensional and noisy; if clustering or region division is based solely on molecular expression, it often ignores tissue morphology and structural information, leading to unclear region boundaries or insufficient spatial continuity.

[0007] Currently, there is a lack of mature and effective techniques to deeply integrate the morphological information contained in high-resolution pathological images with the molecular expression information provided by spatial transcriptomics within the same analytical framework, thereby achieving tissue region segmentation that combines morphological accuracy with molecular specificity. Existing research largely remains at the level of simple feature splicing or posterior association analysis, failing to fully incorporate molecular information as a biological constraint during model training. Consequently, the segmentation results still suffer from shortcomings in accuracy, robustness, and biological interpretability. Therefore, developing a multimodal tissue region segmentation method that can integrate pathological images, spatial transcriptomics, and spatial structural information has become an urgent technical problem to be solved in this field. Summary of the Invention

[0008] To address the aforementioned shortcomings in existing technologies, the tissue region segmentation method and system based on pathological images and spatial transcriptomes provided by this invention solves the problems of existing tissue region segmentation methods based on spatial transcriptome data, such as difficulty in accurately depicting tissue regions with complex spatial structures, easy occurrence of region confusion or misjudgment, easy generation of blurred, broken, or unreasonable regional boundaries, insufficient biological interpretability of tissue region segmentation results, and difficulty in identifying regions with low prediction reliability.

[0009] To achieve the above-mentioned objectives, this invention provides a tissue region segmentation method based on pathological images and spatial transcriptomics, comprising: Obtain the pathological images of H&E to be analyzed; Preprocess the H&E pathology images to be analyzed; The preprocessed H&E pathological images are input into a pre-trained tissue region segmentation model to obtain the tissue region segmentation results corresponding to the H&E pathological images.

[0010] Secondly, the present invention also provides a tissue region segmentation system based on pathological images and spatial transcriptomics, comprising: The data acquisition module is used to acquire the H&E pathological images to be analyzed; The preprocessing module is used to preprocess the H&E pathological images to be analyzed; The segmentation module is used to input the preprocessed image into a pre-trained tissue region segmentation model to obtain the tissue region segmentation results corresponding to the H&E pathological image.

[0011] The beneficial effects of this invention are as follows: (1) Significantly improve the overall accuracy of complex tissue region segmentation: By introducing spatial transcriptome molecular expression to constrain pathological image features during the training phase, the model learns potential representations consistent with the real tissue functional state, thus enabling high-precision segmentation of complex tissue regions even when relying solely on H&E pathological images during the inference phase.

[0012] (2) Improve the ability to locate the boundaries of tissue regions and spatial continuity: This invention introduces explicit spatial consistency constraints through spatial adjacency graph modeling and regional boundary prediction mechanism, which effectively reduces the common problems of regional fragmentation, boundary blurring and prediction jump in high-resolution slices, making the segmentation results more continuous in space and consistent with the tissue anatomy.

[0013] (3) Enhance the ability to distinguish regions with similar morphology but different molecular functions: Through the morphology-molecular consistency learning mechanism, the model can use molecular supervision information to correct misjudgments caused by morphological features alone, and improve the accuracy of identifying regions with similar morphology but different biological properties, such as tumor margins, immune cell enrichment areas, and necrotic areas.

[0014] (4) Improve the stability and convergence reliability of the model training process: The present invention adopts a course-based training and dynamic loss scheduling strategy to gradually introduce different learning objectives into the training process according to their difficulty, thereby reducing the risk of gradient conflict in multi-task joint training and thus improving the training stability of the model under high-dimensional and multi-modal data conditions.

[0015] (5) Enhance the model’s generalization ability across samples and sources: By integrating multimodal information during the training phase and relying solely on pathological image input during the inference phase, the model learns more robust, sequencing-independent tissue representations in the latent space, which helps the model maintain stable performance under different sample sources and staining conditions.

[0016] (6) Support for uncertainty assessment of tissue region prediction results: This invention explicitly models and predicts uncertainty in the process of tissue region discrimination, which can identify regions with low model confidence and provide technical basis for subsequent manual review, key analysis or clinical auxiliary decision-making.

[0017] (7) Expanding the functional boundaries of traditional tissue segmentation methods: Based on tissue region segmentation, this invention further supports disease state stratification and progression trend modeling, providing a unified data foundation for tissue spatial heterogeneity analysis and disease evolution research, and enhancing the application value of the system in scientific research and auxiliary diagnosis scenarios. Attached Figure Description

[0018] Figure 1 Flowchart of a tissue region segmentation method based on pathological images and spatial transcriptomics provided for an embodiment; Figure 2 This is a schematic diagram illustrating the process of training the tissue region segmentation model in this embodiment. Detailed Implementation

[0019] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0020] like Figure 1 As shown, in one embodiment of the present invention, the tissue region segmentation method based on pathological images and spatial transcriptomics includes the following steps: S1. Obtain the H&E pathological images to be analyzed.

[0021] S2. Preprocess the H&E pathological images to be analyzed.

[0022] Preprocessing includes: The obtained H&E pathological images were cropped into fixed-size patches using a sliding window method, and the two-dimensional spatial coordinates of each patch in its corresponding H&E pathological image were recorded. The tissueless regions in the patches were automatically removed based on color thresholding or edge detection methods. The H&E pathological images from different batches and sources were color-normalized using a staining normalization method.

[0023] S3. Input the preprocessed H&E pathological image into the pre-trained tissue region segmentation model to obtain the tissue region segmentation result corresponding to the H&E pathological image.

[0024] The preprocessed patch blocks are input into the tissue region segmentation model one by one to obtain the latent space features and tissue region prediction results of each patch. Based on the spatial coordinates of the patch, the local prediction results are mapped back to the original slice space to generate a whole slice-level tissue region segmentation map.

[0025] During the model inference phase, the model does not rely on spatial transcriptome data; it only needs to input H&E pathological images to obtain tissue region segmentation results. This is because during the model training phase, the model has already used spatial transcriptome data as auxiliary information, which enhances the tissue region segmentation model's understanding of H&E pathological images. Thus, during the inference phase, the model can use the knowledge learned from spatial transcriptome data during training to segment tissue regions of H&E pathological images.

[0026] Specifically, the training process for the organization region segmentation model is as follows: Figure 2 As shown, it includes: Obtain high-resolution digitally scanned H&E pathology image samples: Obtain high-resolution digitally scanned whole slide images of H&E stained pathology tissue from clinical samples or public databases. The original slide images are typically scanned under 20× or 40× objectives, with image sizes reaching tens of thousands of pixels to ensure complete preservation of cellular morphology and structure.

[0027] The obtained H&E pathological image samples were preprocessed. The original H&E whole-slice images were cropped into fixed-size image patches (e.g., 224×224 pixels) using a sliding window method, and the two-dimensional spatial coordinates of each patch in the original slice were recorded. To avoid background interference, the system automatically removed tissue-free areas based on color thresholding or edge detection methods; at the same time, staining normalization methods (such as the Macenko method) were used to standardize the colors of slices from different batches and sources.

[0028] Spatial transcriptome data alignment: The spatial expression matrix corresponding to H&E pathological image samples is exported from the spatial transcriptome platform; based on the spatial location of each patch, the spatial expression matrix is ​​mapped or aggregated into a target spatial resolution grid, so that each patch corresponds to a gene expression vector; its expression is:

[0029] in, Indicates the number of selected genes. Represents the real number range, Indicates the first Each patch corresponds to a gene expression vector. This invention enhances the ability to distinguish functionally heterogeneous tissue regions by fusing spatial transcriptomic molecular expression information.

[0030] Spatial adjacency graph construction: Based on the spatial coordinates of all patches, a spatial adjacency graph is constructed using either the K-nearest neighbor method or a fixed radius method; its expression is:

[0031] in, Represents the spatial diagram structure. This represents a set of graph nodes, where each node is a patch in an H&E pathological image. This represents the graph edge set. This invention introduces a structural modeling mechanism based on spatial adjacency relationships to impose overall constraints on the organization region, thereby improving the spatial continuity and structural consistency of the segmentation results.

[0032] Based on the gene expression vector and spatial adjacency graph corresponding to each patch, the tissue region segmentation model is sequentially trained through representation learning, spatial modeling, region discrimination, and disease progression modeling.

[0033] A. The specific methods for representation learning training are as follows: Each patch is mapped to a high-dimensional representation vector using an image encoder, the expression of which is:

[0034] in, Represents the first in the H&E image One patch, Indicates an image encoder. Indicates pathological morphological characteristics; The gene encoder maps the gene expression vector corresponding to each patch to a latent space representation, the expression of which is:

[0035] in, Represents the gene encoder. This invention represents molecular expression characteristics. By simultaneously introducing pathological morphological information and spatial transcriptomic molecular information during the segmentation process, the resulting tissue regions correspond to specific molecular characteristics and biological functional states.

[0036] Introducing morphology-molecular flow matching loss, we learn the velocity fields for "predicting molecular changes from morphology" and "predicting morphology changes from molecules." The expression for the morphology-molecular flow matching loss is:

[0037] in, Representation of morphology-molecular flow matching loss, This represents the predicted velocity field that predicts molecular changes based on morphology. This represents the true velocity field used to predict molecular changes based on morphology. This represents the predicted velocity field that predicts morphological changes based on molecular patterns. This represents the true velocity field used to predict morphological changes from molecular perspectives. Denotes the Euclidean norm; A cycle consistency loss is introduced to constrain the consistency of the "image→gene→image" and "gene→image→gene" paths, preventing latent space collapse. Its expression is:

[0038] in, Indicates the loss of cycle consistency. This represents the reconstruction result of pathological morphological features after cross-modal cyclic mapping. This represents the reconstruction result of molecular expression features after cross-modal cyclic mapping; Image encoders and gene encoders are trained based on morphology-molecular flow matching loss and cycle consistency loss until the loss function value no longer decreases, thus obtaining a stable multimodal representation.

[0039] The goal of this phase is to learn stable potential representations that are simultaneously constrained by pathological morphology and molecular expression supervision, laying the foundation for subsequent spatial and regional modeling.

[0040] B. The specific methods for spatial modeling training are as follows: Multimodal feature fusion: This involves fusing the pathological morphological features and molecular expression features of the same node in a spatial adjacency graph. The expression is as follows:

[0041] in, Indicates the first Multimodal fusion features of spatial nodes; Region type prediction: The organizational region type of each node is predicted using a region classification network. The expression is as follows:

[0042] in, Represents a regional classification network. Represents a node The corresponding region category probability distribution; the region classification network is a trainable neural network module consisting of at least one fully connected layer, a non-activated linear layer, and a normalization layer, used to map the input multimodal fusion features to the corresponding tissue region category probability distribution; Region boundary detection: Predicts whether any pair of spatially adjacent nodes is located at the region boundary, and its expression is:

[0043] in, Represents the boundary probability. This represents the Sigmoid function. Represents a boundary prediction network. Indicates the first The multimodal fusion features of spatial nodes; the boundary prediction network is a trainable neural network module consisting of at least one feature transformation layer and one nonlinear activation layer, used to predict whether adjacent spatial nodes are located at the boundary position of different tissue regions based on their multimodal fusion features. Construct physical consistency space constraints; Specifically: For nodes The neighborhood of , the discrete Laplace, is expressed as:

[0044] in, Represents the neighborhood set, Representing an edge The weight, Represents a node The term "spatial second-order difference"; The expression for predicting source terms based on pathological morphological features is as follows:

[0045] in, Represents the source network. This represents the morphology-driven source term vector predicted from pathological morphological features; the source term network is a trainable neural network module consisting of at least one feature mapping layer and one nonlinear activation layer, used to generate the corresponding morphology-driven source vector based on pathological morphological features. opposite side The direction weights are calculated using the following expression:

[0046]

[0047] in, Represents a spatial displacement vector. This indicates a direction prediction network. Indicates the first The two-dimensional spatial coordinates of each spatial node in the original pathological slide. Indicates the first The spatial coordinates of each spatial node in the original pathological slice; the orientation prediction network is a trainable neural network module consisting of at least one feature fusion layer and one nonlinear activation layer, used to predict the directional propagation weights between nodes based on the pathological morphological features of adjacent spatial nodes and their spatial displacement relationships. By introducing a learnable diffusion coefficient for each gene dimension, the updated gene expression vector is obtained, and its expression is:

[0048] in, Indicates time step Gene expression vectors below, Indicates time step Gene expression vectors below, Indicates the number of iterations, ⊙ represents element-wise multiplication. Indicates the time step. This represents the diffusion coefficient of each gene; The physical consistency loss is constructed as follows:

[0049] in, Indicates the loss of physical consistency. Indicates the number of nodes. Indicates the total number of iterations; The region classification network, boundary prediction network, source term network, and orientation prediction network are trained based on physical consistency loss until the physical consistency loss value no longer decreases. By utilizing the spatial graph structure to introduce constraints of "diffusion + morphological source term + directionality", the representation within regions is more continuous and the boundaries between regions are more stable, avoiding spatial noise caused by classification alone.

[0050] After obtaining stable multimodal representations, this stage focuses on solving the problems of organizational region continuity, region boundary identification, and spatial consistency modeling.

[0051] C. Region discrimination training includes: Regional-level organization type determination: Based on multimodal fusion features, organization type determination is performed for each node, and the expression is as follows:

[0052] in, Represents an organization type classifier. The classification logits represent the organizational type classification network, which is a multilayer perceptron structure including fully connected layers and nonlinear activation functions, used to distinguish the organizational region type based on multimodal fusion features. Prediction uncertainty calculation: The prediction entropy is calculated based on the probability distribution of region categories, and its expression is:

[0053] in, Indicates the first The uncertainty measure of each node Let e ​​be a logarithmic function with the natural constant e as the base. Indicates the first The node belongs to the node. The predicted probability of a class; high uncertainty regions can be used to indicate areas where the model's confidence is insufficient, assisting in manual review or downstream analysis.

[0054] The cross-entropy loss for organizational region discrimination is expressed as follows:

[0055] In the formula, The cross-entropy loss represents the region discrimination of the organization. Indicates the number of organizational region categories. Indicates the first The actual label of each node; The tissue type classifier is trained using cross-entropy loss based on tissue region discrimination until the cross-entropy loss for tissue region discrimination no longer decreases. By introducing a reliability assessment mechanism into the segmentation framework, the stability and application security of tissue region segmentation results are improved.

[0056] After completing the regional modeling, this phase aims to further improve the reliability of the model in the task of organizing regional discrimination and to explicitly model and predict uncertainties.

[0057] In a preferred embodiment, based on tissue region segmentation and region-level discrimination, the present invention further introduces disease progression modeling to characterize the potential spatial evolutionary state and progression direction of complex tissues (especially tumor tissues). This stage, as an independent modeling module decoupled from the previous three stages, can be trained later in the training process or independently without affecting the stability of the main task of tissue region segmentation.

[0058] D. Disease progression training includes: Disease state prediction: Define a finite set of disease states, expressed as:

[0059] in, Indicates the number of disease progression stages. Indicates the first The disease state at each stage; in a preferred embodiment, These can be represented as discrete stages such as normal tissue, dysplasia, in situ lesions, invasive lesions, and metastasis-related states.

[0060] Based on multimodal fusion features, a disease state prediction network predicts the disease state, and its expression is as follows:

[0061] in, This represents a disease state prediction network. Indicates the predicted first The disease state is labeled with discrete disease status at each node; the disease state prediction network is a multi-layer fully connected neural network module used to map multimodal fusion features to disease state classification results. This step ensures that each spatial region not only has a histological partition label, but also a clear disease progression stage attribute.

[0062] Pseudo-temporal modeling of disease: Since disease progression is usually not a strictly discrete jump but rather exhibits continuous evolutionary characteristics, this invention further introduces a continuous pseudo-temporal variable to characterize the relative position of a region within the disease progression spectrum. For each region, the relative position of the region in disease progression is predicted by a pseudo-temporal prediction network, expressed as:

[0063] in, Representing spatial nodes The degree of disease progression in the corresponding tissue region, This refers to a pseudo-temporal prediction network, which is a regression-based multilayer fully connected neural network used to predict the continuous pseudo-temporal location of tissue regions during disease progression. Pseudo-temporal variables can distinguish different degrees of progression within the same disease state, such as differentiating between early and late infiltrative areas, thereby improving the ability to characterize tissue heterogeneity.

[0064] Spatial state transition modeling and monotonicity constraints: At the spatial level, disease progression often exhibits a certain degree of spatial continuity along tissue structures or infiltration directions. Therefore, based on a spatial adjacency graph, a state transition prediction network is used to model the disease state transition relationships between adjacent regions, with the following expression:

[0065] Based on prior biological knowledge of disease progression, this invention imposes a monotonic constraint on the state transition process:

[0066] This means prohibiting unreasonable spatial shifts in disease states that represent "reverse regression," allowing only the state to remain unchanged or evolve to a higher stage of progression. This constraint can be implemented during training using a mask matrix or normalization operations.

[0067] in, This represents a state transition prediction network. Representing nodes in a spatial adjacency graph The corresponding tissue area is in a disease state. At that time, its adjacent nodes The corresponding tissue area shifts to the disease state. The probability; the state transition prediction network is a multi-layer fully connected neural network used to model the disease state transition relationships between adjacent regions in the spatial adjacency graph.

[0068] The disease state prediction network, pseudo-temporal prediction network, and state transition prediction network are trained based on the total loss function until the loss function value no longer decreases, thus obtaining the trained tissue region segmentation model.

[0069] To address the challenges of complex optimization objectives, diverse and coupled loss function types in the joint modeling of multimodal, high-resolution pathological images and spatial transcriptomes, this invention further proposes a curriculum-based training and dynamic loss scheduling mechanism. This mechanism adaptively adjusts the participation level and weight ratio of different loss terms based on the training stage and model convergence status during training, thereby improving training stability and final model performance.

[0070] The specific expression for the total loss function during model training is as follows:

[0071] in, Represents the total loss function. Indicates the current training round. The cross-entropy loss representing the organization region discrimination is in the first... Dynamic weight coefficients in round training, The cross-entropy loss represents the region discrimination of the organization. Representation of morphology – molecular flux matching loss in the first... Dynamic weight coefficients in round training, The pathological perception self-supervised learning loss represents the loss at the 1st... Dynamic weight coefficients in round training, This indicates the loss in pathological perception through self-supervised learning. The cycle consistency loss is represented in the th... Dynamic weight coefficients in round training, The regional organizational structure classification loss is represented in the first... Dynamic weight coefficients in round training, This indicates the loss in the regional organizational structure classification. The physical consistency loss is represented in the th... Dynamic weight coefficients in round training, The weighted discriminant loss guided by uncertainty is in the first... Dynamic weight coefficients in round training, This represents the weighted discriminant loss guided by uncertainty.

[0072] Among them, regional organizational structure classification loss Specifically:

[0073] In the formula, Indicates the total number of regions. Indicates the first The region is the first Real labels for each category Indicates the first The region belongs to the first The predicted probabilities of each category; Uncertainty-guided weighted discriminant loss Specifically:

[0074] In the formula, This represents the node-level weighting coefficients derived from uncertainty, used to assign greater loss weights to nodes with higher prediction uncertainty. Indicates the first The node at the th Real labels on the tissue-like regions; Among them, self-supervised learning loss Specifically:

[0075]

[0076]

[0077] in, This represents the gene expression gradient prediction loss, used to constrain the direction and magnitude of molecular expression changes between spatially adjacent nodes. This represents the contrastive learning loss, used to bring together the feature distributions of similar tissue regions and widen the feature distributions of dissimilar regions. Represents the set of edges in a spatial adjacency graph. , Representing nodes respectively and nodes The corresponding actual gene expression vector; This represents the gene expression gradient predicted by the gradient prediction network; the gradient prediction network is a multi-layer fully connected neural network. This represents a set of positive sample pairs, containing node pairs with the same or similar organizational region attributes; Represented by node The set of comparison samples for anchor points; Represents the feature similarity function; represents the temperature coefficient, used to adjust the smoothness of the similarity distribution; exp represents an exponential function with the natural constant e as the base. Represents a node Multimodal fusion features.

[0078] This combined loss is used to enhance the model's ability to perceive pathological structures and molecular changes under conditions of no or weak manual annotation.

[0079] In this embodiment, the dynamic adjustment of the various loss weight coefficients is achieved using any combination of the following methods: linear growth or decay function, constant maintenance after warm-up, sigmoid smoothing function, cosine decay function, and delayed onset function. This approach allows different loss terms to gradually enter or exit the optimization process smoothly and controllably during training, avoiding training instability caused by abrupt loss switching.

[0080] This invention improves the overall performance of tissue region segmentation by constructing a unified multimodal learning framework to achieve collaborative modeling of pathological image features, spatial transcriptome features, and spatial structural information.

[0081] This invention provides a method and system for tissue region segmentation based on pathological images and spatial transcriptomics. By constructing a phased, multimodal, and scalable deep learning framework, the method integrates Hematologic and Escherichia coli (H&E) pathological images, spatial transcriptomics molecular expression information, and spatial structural information during the training phase. During the inference phase, it relies solely on H&E pathological images to achieve accurate segmentation of tissue regions such as tumor core, tumor margin, immune cell-rich areas, necrotic areas, normal parenchyma, and stroma in complex tissue sections. Furthermore, it supports tissue region uncertainty assessment and disease progression modeling. The invention employs a curriculum-based training strategy, dividing the model training process into four interconnected but functionally decoupled stages: phenotypic representation learning, spatial and tissue region modeling, discriminative learning and uncertainty assessment, and disease progression modeling. Each stage addresses different levels of technical problems and works collaboratively within a unified framework to form a complete end-to-end tissue region segmentation and analysis system.

[0082] In summary, this invention, through a multi-stage, multi-modal, and spatially consistent system design, effectively solves the problems of insufficient accuracy, poor stability, and lack of progress modeling capability in existing organizational region segmentation methods under complex organizational scenarios, and has good practicality and promotion value.

Claims

1. A tissue region segmentation method based on pathological images and spatial transcriptomics, characterized in that, include: Obtain the pathological images of H&E to be analyzed; Preprocess the H&E pathology images to be analyzed; The preprocessed H&E pathological images are input into a pre-trained tissue region segmentation model to obtain the tissue region segmentation results corresponding to the H&E pathological images.

2. The method according to claim 1, characterized in that, Preprocessing includes: The obtained H&E pathological images were cropped into fixed-size patches using a sliding window method, and the two-dimensional spatial coordinates of each patch in its corresponding H&E pathological image were recorded. Automatically remove disorganized areas from a patch based on color thresholding or edge detection methods; The staining normalization method was used to standardize the colors of H&E pathological images from different batches and sources.

3. The method according to claim 2, characterized in that, The training process for the tissue region segmentation model includes: Acquire high-resolution digitally scanned H&E pathological image samples; The obtained H&E pathological image samples were preprocessed; Spatial transcriptome data alignment: The spatial expression matrix corresponding to H&E pathological image samples is exported from the spatial transcriptome platform; based on the spatial location of each patch, the spatial expression matrix is ​​mapped or aggregated into a target spatial resolution grid, so that each patch corresponds to a gene expression vector; its expression is: in, Indicates the number of selected genes. Represents the real number range, For dimension A real vector matrix, Indicates the first The gene expression vector corresponding to each patch; Spatial adjacency graph construction: Based on the spatial coordinates of all patches, a spatial adjacency graph is constructed using either the K-nearest neighbor method or a fixed radius method; its expression is: in, Represents the spatial diagram structure. This represents a set of graph nodes, where each node is a patch in an H&E pathological image. Represents the set of edges in a graph; Based on the gene expression vector and spatial adjacency graph corresponding to each patch, the tissue region segmentation model is sequentially trained through representation learning, spatial modeling, region discrimination, and disease progression modeling.

4. The method according to claim 3, characterized in that, The specific methods for representation learning training are as follows: Each patch is mapped to a high-dimensional representation vector using an image encoder, the expression of which is: in, Represents the first in the H&E image One patch, Indicates an image encoder. Indicates pathological morphological characteristics; The gene encoder maps the gene expression vector corresponding to each patch to a latent space representation, the expression of which is: in, Represents the gene encoder. Indicates molecular expression characteristics; Introducing morphology-molecular flux matching loss, we learn the velocity fields for "predicting molecular changes from morphology" and "predicting morphology changes from molecules." The expression for the morphology-molecular flux matching loss is: in, Representation of morphology-molecular flow matching loss, This represents the predicted velocity field that predicts molecular changes based on morphology. This represents the true velocity field used to predict molecular changes based on morphology. This represents the predicted velocity field that predicts morphological changes based on molecular patterns. This represents the true velocity field used to predict morphological changes from molecular perspectives. Denotes the Euclidean norm; Introducing the cycle consistency loss, its expression is: in, Indicates the loss of cycle consistency. This represents the reconstruction result of pathological morphological features after cross-modal cyclic mapping. This represents the reconstruction result of molecular expression features after cross-modal cyclic mapping; Image encoders and gene encoders are trained based on morphology-molecular flow matching loss and cycle consistency loss until the loss function value no longer decreases, thus obtaining a stable multimodal representation.

5. The method according to claim 4, characterized in that, The specific methods for spatial modeling training are as follows: The pathological morphological features and molecular expression features of the same node in the spatial adjacency graph are fused together; Its expression is: in, Indicates the first Multimodal fusion features of spatial nodes; The expression for predicting the organizational region type of each node using a region classification network is as follows: in, Represents a regional classification network. Represents a node The corresponding region category probability distribution; the region classification network is a trainable neural network module consisting of at least one fully connected layer, a non-activated linear layer, and a normalization layer, used to map the input multimodal fusion features to the corresponding tissue region category probability distribution; The expression for predicting whether any spatially adjacent pair of nodes is located on the region boundary is: in, Represents the boundary probability. This represents the Sigmoid function. Represents a boundary prediction network. Indicates the first The multimodal fusion features of spatial nodes; the boundary prediction network is a trainable neural network module consisting of at least one feature transformation layer and one nonlinear activation layer, used to predict whether adjacent spatial nodes are located at the boundary position of different tissue regions based on their multimodal fusion features. Construct physical consistency space constraints; Specifically: For nodes The neighborhood of , the discrete Laplace, is expressed as: in, Represents the neighborhood set, Representing an edge The weight, Represents a node The term "spatial second-order difference"; The expression for predicting source terms based on pathological morphological features is as follows: in, Represents the source network. This represents the morphology-driven source term vector predicted from pathological morphological features; the source term network is a trainable neural network module consisting of at least one feature mapping layer and one nonlinear activation layer, used to generate the corresponding morphology-driven source vector based on pathological morphological features. opposite side The direction weights are calculated using the following expression: in, Represents a spatial displacement vector. This indicates a direction prediction network. Indicates the first The two-dimensional spatial coordinates of each spatial node in the original pathological slide. Indicates the first The spatial coordinates of each spatial node in the original pathological slice; the orientation prediction network is a trainable neural network module consisting of at least one feature fusion layer and one nonlinear activation layer, used to predict the directional propagation weights between nodes based on the pathological morphological features of adjacent spatial nodes and their spatial displacement relationships. By introducing a learnable diffusion coefficient for each gene dimension, the updated gene expression vector is obtained, and its expression is: in, Indicates time step Gene expression vectors below, Indicates time step Gene expression vectors below, Indicates the number of iterations, ⊙ represents element-wise multiplication. Indicates the time step. This represents the diffusion coefficient of each gene; The physical consistency loss is constructed as follows: in, Indicates the loss of physical consistency. Indicates the number of nodes. Indicates the total number of iterations; The region classification network, boundary prediction network, source term network, and orientation prediction network are trained based on the physical consistency loss until the physical consistency loss value no longer decreases.

6. The method according to claim 5, characterized in that, Region discrimination training includes: Based on multimodal fusion features, the organization type is determined for each node, and its expression is as follows: in, Represents a network for classifying organizational types. The classification logits represent the organizational type classification network, which is a multilayer perceptron structure including fully connected layers and nonlinear activation functions, used to distinguish the organizational region type based on multimodal fusion features. The prediction entropy is calculated based on the probability distribution of region categories, and its expression is as follows: in, Indicates the first The uncertainty measure of each node Let e ​​be a logarithmic function with the natural constant e as the base. Indicates the first The node belongs to the node. The predicted probability of a class; The cross-entropy loss for organizational region discrimination is expressed as follows: In the formula, The cross-entropy loss represents the region discrimination of the organization. Indicates the number of organizational region categories. Indicates the first The real label of each node; The tissue type classifier is trained based on the cross-entropy loss for tissue region discrimination until the cross-entropy loss for tissue region discrimination no longer decreases.

7. The method according to claim 6, characterized in that, Disease progression training includes: Define a finite set of disease states Its expression is: in, Indicates the number of disease progression stages. Indicates the first The disease state at each stage; Based on multimodal fusion features, a disease state prediction network predicts the disease state, and its expression is as follows: in, This represents a disease state prediction network. Indicates the predicted first The disease state label of each node; the disease state prediction network is a multi-layer fully connected neural network module used to map multimodal fusion features to the classification result of disease state; Based on multimodal fusion features, a pseudo-temporal prediction network is used to predict the relative position of a region during disease progression. The expression for this prediction is: in, Representing spatial nodes The degree of disease progression in the corresponding tissue region, This represents a pseudo-temporal prediction network; a pseudo-temporal prediction network is a regression-type multilayer fully connected neural network used to predict the continuous pseudo-temporal location of tissue regions during disease progression. Based on a spatial adjacency graph, a state transition prediction network is used to model the disease state transition relationships between adjacent regions. The expression is as follows: in, This represents a state transition prediction network. Representing nodes in a spatial adjacency graph The corresponding tissue area is in a disease state. At that time, its adjacent nodes The corresponding tissue area shifts to the disease state. The probability; the state transition prediction network is a multi-layer fully connected neural network used to model the disease state transition relationships between adjacent regions in the spatial adjacency graph; The disease state prediction network, pseudo-temporal prediction network, and state transition prediction network are jointly trained based on the total loss function until the loss function value no longer decreases, thus obtaining the trained tissue region segmentation model.

8. The method according to claim 7, characterized in that, The specific expression for the total loss function is: in, Represents the total loss function. Indicates the current training round. The cross-entropy loss representing the organization region discrimination is in the first... Dynamic weight coefficients in round training, The cross-entropy loss represents the region discrimination of the organization. Representation of morphology – molecular flux matching loss in the first... Dynamic weight coefficients in round training, The pathological perception self-supervised learning loss represents the loss at the 1st... Dynamic weight coefficients in round training, This indicates the loss in pathological perception through self-supervised learning. The cycle consistency loss is represented in the th... Dynamic weight coefficients in round training, The regional organizational structure classification loss is represented in the first... Dynamic weight coefficients in round training, This indicates the loss in the regional organizational structure classification. The physical consistency loss is represented in the th... Dynamic weight coefficients in round training, The weighted discriminant loss guided by uncertainty is in the first... Dynamic weight coefficients in round training, This represents the weighted discriminant loss guided by uncertainty. Among them, regional organizational structure classification loss Specifically: In the formula, Indicates the total number of regions. Indicates the first The region is the first Real labels for each category Indicates the first The region belongs to the first The predicted probabilities of each category; Uncertainty-guided weighted discriminant loss Specifically: In the formula, This represents the node-level weighting coefficients derived from uncertainty, used to assign greater loss weights to nodes with higher prediction uncertainty. Indicates the first The node at the th Real labels on the tissue-like regions; Among them, self-supervised learning loss Specifically: in, This represents the gene expression gradient prediction loss, used to constrain the direction and magnitude of molecular expression changes between spatially adjacent nodes. This represents the contrastive learning loss, used to bring together the feature distributions of similar tissue regions and widen the feature distributions of dissimilar regions. Represents the set of edges in a spatial adjacency graph. , Representing nodes respectively and nodes The corresponding actual gene expression vector; This represents the gene expression gradient predicted by the gradient prediction network; the gradient prediction network is a multi-layer fully connected neural network. This represents a set of positive sample pairs, containing node pairs with the same or similar organizational region attributes; Represented by node The set of comparison samples for anchor points; Represents the feature similarity function; represents the temperature coefficient, used to adjust the smoothness of the similarity distribution; exp represents an exponential function with the natural constant e as the base. Represents a node Multimodal fusion features.

9. The method according to claim 7, characterized in that, The following methods, in combination, can be used to achieve dynamic adjustment of the weighting coefficients: Linear growth or decay function, constant after warm-up, sigmoid smoothing enable function, cosine decay function, delayed enable function.

10. A tissue region segmentation system based on pathological images and spatial transcriptomics, characterized in that, include: The data acquisition module is used to acquire the H&E pathological images to be analyzed; The preprocessing module is used to preprocess the H&E pathological images to be analyzed; The segmentation module is used to input the preprocessed image into a pre-trained tissue region segmentation model to obtain the tissue region segmentation results corresponding to the H&E pathological image.