Pathological image staining normalization method and system based on global and local feature interaction

By employing a pathological image staining normalization method based on global-local feature interaction, and utilizing a lightweight multi-scale Transformer and a multi-scale hierarchical perception module, a pathological image staining normalization model is constructed. This solves the problem of ignoring the local and global feature dependencies in existing methods, achieving efficient staining normalization of pathological images and reducing computational complexity.

CN121837846APending Publication Date: 2026-04-10GUANGDONG GENERAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing pathological image staining normalization methods ignore the dependencies between local and global features, resulting in local over-staining and blurred areas in the processed images. Furthermore, supervised learning-based methods rely on paired pathological images, which limits the application of computer-aided diagnostic systems.

Method used

A pathological image staining normalization method based on global-local feature interaction is adopted. By introducing a lightweight multi-scale Transformer structure and a multi-scale hierarchical perception module, a pathological image staining normalization model is constructed. Combined with a global-local feature interaction perception network and a fusion network, the system can comprehensively perceive pathological image features, thereby reducing the number of parameters and computational complexity.

Benefits of technology

It achieves staining normalization for pathological images from different centers and batches, improving the performance and generalization of pathological image staining normalization, enabling comprehensive perception of pathological image features and reducing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121837846A_ABST
    Figure CN121837846A_ABST
Patent Text Reader

Abstract

The invention discloses a pathological image dyeing normalization method and system based on global and local feature interaction. The method comprises the following steps: acquiring a pathological image; the method comprises the following steps: introducing a lightweight multi-scale Transform structure and a multi-scale hierarchical sensing module, and constructing a pathological image dyeing normalization model; and carrying out dyeing normalization processing on the pathological image based on the pathological image dyeing normalization model to obtain a normalized pathological image. According to the method, pathological image features can be comprehensively perceived, the parameter quantity and the calculation complexity of the pathological image features can be reduced, and dyeing normalization of different centers and different batches of pathological images is realized. The pathological image dyeing normalization method and system based on global and local feature interaction can be widely applied to the technical field of image dyeing normalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image staining normalization technology, and in particular to a method and system for pathological image staining normalization based on global-local feature interaction. Background Technology

[0002] Pathological examination is the gold standard for cancer diagnosis. In clinical diagnosis, hematoxylin and eosin are often used to stain cell nuclei and cytoplasm purple-blue and pink, respectively, to reveal hidden pathological features in biopsy tissue samples, such as tumor epithelium, tumor-infiltrating lymphocytes (TILs), tumor-associated stroma, and other tumor microenvironments. However, due to factors such as the staining process and staining type, inconsistencies in the color of stained pathological images can occur, affecting the pathologist's diagnostic results. Histogram matching corrects color deviations by processing pixel values ​​in the RGB color space, but this can lead to a loss of detail in the pathological image. Color normalization achieves staining normalization by analyzing the pathological image pixel by pixel. Color enhancement methods reduce the color difference between the source and target domains of the pathological image through color separation and normalization. In recent years, artificial intelligence (AI) technology, with deep learning at its core, has been widely applied in fields such as medical image processing and disease diagnosis. In the task of pathological image staining normalization, researchers have proposed learning-based pathological staining normalization models. These models utilize convolutional neural networks (CNNs), Transformers, and generative adversarial networks (GANs) to mine the mapping relationship between the source and target domains of pathological images through supervised, unsupervised, and supervised learning strategies. This aims to alleviate staining inconsistencies between pathological images from different centers and batches, and further develop computer-aided diagnostic (CAD) systems to help pathologists make more accurate diagnoses. While these methods can mitigate the impact of color deviations in histopathological image analysis, supervised learning-based pathological images heavily rely on paired pathological images, limiting their application in CAD systems. Conversely, the latter can effectively remove staining differences in pathological images without paired inputs. Generative adversarial network (GAN)-based methods such as StainGAN and CytoGAN are the most typical examples. However, existing GAN-based staining normalization models ignore the dependencies between local and global features, introducing local over-staining and blurred regions into the processed images. Summary of the Invention

[0003] To address the aforementioned technical problems, the present invention aims to provide a method and system for normalizing pathological image staining based on global-local feature interaction. This method can comprehensively perceive the features of pathological images and reduce their parameter count and computational complexity, thereby achieving staining normalization of pathological images from different centers and batches.

[0004] The first technical solution adopted in this invention is: a pathological image staining normalization method based on global-local feature interaction, comprising the following steps:

[0005] Acquire pathological images;

[0006] A lightweight multi-scale Transformer structure and a multi-scale hierarchical perception module are introduced to construct a staining normalization model for pathological images.

[0007] Based on the pathological image staining normalization model, the pathological images are subjected to staining normalization processing to obtain normalized pathological images.

[0008] Furthermore, the pathological image staining normalization model specifically includes a global-local feature interaction perception network and a global-local feature fusion network, wherein the output of the global-local feature interaction perception network is connected to the input of the global-local feature fusion network, and wherein:

[0009] The global-local feature interaction perception network includes a global feature extraction branch, a local feature perception branch, and a feature interaction module. The global feature extraction branch includes a large convolutional layer and several stacked lightweight multi-scale Transformer modules. The local feature perception branch includes a shallow feature extractor and several stacked multi-scale hierarchical perception modules. The first output of the lightweight multi-scale Transformer module is connected to the input of the next stacked lightweight multi-scale Transformer module, and the second output of the lightweight multi-scale Transformer module is connected to the input of the feature interaction module. The first output of the multi-scale hierarchical perception module is connected to the input of the next stacked multi-scale hierarchical perception module, and the second output of the multi-scale hierarchical perception module is connected to the input of the feature interaction module.

[0010] The global-local feature fusion network includes a global pooling layer, an average pooling layer, and a post-processing module. The post-processing module includes a compression and activation network, a first post-processing block, and a second post-processing block. Both the first and second post-processing blocks include a convolutional layer, a normalization layer, and a ReLU function.

[0011] Furthermore, the lightweight multi-scale Transformer module specifically includes a first normalization layer, a first convolutional module, a second convolutional module, a third convolutional module, a first concatenated feature mining module, a second concatenated feature mining module, a third concatenated feature mining module, a channel attention mechanism, a second normalization layer, a first convolutional layer, and a second convolutional layer. The output of the first normalization layer is connected to the inputs of the first convolutional module, the second convolutional module, and the third convolutional module, respectively. The output of the first convolutional module is connected to the input of the first concatenated feature mining module, the output of the second convolutional module is connected to the input of the second concatenated feature mining module, and the output of the third convolutional module is connected to the input of the second concatenated feature mining module. The first convolutional module is connected to the input of the third splicing feature mining module. The outputs of the first and second splicing feature mining modules are both connected to the channel attention mechanism. The outputs of the channel attention mechanism and the third splicing feature mining module are both connected to the input of the second normalization layer. The output of the second normalization layer is connected to the inputs of the first and second convolutional layers, respectively. The first, second, and third convolutional modules each include several convolutional layers with different dilation rates. The first, second, and third splicing feature mining modules each include a splicing layer and a compression excitation network.

[0012] Furthermore, the multi-scale hierarchical perception module specifically includes a first convolutional layer, a ReLU function, a second convolutional layer, a pixel attention layer, a first global pooling layer, a first average pooling layer, a third convolutional layer, a channel attention layer, and a lightweight multi-scale structure. The output of the first convolutional layer is connected to the input of the second convolutional layer, the input of the third convolutional layer, and the input of the ReLU function. The output of the ReLU function is connected to the input of the first global pooling layer and the input of the first average pooling layer. The output of the second convolutional layer is connected to the input of the pixel attention layer. The output of the third convolutional layer is connected to the input of the channel attention layer. The outputs of the first global pooling layer and the first average pooling layer are both connected to the input of the lightweight multi-scale structure. The lightweight multi-scale structure includes several convolutional and compression / excitation network blocks with different dilation rates.

[0013] Furthermore, the feature interaction module specifically includes a first convolutional layer, a second convolutional layer, a first self-attention module, a second self-attention module, a third convolutional layer, a fourth convolutional layer, a first sigmoid function, a second sigmoid function, and a cross-attention module. The output of the first convolutional layer is connected to the input of the first self-attention module, the output of the second convolutional layer is connected to the input of the second self-attention module, the output of the first self-attention module is connected to the inputs of the third convolutional layer and the cross-attention module, the output of the second self-attention module is connected to the inputs of the fourth convolutional layer and the cross-attention module, the third convolutional layer is connected to the first sigmoid function, and the fourth convolutional layer is connected to the second sigmoid function.

[0014] Furthermore, the step of performing staining normalization processing on the pathological image based on the pathological image staining normalization model to obtain the normalized pathological image specifically includes:

[0015] Input the pathological images into the pathological image staining normalization model;

[0016] The global feature extraction branch in the global-local feature interaction perception network based on the pathological image staining normalization model is used to extract global feature information from the pathological image to obtain the global features of the pathological image.

[0017] Based on the local feature perception branch in the global-local feature interaction perception network of the pathological image staining normalization model, local feature information is extracted from the pathological image to obtain the local features of the pathological image.

[0018] The feature interaction module in the global-local feature interaction perception network based on the pathological image staining normalization model performs feature interaction processing on the global features and local features of the pathological image to obtain the autocorrelation of the global features and local features of the pathological image and the cross-correlation of the global-local features of the pathological image.

[0019] The global pooling layer and average pooling layer in the global-local feature fusion network based on the pathological image staining normalization model pool and concatenate the autocorrelation of global and local features of the pathological image and the cross-correlation of global and local features of the pathological image to obtain the filtered pathological image.

[0020] The post-processing module in the global-local feature fusion network based on the pathological image staining normalization model explores and normalizes the correlation between different feature channels of the filtered pathological image to obtain the normalized pathological image.

[0021] Furthermore, the global feature extraction branch in the global-local feature interaction perception network based on the pathological image staining normalization model, which performs global feature information extraction processing on the pathological image to obtain the global features of the pathological image, specifically includes:

[0022] The pathological image is input into the global feature extraction branch of the global-local feature interaction perception network of the pathological image staining normalization model;

[0023] Based on a large convolutional layer with a global feature extraction branch, feature extraction is performed on pathological images to obtain global information of the pathological images;

[0024] The first normalization layer in the lightweight multi-scale Transformer module based on the global feature extraction branch normalizes the global information of the pathological image to obtain the normalized global information of the pathological image.

[0025] The first convolutional module, second convolutional module, third convolutional module, first stitching feature mining module, second stitching feature mining module, and third stitching feature mining module in the lightweight multi-scale Transformer module based on the global feature extraction branch mine the correlation between feature channels of the normalized pathological image global information to obtain the query value, key value, and value of the Transformer block.

[0026] The channel attention mechanism in the lightweight multi-scale Transformer module based on the global feature extraction branch calculates the query value and key value of the Transformer block to obtain the feature map.

[0027] The feature map and the value of the Transformer block are multiplied and added to the global information of the pathological image. The result is then fed into the second normalization layer, the first convolutional layer, and the second convolutional layer for normalization and dot product calculations to obtain the global features of the pathological image.

[0028] Furthermore, the local feature perception branch in the global-local feature interaction perception network based on the pathological image staining normalization model extracts local feature information from the pathological image to obtain local features of the pathological image. This step specifically includes:

[0029] The pathological image is input into the local feature perception branch of the global-local feature interaction perception network of the pathological image staining normalization model.

[0030] A shallow feature extractor based on a local feature perception branch is used to perform shallow feature extraction processing on pathological images to obtain shallow features of the pathological images.

[0031] The first convolutional layer in the multi-scale hierarchical perception module based on the local feature perception branch performs convolution processing on the shallow features of the pathological image to obtain preliminary local information of the pathological image.

[0032] Based on the ReLU function, the first global pooling layer and the first average pooling layer in the multi-scale hierarchical perception module of the local feature perception branch, feature enhancement processing is performed on the preliminary local information of the pathological image to obtain the enhanced local information of the pathological image.

[0033] Based on the second convolutional layer and pixel attention layer in the multi-scale hierarchical perception module of the local feature perception branch, spatial feature correlation mining of pathological image pixels is performed on the preliminary local information of the pathological image to obtain the spatial features of the pathological image pixels.

[0034] Based on the third convolution and channel attention layer in the multi-scale hierarchical perception module of the local feature perception branch, the correlation of different channel features of the preliminary local information of the pathological image is mined to obtain the channel features of the pathological image.

[0035] Based on the lightweight multi-scale structure in the multi-scale hierarchical perception module of the local feature perception branch, the local information of the enhanced pathological image is mined to obtain multi-scale features of the pathological image at different scales.

[0036] The spatial features of pixels in a pathological image, the channel features of a pathological image, and the multi-scale features of a pathological image are combined to obtain the local features of the pathological image.

[0037] Furthermore, the feature interaction module in the global-local feature interaction perception network based on the pathological image staining normalization model performs feature interaction processing on the global and local features of the pathological image to obtain the autocorrelation of the global and local features of the pathological image and the cross-correlation of the global and local features of the pathological image. This step specifically includes:

[0038] The global and local features of the pathological image are input into the feature interaction module of the global-local feature interaction perception network of the pathological image staining normalization model.

[0039] Based on the first convolutional layer, the second convolutional layer, the first self-attention module, and the second self-attention module of the feature interaction module, autocorrelation mining is performed on the global and local features of the pathological image to obtain the preliminary autocorrelation of the global and local features of the pathological image.

[0040] Based on the third and fourth convolutional layers, the first and second Sigmoid functions of the feature interaction module, the autocorrelation of the preliminary global and local features of the pathological image is preprocessed to obtain the autocorrelation of the global and local features of the pathological image.

[0041] The cross-attention module based on the feature interaction module performs feature interaction between global and local features of pathological images to obtain the cross-correlation between global and local features of pathological images.

[0042] The second technical solution adopted in this invention is: a pathological image staining normalization system based on global-local feature interaction, comprising:

[0043] The first module is used to acquire pathological images;

[0044] The second module is used to introduce a lightweight multi-scale Transformer structure and a multi-scale hierarchical perception module to construct a pathological image staining normalization model.

[0045] The third module is used to perform staining normalization processing on pathological images based on the pathological image staining normalization model to obtain normalized pathological images.

[0046] The beneficial effects of the method and system of this invention are as follows: This invention acquires pathological images, then introduces a lightweight multi-scale Transformer structure and a multi-scale hierarchical perception module to construct a pathological image staining normalization model. It mines local features of the pathological image through a local feature extraction branch centered on a convolutional neural network and global feature extraction branch centered on a Transformer to mine global features. Furthermore, it achieves deep fusion of global and local features through a fusion sub-network. The lightweight multi-scale Transformer block enhances the Transformer's multi-scale global feature extraction capability while maintaining low computational complexity. A feature interaction module based on cross-attention enables interaction between global and local features, fully utilizing the complementarity and correlation of global and local features. Finally, based on the pathological image staining normalization model, it performs staining normalization processing on the pathological image to obtain a normalized pathological image. This model can comprehensively perceive pathological image features while reducing its parameter count and computational complexity, achieving staining normalization of pathological images from different centers and batches, with good performance and strong generalization. Attached Figure Description

[0047] Figure 1 This is a flowchart of the steps of the pathological image staining normalization method based on global-local feature interaction of the present invention;

[0048] Figure 2This is a structural block diagram of the pathological image staining normalization system based on global-local feature interaction of the present invention;

[0049] Figure 3 This is a schematic diagram of the pathological image staining normalization model provided in a specific embodiment of the present invention;

[0050] Figure 4 This is a schematic diagram of a multi-scale hierarchical perception module provided in a specific embodiment of the present invention;

[0051] Figure 5 This is a schematic diagram of a lightweight multi-scale Transformer block provided in a specific embodiment of the present invention;

[0052] Figure 6 This is a schematic diagram of the global-local feature interaction module provided in a specific embodiment of the present invention. Detailed Implementation

[0053] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are only for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adapted according to the understanding of those skilled in the art.

[0054] First, it's important to note that histopathological images contain tumor microenvironment information closely related to cancer diagnosis, prognosis prediction, and treatment efficacy evaluation. Therefore, histopathological examination is the gold standard for clinical cancer diagnosis. However, inconsistencies in staining patterns in histopathological images from different medical centers, caused by factors such as staining agents, scanning equipment, and scanning protocols, limit the performance of AI models for pathology images. Existing pathology image staining normalization methods often use a single convolutional neural network or Transformer as the backbone network. However, convolutional neural networks focus only on mining local features of pathology images, while Transformers focus on mining global features, making it difficult for existing models to comprehensively mine feature information at different levels of the pathology image.

[0055] Based on this, embodiments of the present invention provide a pathological image staining normalization method guided by global-local feature interaction. It combines the advantages of convolutional neural networks and Transformers to fully mine global-local features of pathological images, while the interaction of global-local features is realized by the feature interaction module. Furthermore, pathological image staining normalization is realized through a global-local feature fusion network.

[0056] Reference Figure 1 This invention provides a staining normalization method for pathological images based on global-local feature interaction, which includes the following steps:

[0057] S100, Obtain pathological images;

[0058] S200 introduces a lightweight multi-scale Transformer structure and a multi-scale hierarchical perception module to construct a pathological image staining normalization model.

[0059] Specifically, such as Figure 3 As shown, the pathological image staining normalization model specifically includes a global-local feature interaction perception network and a global-local feature fusion network. The output of the global-local feature interaction perception network is connected to the input of the global-local feature fusion network, wherein:

[0060] The global-local feature interaction perception network includes a global feature extraction branch, a local feature perception branch, and a feature interaction module. The global feature extraction branch includes a large convolutional layer and several stacked lightweight multi-scale Transformer modules. The local feature perception branch includes a shallow feature extractor and several stacked multi-scale hierarchical perception modules. The first output of the lightweight multi-scale Transformer module is connected to the input of the next stacked lightweight multi-scale Transformer module, and the second output of the lightweight multi-scale Transformer module is connected to the input of the feature interaction module. The first output of the multi-scale hierarchical perception module is connected to the input of the next stacked multi-scale hierarchical perception module, and the second output of the multi-scale hierarchical perception module is connected to the input of the feature interaction module.

[0061] More specifically, such as Figure 5As shown, the lightweight multi-scale Transformer module specifically includes a first normalization layer, a first convolutional module, a second convolutional module, a third convolutional module, a first concatenated feature mining module, a second concatenated feature mining module, a third concatenated feature mining module, a channel attention mechanism, a second normalization layer, a first convolutional layer, and a second convolutional layer. The output of the first normalization layer is connected to the inputs of the first convolutional module, the second convolutional module, and the third convolutional module, respectively. The output of the first convolutional module is connected to the input of the first concatenated feature mining module, the output of the second convolutional module is connected to the input of the second concatenated feature mining module, and the output of the third convolutional module is connected to the input of the third convolutional module. The first and second convolutional modules are connected to the input of the third splicing feature mining module. The outputs of both the first and second splicing feature mining modules are connected to the channel attention mechanism. The outputs of both the channel attention mechanism and the third splicing feature mining module are connected to the input of the second normalization layer. The output of the second normalization layer is connected to the inputs of both the first and second convolutional layers. Each of the first, second, and third convolutional modules includes several convolutional layers with different dilation rates. Each of the first, second, and third splicing feature mining modules includes a splicing layer and a compression excitation network.

[0062] More specifically, such as Figure 4 As shown, the multi-scale hierarchical perception module specifically includes a first convolutional layer, a ReLU function, a second convolutional layer, a pixel attention layer, a first global pooling layer, a first average pooling layer, a third convolutional layer, a channel attention layer, and a lightweight multi-scale structure. The output of the first convolutional layer is connected to the input of the second convolutional layer, the input of the third convolutional layer, and the input of the ReLU function. The output of the ReLU function is connected to the input of the first global pooling layer and the input of the first average pooling layer. The output of the second convolutional layer is connected to the input of the pixel attention layer. The output of the third convolutional layer is connected to the input of the channel attention layer. The outputs of the first global pooling layer and the first average pooling layer are both connected to the input of the lightweight multi-scale structure. The lightweight multi-scale structure includes several convolutional and compression / excitation network blocks with different dilation rates.

[0063] More specifically, such as Figure 6As shown, the feature interaction module specifically includes a first convolutional layer, a second convolutional layer, a first self-attention module, a second self-attention module, a third convolutional layer, a fourth convolutional layer, a first sigmoid function, a second sigmoid function, and a cross-attention module. The output of the first convolutional layer is connected to the input of the first self-attention module, the output of the second convolutional layer is connected to the input of the second self-attention module, the output of the first self-attention module is connected to the inputs of the third convolutional layer and the cross-attention module, the output of the second self-attention module is connected to the inputs of the fourth convolutional layer and the cross-attention module, the third convolutional layer is connected to the first sigmoid function, and the fourth convolutional layer is connected to the second sigmoid function.

[0064] The global-local feature fusion network includes a global pooling layer, an average pooling layer, and a post-processing module. The post-processing module includes a compression and activation network, a first post-processing block, and a second post-processing block. Both the first and second post-processing blocks include a convolutional layer, a normalization layer, and a ReLU function.

[0065] S300. Based on the pathological image staining normalization model, the pathological image is subjected to staining normalization processing to obtain the normalized pathological image.

[0066] S310. Input the pathological image into the pathological image staining normalization model;

[0067] Specifically, based on the fact that image features have different manifestations and complementarity and correlation at different scales, a global-local feature interaction perception network is constructed to analyze pathological images. A local feature perception branch is constructed by densely connected multi-scale hierarchical perception modules to mine local features of pathological images. At the same time, a global feature mining branch constructed by stacked lightweight multi-scale Transformer blocks extracts global features of pathological images. Furthermore, the feature interaction model perceives the autocorrelation and cross-correlation of global-local features of pathological images, enabling the model to fully perceive the features of pathological images and reduce its parameter count and computational complexity.

[0068] S320. The global feature extraction branch in the global-local feature interaction perception network based on the pathological image staining normalization model performs global feature information extraction processing on the pathological image to obtain the global features of the pathological image.

[0069] Specifically, the pathological image is input into the global feature extraction branch of the global-local feature interaction perception network of the pathological image staining normalization model; based on the large convolutional layer of the global feature extraction branch, feature extraction is performed on the pathological image to obtain global information of the pathological image; based on the first normalization layer in the lightweight multi-scale Transformer module of the global feature extraction branch, the global information of the pathological image is normalized to obtain normalized global information of the pathological image; based on the first convolutional module, second convolutional module, third convolutional module, first stitching feature mining module, and third convolutional module in the lightweight multi-scale Transformer module of the global feature extraction branch, the global information is normalized to obtain normalized global information of the pathological image. The second and third stitching feature mining modules mine the correlation between feature channels in the normalized global information of the pathological image to obtain the query value, key value, and value of the Transformer block. Based on the channel attention mechanism in the lightweight multi-scale Transformer module of the global feature extraction branch, the query value and key value of the Transformer block are calculated to obtain the feature map. The feature map and the value of the Transformer block are multiplied and added to the global information of the pathological image, and then input into the second normalization layer, the first convolutional layer, and the second convolutional layer for normalization and dot product calculation to obtain the global features of the pathological image.

[0070] In this embodiment, the Transformer structure can analyze the global dependencies of features through a self-attention mechanism and has strong parallel computing capabilities. However, existing Transformer structures mostly only consider single-scale pathological image features and have weak multi-scale feature representation capabilities. Therefore, this embodiment of the invention provides a global feature extraction branch with a lightweight multi-scale Transformer block as its core to mine multi-scale global features of pathological images. Specifically,

[0071] The global feature extraction branch is centered around stacked lightweight multi-scale Transformer blocks. It first extracts more global information at once using a large 27×27 convolutional kernel, and then mines long-distance dependencies between pixels through stacked lightweight multi-scale Transformer blocks to comprehensively uncover global features of the pathological image. The multi-scale lightweight Transformer blocks first normalize the global information extracted by the 27×27 large convolutional kernel, and then concatenate different dilation rates. The 3×3 convolution extracts features at different scales, and the compression and activation of the network block fully utilizes the correlation between feature channels to generate the query value (Q), key value (K), and value (V) of the Transformer block. The expression is as follows: ; in, Indicates different expansion rates 3×3 convolution, This indicates a normalization operation. This represents a 27×27 large convolution kernel. This represents the compression and activation of network blocks. Subsequently, channel attention is used to process the query value (Q) and key value (K) to obtain the feature map. Its expression is: ; in, This represents channel attention. Next, the feature map is multiplied by the value (V) and added to the global information extracted by the large 27×27 convolutional kernel. The expression is: ; Finally, intermediate features The system utilizes 3×3 convolutional layers with dilation rates of 1 and 2, followed by normalization and dot product. This data is then added to intermediate features to obtain the final global pathological features. Its expression is: ; in, and This represents a 3×3 convolution with dilation rates of 1 and 2.

[0072] S330. The local feature perception branch in the global-local feature interaction perception network based on the pathological image staining normalization model performs local feature information extraction processing on the pathological image to obtain the local features of the pathological image.

[0073] Specifically, the pathological image is input into the local feature perception branch of the global-local feature interaction perception network of the pathological image staining normalization model; based on the shallow feature extractor of the local feature perception branch, shallow feature extraction processing is performed on the pathological image to obtain shallow features of the pathological image; based on the first convolutional layer of the multi-scale hierarchical perception module of the local feature perception branch, convolution processing is performed on the shallow features of the pathological image to obtain preliminary local information of the pathological image; based on the ReLU function, the first global pooling layer and the first average pooling layer in the multi-scale hierarchical perception module of the local feature perception branch, feature enhancement processing is performed on the preliminary local information of the pathological image to obtain enhanced local information of the pathological image; based on the multi-scale hierarchical perception module of the local feature perception branch... The second convolutional layer and pixel attention layer perform spatial feature correlation mining on the preliminary local information of the pathological image pixels to obtain the spatial features of the pathological image pixels. The third convolutional layer and channel attention layer in the multi-scale hierarchical perception module based on the local feature perception branch perform correlation mining on the different channel features of the preliminary local information of the pathological image to obtain the channel features of the pathological image. The lightweight multi-scale structure in the multi-scale hierarchical perception module based on the local feature perception branch mines multi-scale features of different scale spaces on the enhanced local information of the pathological image to obtain the multi-scale features of the pathological image. The spatial features of the pathological image pixels, the channel features of the pathological image, and the multi-scale features of the pathological image are merged to obtain the local features of the pathological image.

[0074] In this embodiment, convolutional neural networks have strong local feature extraction and perception capabilities, but their multi-scale representation capabilities are weak and the number of model parameters is large. Therefore, this embodiment of the invention provides a local feature perception branch centered on a multi-scale hierarchical perception module to mine local features in pathological images.

[0075] First, shallow features of pathological images are extracted using consecutive 3×3 and 1×1 convolutions. Its expression is: ; in, This represents the input pathological image. and These represent 3×3 and 1×1 convolutions, respectively. Further, pixel attention and spatial attention are introduced to mine the spatial and feature channel correlations of pixels in pathological images. Simultaneously, a lightweight multi-scale structure is introduced to enhance the multi-scale representation capabilities of the model, constructing a multi-scale hierarchical perception module. This module has a three-branch structure. The top branch contains a 3×3 convolution and a pixel spatial attention module to mine the spatial feature correlations of pixels; the bottom branch contains a 3×3 convolution and a channel attention module to fully explore the correlations of features from different channels. This process can be represented as: ; in, and These represent the features extracted from the top and bottom branches, respectively. This represents the pixel-space attention module. This indicates the channel attention module.

[0076] In the middle branch, lightweight multi-scale structure-aware pathological image multi-scale features are introduced. First, downsampling is used to reduce the dimensionality of shallow features, and the ReLU function is used to enhance nonlinear representation. Global pooling and average pooling are also used to enhance feature representation. Then, different dilation rates are utilized... The Lightweight Multiscale Structure (LMS2) is constructed using 3×3 convolutional and compression / activation network blocks to mine multiscale features of pathological images at different scales. Pathological features extracted from the low-dilation-rate branch are sequentially fed into the high-dilation-rate branch. To fully utilize the complementarity of features at different scales, its expression is: ; in, and These represent downsampling and upsampling operations, respectively. Represents the ReLU activation function. This represents global pooling and average pooling. This represents a lightweight multi-scale structural module. Finally, the features extracted from the three branches are merged using pixel multiplication, and shallow features are further embedded to obtain local features of the pathological image. ,Right now ; in, This indicates pixel multiplication. This indicates pixel addition.

[0077] S340. The feature interaction module in the global-local feature interaction perception network based on the pathological image staining normalization model performs feature interaction processing on the global features and local features of the pathological image to obtain the autocorrelation of the global features and local features of the pathological image and the cross-correlation of the global-local features of the pathological image.

[0078] Specifically, the global and local features of the pathological image are input into the feature interaction module of the global-local feature interaction perception network of the pathological image staining normalization model. Based on the first convolutional layer, the second convolutional layer, the first self-attention module, and the second self-attention module of the feature interaction module, autocorrelation mining is performed on the global and local features of the pathological image to obtain the preliminary autocorrelation between the global and local features of the pathological image. Based on the third and fourth convolutional layers, the first sigmoid function, and the second sigmoid function of the feature interaction module, the preliminary autocorrelation between the global and local features of the pathological image is preprocessed to obtain the autocorrelation between the global and local features of the pathological image. Based on the cross-attention module of the feature interaction module, feature interaction is performed on the global and local features of the pathological image to obtain the cross-correlation between the global and local features of the pathological image.

[0079] In this embodiment, global features are a macroscopic, generalized description of the pathological image, covering the entire image and weakly dependent on local locations; local features are precise depictions of the microscopic details (such as texture details and edge contours) of specific regions of the pathological image, strongly dependent on local locations and with higher information density. Therefore, this embodiment provides a global-local feature interaction module for pathological images. This module uses a self-attention mechanism to mine feature autocorrelation and cross-attention to mine feature cross-correlation, enabling information interaction between the local feature perception branch and the global feature extraction branch. This fully explores the correlation and complementarity of global and local features at different levels. Specifically,

[0080] First, global and local features extracted from the two branches are processed using 3×3 convolutions and a self-attention module to uncover feature autocorrelation. This is further enhanced by continuous 3×3 convolutions and the sigmoid function. Furthermore, the features extracted by the self-attention module are processed through a cross-attention module to achieve interaction between global and local features. The entire process can be represented as follows: ; in, Indicates the first Local features extracted by a multi-scale hierarchical perception module Indicates the first Global features extracted from a lightweight multi-scale Transformer block and These represent the cross-attention and self-attention mechanisms, respectively. Subsequently, the output of the cross-attention module is embedded into the two branches through pixel multiplication, and further added to the global-local features processed by 3×3 convolution to obtain the final result, i.e. ; In the above formula, The output represents the autocorrelation of global and local features in a pathological image. This indicates the cross-correlation between global and local features in pathological images.

[0081] S350. The global pooling layer and average pooling layer in the global-local feature fusion network based on the pathological image staining normalization model pool and concatenate the autocorrelation of global and local features of the pathological image and the cross-correlation of global and local features of the pathological image to obtain the filtered pathological image.

[0082] S360, a post-processing module in a global-local feature fusion network based on a pathological image staining normalization model, explores and normalizes the correlation between different feature channels of the filtered pathological image to obtain a normalized pathological image.

[0083] Therefore, this embodiment of the invention first mines local features of pathological images at different scales through a local feature perception branch centered on a densely connected multi-scale hierarchical perception module, and simultaneously mines global features of pathological images through a global feature extraction branch centered on a lightweight multi-scale Transformer block. Furthermore, it mines the autocorrelation and cross-correlation between global and local features of the pathological images through a feature interaction module. The global and local features of the pathological images are then simultaneously processed by global pooling and average pooling before being concatenated and fed into a lightweight post-processing module to fully exploit the correlation and complementarity between global and local features of the pathological images. This achieves staining normalization of pathological images from different centers and batches, exhibiting good performance and strong generalization.

[0084] In this embodiment, global-local feature fusion can simultaneously capture both the overall and detailed information of pathological images, thereby significantly improving the performance of pathological image staining normalization models. To this end, this embodiment of the invention provides a global-local feature fusion subnetwork that effectively aggregates global and local features of pathological images, effectively mitigating staining inconsistencies between different batches and different centers of pathological images while better preserving local image details. Specifically,

[0085] The global-local features extracted in step one are fed into the fusion sub-network to aggregate features from different levels, thereby enhancing the feature representation capability of the pathological image staining normalization model. The global-local features are simultaneously processed by global pooling and average pooling before being concatenated to preserve texture information while filtering out noise. The expression is as follows: ; Subsequently, the correlation between feature channels was analyzed using compressed and activated network blocks, and the normalized pathological image was obtained through post-processing operations including 3×3 convolutions, normalization layers, and ReLU activation functions. Its expression is: ; in, Indicates post-processing operation. This represents a pathological image after staining normalization.

[0086] In summary, the embodiments of the present invention have the following advantages over the prior art:

[0087] 1) Existing pathological image staining normalization methods mostly use convolutional neural networks (CNNs) or Transformer structures as their core architecture. However, CNNs focus more on local image features, while Transformers focus more on long-distance dependencies of image features to mine global image features. This invention combines the advantages of CNNs and Transformers to design a pathological image staining normalization method guided by global-local feature interaction. It mines local features of pathological images through a local feature extraction branch with a CNN as the core, and mines global features of pathological images through a global feature extraction branch with a Transformer as the core. Furthermore, it achieves deep fusion of global and local features through a fusion sub-network.

[0088] 2) Existing pathological image staining normalization methods mostly process pathological images in a single-scale space, ignoring the different representations of image features at different scales. However, this invention enhances the multi-scale extraction and hierarchical representation capabilities of convolutional neural networks by densely connecting multi-scale hierarchical perception modules; it also improves the multi-scale global feature extraction capabilities of Transformers while maintaining low computational complexity by using lightweight multi-scale Transformer blocks. Furthermore, it achieves global-local feature interaction through a cross-attention-based feature interaction module to fully utilize the complementarity and correlation between global and local features.

[0089] Reference Figure 2 A pathological image staining normalization system based on global-local feature interaction includes:

[0090] The first module 201 is used to acquire pathological images;

[0091] The second module 202 is used to introduce a lightweight multi-scale Transformer structure and a multi-scale hierarchical perception module to construct a pathological image staining normalization model.

[0092] The third module 203 is used to perform staining normalization processing on pathological images based on the pathological image staining normalization model to obtain normalized pathological images.

[0093] The content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0094] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A pathological image staining normalization method based on global-local feature interaction, characterized in that, Includes the following steps: Acquire pathological images; A lightweight multi-scale Transformer structure and a multi-scale hierarchical perception module are introduced to construct a staining normalization model for pathological images. Based on the pathological image staining normalization model, the pathological images are subjected to staining normalization processing to obtain normalized pathological images.

2. The pathological image staining normalization method based on global-local feature interaction according to claim 1, characterized in that, The pathological image staining normalization model specifically includes a global-local feature interaction perception network and a global-local feature fusion network. The output of the global-local feature interaction perception network is connected to the input of the global-local feature fusion network, wherein: The global-local feature interaction perception network includes a global feature extraction branch, a local feature perception branch, and a feature interaction module. The global feature extraction branch includes a large convolutional layer and several stacked lightweight multi-scale Transformer modules. The local feature perception branch includes a shallow feature extractor and several stacked multi-scale hierarchical perception modules. The first output of the lightweight multi-scale Transformer module is connected to the input of the next stacked lightweight multi-scale Transformer module, and the second output of the lightweight multi-scale Transformer module is connected to the input of the feature interaction module. The first output of the multi-scale hierarchical perception module is connected to the input of the next stacked multi-scale hierarchical perception module, and the second output of the multi-scale hierarchical perception module is connected to the input of the feature interaction module. The global-local feature fusion network includes a global pooling layer, an average pooling layer, and a post-processing module. The post-processing module includes a compression and activation network, a first post-processing block, and a second post-processing block. Both the first and second post-processing blocks include a convolutional layer, a normalization layer, and a ReLU function.

3. The pathological image staining normalization method based on global-local feature interaction according to claim 2, characterized in that, The lightweight multi-scale Transformer module specifically includes a first normalization layer, a first convolutional module, a second convolutional module, a third convolutional module, a first concatenated feature mining module, a second concatenated feature mining module, a third concatenated feature mining module, a channel attention mechanism, a second normalization layer, a first convolutional layer, and a second convolutional layer. The output of the first normalization layer is connected to the inputs of the first convolutional module, the second convolutional module, and the third convolutional module, respectively. The output of the first convolutional module is connected to the input of the first concatenated feature mining module, the output of the second convolutional module is connected to the input of the second concatenated feature mining module, and the output of the third convolutional module is connected to... The input of the third splicing feature mining module is connected, and the outputs of the first and second splicing feature mining modules are both connected to the channel attention mechanism. The outputs of the channel attention mechanism and the third splicing feature mining module are both connected to the input of the second normalization layer. The output of the second normalization layer is connected to the inputs of the first and second convolutional layers, respectively. The first, second, and third convolutional modules each include several convolutional layers with different dilation rates. The first, second, and third splicing feature mining modules each include a splicing layer and a compression excitation network.

4. The pathological image staining normalization method based on global-local feature interaction according to claim 3, characterized in that, The multi-scale hierarchical perception module specifically includes a first convolutional layer, a ReLU function, a second convolutional layer, a pixel attention layer, a first global pooling layer, a first average pooling layer, a third convolutional layer, a channel attention layer, and a lightweight multi-scale structure. The output of the first convolutional layer is connected to the input of the second convolutional layer, the input of the third convolutional layer, and the input of the ReLU function. The output of the ReLU function is connected to the input of the first global pooling layer and the input of the first average pooling layer. The output of the second convolutional layer is connected to the input of the pixel attention layer. The output of the third convolutional layer is connected to the input of the channel attention layer. The outputs of the first global pooling layer and the first average pooling layer are both connected to the input of the lightweight multi-scale structure. The lightweight multi-scale structure includes several convolutional and compression / excitation network blocks with different dilation rates.

5. The pathological image staining normalization method based on global-local feature interaction according to claim 4, characterized in that, The feature interaction module specifically includes a first convolutional layer, a second convolutional layer, a first self-attention module, a second self-attention module, a third convolutional layer, a fourth convolutional layer, a first sigmoid function, a second sigmoid function, and a cross-attention module. The output of the first convolutional layer is connected to the input of the first self-attention module, the output of the second convolutional layer is connected to the input of the second self-attention module, the output of the first self-attention module is connected to the inputs of the third convolutional layer and the cross-attention module, the output of the second self-attention module is connected to the inputs of the fourth convolutional layer and the cross-attention module, the third convolutional layer is connected to the first sigmoid function, and the fourth convolutional layer is connected to the second sigmoid function.

6. The pathological image staining normalization method based on global-local feature interaction according to claim 5, characterized in that, The step of performing staining normalization processing on pathological images based on a pathological image staining normalization model to obtain normalized pathological images specifically includes: Input the pathological images into the pathological image staining normalization model; The global feature extraction branch in the global-local feature interaction perception network based on the pathological image staining normalization model is used to extract global feature information from the pathological image to obtain the global features of the pathological image. Based on the local feature perception branch in the global-local feature interaction perception network of the pathological image staining normalization model, local feature information is extracted from the pathological image to obtain the local features of the pathological image. The feature interaction module in the global-local feature interaction perception network based on the pathological image staining normalization model performs feature interaction processing on the global features and local features of the pathological image to obtain the autocorrelation of the global features and local features of the pathological image and the cross-correlation of the global-local features of the pathological image. The global pooling layer and average pooling layer in the global-local feature fusion network based on the pathological image staining normalization model pool and concatenate the autocorrelation of global and local features of the pathological image and the cross-correlation of global and local features of the pathological image to obtain the filtered pathological image. The post-processing module in the global-local feature fusion network based on the pathological image staining normalization model explores and normalizes the correlation between different feature channels of the filtered pathological image to obtain the normalized pathological image.

7. The pathological image staining normalization method based on global-local feature interaction according to claim 6, characterized in that, The global feature extraction branch in the global-local feature interaction perception network based on the pathological image staining normalization model performs global feature information extraction processing on the pathological image to obtain the global features of the pathological image. This step specifically includes: The pathological image is input into the global feature extraction branch of the global-local feature interaction perception network of the pathological image staining normalization model; Based on a large convolutional layer with a global feature extraction branch, feature extraction is performed on pathological images to obtain global information of the pathological images; The first normalization layer in the lightweight multi-scale Transformer module based on the global feature extraction branch normalizes the global information of the pathological image to obtain the normalized global information of the pathological image. The first convolutional module, second convolutional module, third convolutional module, first stitching feature mining module, second stitching feature mining module, and third stitching feature mining module in the lightweight multi-scale Transformer module based on the global feature extraction branch mine the correlation between feature channels of the normalized pathological image global information to obtain the query value, key value, and value of the Transformer block. The channel attention mechanism in the lightweight multi-scale Transformer module based on the global feature extraction branch calculates the query value and key value of the Transformer block to obtain the feature map. The feature map and the value of the Transformer block are multiplied and added to the global information of the pathological image. The result is then fed into the second normalization layer, the first convolutional layer, and the second convolutional layer for normalization and dot product calculations to obtain the global features of the pathological image.

8. The pathological image staining normalization method based on global-local feature interaction according to claim 7, characterized in that, The local feature perception branch in the global-local feature interaction perception network based on the pathological image staining normalization model extracts local feature information from the pathological image to obtain the local features of the pathological image. This step specifically includes: The pathological image is input into the local feature perception branch of the global-local feature interaction perception network of the pathological image staining normalization model. A shallow feature extractor based on a local feature perception branch is used to perform shallow feature extraction processing on pathological images to obtain shallow features of the pathological images. The first convolutional layer in the multi-scale hierarchical perception module based on the local feature perception branch performs convolution processing on the shallow features of the pathological image to obtain preliminary local information of the pathological image. Based on the ReLU function, the first global pooling layer and the first average pooling layer in the multi-scale hierarchical perception module of the local feature perception branch, feature enhancement processing is performed on the preliminary local information of the pathological image to obtain the enhanced local information of the pathological image. Based on the second convolutional layer and pixel attention layer in the multi-scale hierarchical perception module of the local feature perception branch, spatial feature correlation mining of pathological image pixels is performed on the preliminary local information of the pathological image to obtain the spatial features of the pathological image pixels. Based on the third convolution and channel attention layer in the multi-scale hierarchical perception module of the local feature perception branch, the correlation of different channel features of the preliminary local information of the pathological image is mined to obtain the channel features of the pathological image. Based on the lightweight multi-scale structure in the multi-scale hierarchical perception module of the local feature perception branch, the local information of the enhanced pathological image is mined to obtain multi-scale features of the pathological image at different scales. The spatial features of pixels in a pathological image, the channel features of a pathological image, and the multi-scale features of a pathological image are combined to obtain the local features of the pathological image.

9. The pathological image staining normalization method based on global-local feature interaction according to claim 8, characterized in that, The feature interaction module in the global-local feature interaction perception network based on the pathological image staining normalization model performs feature interaction processing on the global and local features of the pathological image to obtain the autocorrelation of the global and local features of the pathological image and the cross-correlation of the global and local features of the pathological image. This step specifically includes: The global and local features of the pathological image are input into the feature interaction module of the global-local feature interaction perception network of the pathological image staining normalization model. Based on the first convolutional layer, the second convolutional layer, the first self-attention module, and the second self-attention module of the feature interaction module, autocorrelation mining is performed on the global and local features of the pathological image to obtain the preliminary autocorrelation of the global and local features of the pathological image. Based on the third and fourth convolutional layers, the first and second Sigmoid functions of the feature interaction module, the autocorrelation of the preliminary global and local features of the pathological image is preprocessed to obtain the autocorrelation of the global and local features of the pathological image. The cross-attention module based on the feature interaction module performs feature interaction between global and local features of pathological images to obtain the cross-correlation between global and local features of pathological images.

10. A pathological image staining normalization system based on global-local feature interaction, characterized in that, Includes the following modules: The first module is used to acquire pathological images; The second module is used to introduce a lightweight multi-scale Transformer structure and a multi-scale hierarchical perception module to construct a pathological image staining normalization model. The third module is used to perform staining normalization processing on pathological images based on the pathological image staining normalization model to obtain normalized pathological images.