Weak supervision tissue pathology image segmentation method based on multi-scale gating fusion
Through the weakly supervised histopathological image segmentation method of multi-scale gating fusion, the problem of pathological image segmentation relies on manual annotation and low accuracy in the prior art is solved. Through multi-level feature extraction and gating fusion, the accuracy and practicality of pathological image segmentation are improved.
Patent Information
- Application Number
- CN202510261482.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-29
AI Technical Summary
The existing pathological image segmentation methods rely on high-cost manual annotation, and the unsupervised learning methods have low accuracy, making it difficult to meet the needs of pathologists for cross-regional information. The existing multi-instance learning methods ignore the connection between hierarchical and global features and fail to make full use of WSI information.
The weakly supervised tissue pathological image segmentation method of multi-scale gating fusion is adopted. Through the multi-level feature extraction module and the gated adaptive fusion module, multi-level features are extracted from different resolutions, and feature information at different levels is dynamically fused through the gating mechanism to capture global semantic connections and realize image segmentation.
It significantly improves the accuracy and practicality of pathological image segmentation, especially when processing complex pathological images, improves the accuracy of pixel-level prediction and the effectiveness of segmentation tasks.
Smart Images

Figure CN120388171A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic information technology, and more specifically, to a weakly supervised tissue pathology image segmentation method based on multi-scale gated fusion. Background Art
[0002] Cancer pathology image analysis is a key research area in computer vision. Thanks to advancements in scanning technology, whole slide imaging (WSI) has become widely used for the storage and analysis of pathology images. Deep learning-based pathology image segmentation technology, by processing, classifying, and segmenting WSI, can provide crucial support for doctors' diagnosis and treatment. This technology holds broad application prospects in cancer diagnosis, pathology analysis, and other fields.
[0003] Supervised learning methods for pathological image segmentation rely on large amounts of high-quality, manually annotated data and have achieved significant progress in WSI segmentation. However, a major challenge for these methods lies in the need for precise annotation of target regions, which is time-consuming and costly for pathologists to perform. In contrast, unsupervised learning methods, while reducing the reliance on annotated data, remain difficult to widely apply in real-world scenarios due to their lower accuracy.
[0004] In recent years, multiple instance learning (MIL) based on WSI has become a research hotspot. MIL is a subclass of weakly supervised learning. Existing MIL methods often rely on single-scale analysis, neglecting the connection between hierarchical and global features and failing to fully explore the relationships between instances within the same bag. Although convolutional neural networks (CNNs) introduce local dependencies, they still fail to meet the pathologist's need for cross-regional information. Therefore, capturing the global relationships between instances within a bag is crucial for improving MIL performance. Some studies have begun to explore the fusion of multi-scale WSI information, demonstrating that combining multi-scale features can significantly improve the effectiveness of WSI analysis. For example, some schemes process WSI image patches of different resolutions in two substreams and then concatenate their features. Some schemes propose a hybrid multi-instance learning model based on Transformer and graph attention networks. However, these methods often ignore the semantic differences between different scales and fail to fully utilize the global feature comparison capabilities of WSI.
[0005] In summary, existing MIL methods are usually limited to single-scale WSI analysis. Although this method simplifies the calculation process, it also ignores the hierarchical information within the WSI and the correlation between global features. At the same time, existing methods overly focus on instance-level prediction and fail to fully explore the potential connections between instances within the same bag. Although MIL methods based on convolutional neural networks (CNNs) introduce local dependencies through convolutional operations, they still struggle to meet the pathologists' need for cross-region information correlation during the diagnosis process. Summary of the Invention
[0006] The objective of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a weakly supervised tissue pathology image segmentation method based on multi-scale gated fusion. This aspect includes the following steps:
[0007] Obtain a pathology image;
[0008] Input the pathology image into an image segmentation model to obtain an image segmentation result;
[0009] Among them, the image segmentation model includes a multi-level feature extraction module and a gated adaptive fusion module. The multi-level feature extraction module is used to extract corresponding multi-level features from different resolutions for the pathology image, and the gated adaptive fusion module is used to fuse the multi-level features based on a gating mechanism to obtain global features, and then predict pixel categories based on the global features to achieve the image segmentation task.
[0010] Compared with the existing technologies, the advantages of the present invention are as follows. In response to the requirements of application scenarios such as cancer diagnosis, pathology analysis, and image segmentation, a weakly supervised tissue pathology image segmentation network based on multi-scale gated fusion is designed. By continuously increasing the number of convolutional layers, the number of channels, and self-attention operations in the processing flow, hierarchical features are extracted, capturing boundary details and global semantic connections. And through gated adaptive fusion, the feature information of different levels is fully integrated, thereby improving the accuracy of cancer tissue segmentation.
[0011] Through the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings, other features and advantages of the present invention will become clear. Brief Description of the Drawings
[0012] The accompanying drawings incorporated in and constituting a part of this specification illustrate embodiments of the present invention and, together with the description, are used to explain the principles of the present invention.
[0013] Figure 1 is a schematic process diagram of a weakly supervised tissue pathology image segmentation method based on multi-scale gated fusion according to an embodiment of the present invention;
[0014] Figure 2It is a flowchart of a weakly supervised histopathological image segmentation method based on multi-scale gated fusion according to an embodiment of the present invention;
[0015] Figure 3 It is a flowchart of multi-level feature extraction according to an embodiment of the present invention;
[0016] Figure 4 It is a schematic diagram of a gated adaptive fusion model according to an embodiment of the present invention; [[ID=,10]]
[0017] Figure 5 It is a schematic diagram of multi-level feature decoding and prediction according to an embodiment of the present invention. Detailed implementation manners
[0018] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present invention.
[0019] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way limits the present invention, its application or use.
[0020] Technologies, methods and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods and devices should be regarded as part of the specification.
[0021] In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values.
[0022] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.
[0023] The present invention proposes a weakly supervised histopathological image segmentation method based on multi-scale gated fusion. Taking whole slide images (WSIs) as an example, each WSI is regarded as a bag, and each pixel in the WSI is regarded as an instance, thereby transforming the weakly supervised semantic segmentation problem based on image-level labels into an instance prediction task based on bag-level labels. Moreover, a multi-scale feature fusion strategy is introduced to focus on the local and global information of WSIs at different resolutions. By extracting hierarchical information from local regions to the entire bag layer by layer, more accurate segmentation results are ensured. In addition, a gated mechanism is adopted to dynamically fuse features at different scales, enhancing the ability to model information at each level, thereby improving the accuracy of pixel-level prediction.
[0024] See Figure 1 As shown, generally speaking, the provided weakly supervised histopathological image segmentation method based on multi-scale gated fusion includes: First, extract low-level features from the pathological image, and respectively input them into the decoder through the attention mechanism, and increase the number of convolutional layers and channels. The decoder outputs the decoded features, and the middle-level features after increasing the channels are obtained; then, perform the same processing on the middle-level features as the low-level features to obtain high-level features, and perform similar operations; subsequently, respectively obtain the decoded low-level, middle-level, and high-level features from the outputs of the three decoders, and perform multi-level feature fusion through the gated adaptive fusion model to obtain global features; finally, perform pixel category prediction through the global features to complete the image segmentation task.
[0025] Specifically, in combination with Figure 1 and Figure 2 As shown, the provided weakly supervised histopathological image segmentation method based on multi-scale gated fusion includes the following steps:
[0026] Step S1, construct an image segmentation model, which includes a multi-level feature extraction module and a gated adaptive fusion module.
[0027] The multi-level feature extraction module is used to extract corresponding multi-level features (or multi-scale features) from different resolutions for the pathological image, and the number of levels can be set according to requirements such as accuracy and efficiency. The gated adaptive fusion module is used to fuse multi-level features based on the gated mechanism to obtain global features.
[0028] Step S2, take the pathological image as the input, and use the multi-level feature extraction module to extract multi-level features.
[0029] Different from existing research that mainly focuses on single-scale pathological image analysis, the multi-level feature extraction module of the present invention extracts the information contained in the image from multiple scales by gradually increasing the convolutional depth and channels. For example, Figure 3 Taking the extraction of features in three stages (Stage) as an example, where Conv(x,y) represents that the input has x channels and the output has y channels, and Decoder represents the decoder.
[0030] Stage 1: The main task of this stage is to extract low-level detail features from the input image, such as basic information such as edges and textures. The convolutional kernel with 64 channels can capture relatively fine image features. This stage includes two convolutional layers, the input is 3 channels (for example, an RGB image), and the output is 64 channels. The decoder output x1 of Stage 1 can be used for prediction at the detail level to capture the basic features of the image.
[0031] Stage 2: This stage is responsible for extracting more complex features. The convolutional kernels with 128 channels can capture more and more local pattern or structural information and start to focus on higher-level semantic content. This stage contains two convolutional layers, with an input of 64 channels and an output of 128 channels. The decoder output x2 of Stage 2 helps to make more refined predictions of local morphology, focusing on the specific structures and local information in the image.
[0032] Stage 3: This stage extracts more advanced features. The convolutional kernels with 256 channels can identify more complex patterns and semantic information, and are usually used to process deeper context information. This stage contains three convolutional layers, with an input of 128 channels and an output of 256 channels. The decoder output x3 of Stage 3 helps to make global-level predictions and capture the high-level semantic features of the entire image.
[0033] It should be noted that the number of levels, as well as the number of channels, the number of convolutional layers, and the size of the convolutional kernels in each level (or stage), can all be adjusted according to actual needs.
[0034] Step S3: For the multi-level features extracted, use the gated adaptive fusion module to fuse them to obtain global features, and then predict the pixel categories based on these global features to achieve the image segmentation task.
[0035] The present invention designs a novel multi-level feature fusion method. Considering the differences and complementarities of different-level features in terms of details and overall morphology, in one embodiment, the fusion strategy learns and dynamically adjusts the contribution of each decoder output (outputs from different levels) to the final result through a gating mechanism, thereby enhancing the flexibility and performance of the model.
[0036] Figure 4 This is the architecture of the gated adaptive fusion module (or gated adaptive fusion model). The shapes of the outputs x1, x2, and x3 of the three decoders are all (B x C x H x W), where B represents the number of training images in each batch, C is the number of channels of the image, and H and W represent the height and width of the image respectively. First, these three outputs are concatenated into a new tensor to obtain the concatenated feature map. Then, global average pooling is performed on the concatenated feature map, and weights for each decoder output are generated through convolutional operations. These weights control the contribution degree of each decoder to the final fusion result. Next, the weights are processed by softmax to ensure that the sum of all weights is 1, and finally, the gating weights gates are obtained.
[0037] Specifically, the outputs of each decoder are fused together through weighted summation, expressed as:
[0038]
[0039] Among them, gates is a gated weight tensor used to determine the contribution degrees of different x. [:,0:1,:,:] represents the weight of the 0th channel of gates in the second dimension (channel dimension), that is, the weight of x1; [:,1:2,:,:] represents the weight of the 1st channel selected, which is the weight of x2; [:,2:3,:,:] selects the weight of the 2nd channel, representing the weight of x3.
[0040] The outputs of each decoder are weighted and fused according to their corresponding gated weights, and finally the fused result x_fused is obtained as the global feature. Furthermore, based on this global feature, the pixel category can be predicted to achieve the image segmentation task.
[0041] After training the above - constructed image segmentation model using the pathological image dataset, it can be used for actual image segmentation.
[0042] To further verify the effect of the present invention, the performance of the image segmentation model was evaluated on widely used pathological image datasets. For example, on the Camelyon 16 and TCGALC datasets, the present invention was compared with the prior art. For a comprehensive comparison, in the experiment, it was compared with a variety of existing methods, including weakly - supervised multi - instance learning (MIL) methods and supervised methods. See the comparison experiment results in Table 1. The experimental results show that although the existing supervised - based methods show the best performance, due to the lack of finely labeled data, their feasibility in practical applications is limited.
[0043] Table 1 Comparison experiment results
[0044]
[0045] As can be seen from Table 1, the F1 score of the present invention (the closer the F1 value is to 1, the higher the model prediction accuracy; the closer it is to 0, the lower the prediction accuracy) is only lower than the best method, and this result is obtained under the weakly - supervised setting, surpassing other weakly - supervised methods.
[0046] Figure 5 The visualization of the prediction results of randomly selected slices in the Camelyon 16 dataset is shown, and the segmentation effects of different - level features are compared. It can be clearly seen that as the level increases, the segmentation results gradually shift from focusing on details to focusing on the overall contour, showing a gradual transition from local features to global semantic information. The experimental results show that the multi - scale gated fusion strategy of the present invention can effectively extract key features of different scales, significantly improve the accuracy of image segmentation, and achieve excellent performance on multiple pathological image datasets, proving its practicality and effectiveness in the segmentation task.
[0047] In summary, compared with the prior art, the present invention has the following advantages:
[0048] 1) By treating each WSI (Whole Slide Image) as a bag and each pixel within the WSI as an instance, the present invention transforms the weakly supervised semantic segmentation problem of image-level labels into an instance prediction task based on bag-level labels. By introducing a multi-scale feature fusion strategy, it focuses on the local and global features of the WSI, extracts hierarchical information from different resolutions, and ensures more accurate segmentation results.
[0049] 2) The present invention adopts a gating mechanism to dynamically fuse features at different scales, effectively enhancing the ability to model information at each level. By dynamically adjusting the contribution of features at each scale to the final prediction, it improves the accuracy of pixel-level prediction, significantly enhances the accuracy of the segmentation task, especially the effectiveness and practicality in processing complex pathological images.
[0050] In summary, the weakly supervised histopathological image segmentation method based on multi-scale gating fusion proposed by the present invention focuses on the local details and global semantic features of the WSI at different resolutions, extracts hierarchical information from the local region to the entire image layer by layer, thereby ensuring more accurate segmentation results. In addition, a gating mechanism is adopted to dynamically fuse features at different scales. The gating mechanism dynamically adjusts the weights according to the contribution degree of each decoder output, enhances the ability to model information at each level, and thus improves the accuracy of pixel-level prediction. Compared with the simple multi-resolution feature aggregation weighted summation method, the present invention can more flexibly and precisely combine information at different resolutions by learning and adjusting the dynamic weights in the feature fusion process, significantly improving the classification accuracy. Especially when processing complex pathological images, it effectively improves the overall performance and robustness of image segmentation.
[0051] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present invention.
[0052] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0053] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0054] The computer program instructions for carrying out the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.
[0055] Aspects of the present invention are described herein with reference to the flowchart and / or block diagram of a method, apparatus (system), and computer program product according to embodiments of the present invention. It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.
[0056] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0057] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0058] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are equivalent.
[0059] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.
Claims
1. A weakly supervised histopathological image segmentation method based on multi-scale gated fusion, comprising the following steps: Obtain a pathological image; Input the pathological image into an image segmentation model to obtain an image segmentation result; Wherein, the image segmentation model includes a multi-level feature extraction module and a gated adaptive fusion module. The multi-level feature extraction module is used to extract corresponding multi-level features from different resolutions for the pathological image, and the gated adaptive fusion module is used to fuse the multi-level features based on a gated mechanism to obtain global features, and then predict pixel categories based on the global features to achieve the image segmentation task.
2. The method according to claim 1, wherein The multi-level feature extraction module includes a first-layer feature extraction sub-module, a second-layer feature extraction sub-module, and a third-layer feature extraction sub-module with the number of channels increasing in sequence. The first-layer feature extraction sub-module is used to extract the underlying features at the detail level from the pathological image, and after passing through an attention mechanism and a first decoder, output the decoded underlying features; The second-layer feature extraction sub-module is used to extract the structure and local information of the pathological image to obtain middle-level features, and after passing through an attention mechanism and a second decoder, output the decoded middle-level features; The third-layer feature extraction sub-module is used to extract high-level semantic features in the pathological image to obtain high-level features, and after passing through an attention mechanism and a third decoder, output the decoded high-level features.
3. The method according to claim 2, wherein The first-layer feature extraction sub-module includes two convolutional layers, with an input of 3 channels and an output of 64 channels; The second-layer feature extraction sub-module includes two convolutional layers, with an input of 64 channels and an output of 128 channels; The third-layer feature extraction sub-module includes three convolutional layers, with an input of 128 channels and an output of 256 channels.
4. The method according to claim 2, wherein The gated adaptive fusion module is used to perform: concatenate the outputs of the first decoder, the second decoder, and the third decoder into a new tensor to obtain a concatenated feature map; perform global average pooling on the concatenated feature map, and generate weights for the outputs of each decoder through convolution operations; Perform softmax processing on the obtained weights to obtain gated weights; Fuse the outputs of the first decoder, the second decoder, and the third decoder based on the gated weights to obtain the global features.
5. The method according to claim 4, wherein The global features are represented as: x_fused = gates[:, 0:1, :, :] * x1 + gates[:, 1:2, :, :] * x2 + gates[:, 2:3, :, :] * x3 + Wherein, gates[:, 0:1, :, :], gates[:, 1:2, :, :], and gates[:, 2:3, :, :] are the weights of the corresponding terms, x1 is the output of the first decoder, x2 is the output of the second decoder, and x3 is the output of the third decoder.
6. The method according to claim 1, wherein The pathological image is a whole-slide pathological image.
7. The method according to claim 1, characterized in that The input of the first-layer feature extraction sub-module is an RGB pathological image.
8. The method according to claim 6, characterized in that Inputting the pathological image into the image segmentation model to obtain an image segmentation result includes: regarding each whole-slide pathological image as a bag, and each pixel in the whole-slide pathological image as an instance, and using the image segmentation model to implement an instance prediction task based on bag-level labels.
9. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer device, comprising a memory and a processor, wherein a computer program capable of running on the processor is stored on the memory, characterized in that When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.