Image semantic segmentation method and electronic device
By cropping local windows and context windows in ultra-high resolution pathological slice images, the method of feature coding and attention calculation is solved, and the problem of how local windows understand and fuse context semantics is achieved, and high-precision semantic segmentation is achieved.
Patent Information
- Application Number
- CN202411650693.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-19
AI Technical Summary
The prior art has difficulty in achieving precise semantic segmentation in ultra-high resolution pathological slice images, especially with challenges in how local windows understand and fuse contextual semantics.
A semantic segmentation method of image is proposed. By cropping local windows and context windows with different resolutions in the image to be processed, feature encoding and attention calculation are performed, and feature maps are fused to weight fuse local and context information to realize semantic segmentation.
Through the neural network model architecture Transformer, a neural network model architecture with multi-level deep feature fusion, the comprehensive depth feature fusion of context and local window semantics is achieved, and the semantic segmentation accuracy of ultra-high resolution pathological slice images is improved.
Smart Images

Figure CN119152507B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image segmentation, and in particular to a method and electronic device for semantic segmentation of an image. Background Art
[0002] Usually, a 20x full-field digital slice image (WSI) has a resolution of 80,000*80,000 and 6.4 billion pixels. In such a high-resolution image, deep learning methods are needed to segment normal areas, benign tumor areas, in situ tumor areas, and invasive tumor areas, or to segment different types of tissue areas. The segmentation accuracy requirements are getting higher and higher, and at the same time, it poses a major challenge to the efficiency of the algorithm. Considering that the graphics processing unit (GPU) has limited memory and cannot process the entire image at once, traditional methods either downsample an ultra-high-resolution image or crop it into small blocks for separate processing. Either way, the loss of local details or global context information will lead to limited segmentation accuracy.
[0003] Regarding the problem of how to understand and integrate contextual semantics in local windows of ultra-high-resolution pathological slice images in related technologies, and the difficulty in achieving accurate semantic segmentation based on local windows, no effective solution has been proposed so far. Summary of the invention
[0004] The embodiments of the present application provide a method for semantic segmentation of an image and an electronic device, which at least solve the problem of how to understand and fuse contextual semantics in a local window of an ultra-high-resolution pathological slice image, and the problem that it is difficult to achieve accurate semantic segmentation based on a local window.
[0005] According to one aspect of an embodiment of the present application, a method for semantic segmentation of an image is provided, comprising: cropping a local window in an image to be processed, and cropping two context windows of different resolution sizes in a downsampled image obtained by downsampling the image to be processed; performing feature encoding on the local window and the two context windows, respectively, to obtain corresponding convolution feature maps; based on the convolution feature map, performing attention calculations on the local window and itself, and on the local window and the two context windows, respectively, and weighted fusion to obtain an attention feature map of the local window; obtaining a fused feature map of the local window based on the convolution feature map and the attention feature map of the local window; and performing semantic segmentation processing on the local window based on the fused feature map of the local window.
[0006] Optionally, before cropping two context windows of different resolution sizes from the downsampled image obtained by downsampling the processed image, the above method also includes: downsampling the processed image according to two different downsampling ratios, respectively, to obtain two downsampled images of different resolution sizes, wherein the downsampled image with a high downsampling ratio is used to crop the first context window, and the downsampled image with a low downsampling ratio is used to crop the second context window, and the resolution of the first context window restored to before downsampling is higher than the resolution of the second context window restored to before downsampling.
[0007] Optionally, based on the convolution feature map, attention calculation is performed on the local window and itself, and on the local window and two context windows respectively, and weighted fusion is performed to obtain an attention feature map of the local window, including: feature embedding calculation is performed on the local window, the first context window and the second context window respectively to obtain the embedded features of the local window, the first context window and the second context window respectively; based on the embedded features, the query, key and value of the local window, the key and value of the first context window, and the key and value of the second context window are calculated respectively; according to the query, key and value of the local window, the local window and the first local window after attention calculation are calculated to update the embedded features; according to the query of the local window, the key and value of the first context window, the local window and the second local window after attention calculation are calculated to update the embedded features; according to the query of the local window, the key and value of the first context window, the local window and the second local window after attention calculation are calculated to update the embedded features; according to the query of the local window, the key and value of the second context window, the local window and the second context window are calculated to update the embedded features of the third local window after attention calculation is calculated; through Concat-Excitation-and-Split The network calculates the weights corresponding to the first local window updated embedding feature, the second local window updated embedding feature and the third local window updated embedding feature respectively, and multiplies the first local window updated embedding feature, the second local window updated embedding feature and the third local window updated embedding feature with their respective corresponding weights and then adds them together to obtain the local window attention feature map.
[0008] Optionally, a fused feature map of the local window is obtained based on the convolution feature map and the attention feature map of the local window, including: splicing the convolution feature map of the local window and the attention feature map of the local window in the channel dimension to obtain the fused feature map of the local window.
[0009] Optionally, semantic segmentation processing is performed on the local window based on the fused feature map of the local window, including: classifying each pixel point in the fused feature map corresponding to the local window to obtain a semantic segmentation result of the fused feature map of the local window; upsampling the semantic segmentation result of the fused feature map of the local window to obtain a semantic segmentation result of the local window; and splicing the semantic segmentation results of each local window to obtain a semantic segmentation result of the image to be processed.
[0010] Optionally, before performing semantic segmentation processing on the local window based on the fused feature map of the local window, the above method also includes: classifying each pixel point in the convolution feature map corresponding to the first context window to obtain the semantic segmentation result of the convolution feature map corresponding to the first context window; upsampling the semantic segmentation result of the convolution feature map corresponding to the first context window to obtain the semantic segmentation result of the first context window.
[0011] Optionally, before performing semantic segmentation processing on the local window based on the fused feature map of the local window, the above method also includes: classifying each pixel point in the convolution feature map corresponding to the second context window to obtain the semantic segmentation result of the convolution feature map corresponding to the second context window; upsampling the semantic segmentation result of the convolution feature map corresponding to the second context window to obtain the semantic segmentation result of the second context window.
[0012] Optionally, performing semantic segmentation processing on the local window based on the fused feature map of the local window also includes: determining a first loss function of a semantic segmentation model for performing semantic segmentation on the first context window; determining a second loss function of the semantic segmentation model for performing semantic segmentation on the second context window; determining a third loss function of the semantic segmentation model for performing semantic segmentation on the local window; performing weighted processing on the sum of the first loss function and the second loss function and adding the sum to the third loss function to obtain a target loss function; and using the target loss function to control the segmentation accuracy of the semantic segmentation processing on the local window.
[0013] Optionally, feature encoding is performed on the local window and the two context windows respectively, including: using a deep residual convolutional neural network model to perform feature encoding on the local window, wherein the deep residual convolutional network model includes 49 convolutional layers, and the convolutional layers are used for feature encoding; using a convolutional neural network model to perform feature encoding on the two context windows respectively, wherein the convolutional neural network model is composed of multiple convolutional layers and multiple pooling layers connected in sequence.
[0014] According to another aspect of an embodiment of the present application, an electronic device is provided, including: a processor, and a memory storing a program, wherein the program includes instructions, and when the instructions are executed by the processor, the processor executes the above method.
[0015] According to another aspect of an embodiment of the present application, a non-transitory machine-readable medium storing computer instructions is also provided, where the computer instructions are used to enable a computer to execute the above method.
[0016] Beneficial effects of the embodiments of the present application:
[0017] In an embodiment of the present application, a multi-level deep feature fusion neural network model architecture Transformer is proposed. Through the Transformer, semantic segmentation of pathological slice images is performed to achieve the purpose of all-round deep feature fusion of context and local window semantics, thereby achieving the technical effect of improving the semantic segmentation accuracy of ultra-high resolution pathological slice images.
[0018] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other embodiments can be obtained based on these drawings without creative work.
[0020] Figure 1 It is a schematic diagram of the principle of a wide-context network WicoNet;
[0021] Figure 2 It is a schematic diagram of the principle of a multi-level visual feature deep fusion converter network MFFT according to an embodiment of the present application;
[0022] Figure 3 is a flowchart of a method for semantic segmentation of an image according to an embodiment of the present application;
[0023] Figure 4 It is a schematic diagram of splicing a convolution feature map and an attention map of a local window according to an embodiment of the present application;
[0024] Figure 5 is a schematic structural diagram of an electronic device of this embodiment;
[0025] Figure 6a 1 is a schematic diagram of the structure of a connection-excitation-and-split network Concat-Excitation-and-Split Network according to an embodiment of the present application;
[0026] Figure 6b It is a schematic diagram of the structure of the excitation operation in the Concat-Excitation-and-Split Network;
[0027] Figure 6cIt is a structural diagram of the weight operation in the Concat-Excitation-and-Split Network. DETAILED DESCRIPTION
[0028] Embodiments of the present embodiment will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present embodiment are shown in the accompanying drawings, it should be understood that the present embodiment can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, which are instead provided for a more thorough and complete understanding of the present embodiment. It should be understood that the drawings and embodiments of the present embodiment are only for exemplary purposes and are not intended to limit the scope of protection of the present embodiment.
[0029] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:
[0030] WSIs, the abbreviation of Whole Slide Images, is a full-view digital slide. Digital slides are not static pictures. They contain all the lesion information on the glass slides. On the computer, they can be observed at different magnifications, just like under a microscope.
[0031] CNN, the abbreviation of Convolutional Neural Networks, is a deep learning model that is often used to analyze visual images.
[0032] RESNET50 is a deep residual network based on residual network. It has a total of 50 layers, including 49 convolutional layers and one fully connected layer. It can train very deep neural networks, avoid the gradient vanishing problem, and improve the expressiveness and performance of the model.
[0033] Transformer is a deep learning model architecture. The Transformer architecture introduces the self-attention mechanism, which is another key innovation after the introduction of CNN. It can capture the relationship between text and image sequences. Compared with CNN, they are two different feature learning methods.
[0034] In the related art, there is a problem that it is difficult to achieve accurate semantic segmentation for ultra-high resolution pathological slice images. In order to solve this problem, a relevant solution is provided in the embodiments of the present application, which is described in detail below.
[0035] An existing wide-context network (WiCoNet) for high-resolution image semantic segmentation can perform semantic segmentation on images by combining contextual information.
[0036] WiCoNet includes two CNNs, which extract features from local and global level images respectively, which enables WiCoNet to consider local details and wide context at the same time. The local branch M1 is the main branch of WiCoNet, which uses the residual network ResNet to extract local features, and the context branch M2, which is introduced to explicitly model wide range context information, uses a simple CNN encoder to roughly learn context information (rather than collecting spatial details). The context information is embedded into the local branch M1 through the context transformer Context Transformer. The final result of WiCoNet is produced by the context-rich local branch M1.
[0037] Figure 1 This is a schematic diagram of the principle of a wide-context network WicoNet, such as Figure 1 As shown, the algorithm mainly includes the following steps:
[0038] 1. Downsampling of large images. Crop the local window and context window in the original image and the downsampled image respectively. The resolution of the local window is 256*256, and the size of the context window is 9 times the size of the local window (w = 3*256, h = 3*256). When the local window is located at the image boundary, the blank area in the context window will be filled with the reflection of the image. The basic idea of reflection filling is to reflect the pixel values near the image boundary along the boundary, so as to maintain the continuity of the edge information and avoid introducing additional artificial edge effects. For example, if a local window is located at the right boundary of the image, the blank part on the right side of the window will be filled with the reflection of the nearest pixel value on the left.
[0039] 2. The CNN encoder is used to encode features for the local window and context window, that is, the M1 and M2 branches. M1 selects ResNet50 as the feature extraction network, which is powerful in utilizing local features. The downsampling ratio is ×1 / 8 to better preserve spatial information. M2 uses a simple convolution block (called a context encoder) to extract context features. It consists of 11 sequentially connected layers, including 8 convolutional layers and 3 maximum pooling layers. According to the encoder design of UNet, each pooling layer is placed after 2 convolutional layers, and the downsampling ratio of the context encoder is the same as that of the residual network ResNet (×1 / 8).
[0040] 3. Transformer with context feature fusion. First, the encoded local window features and context window features are embedded and patch embedding is calculated. The token of the local window is used as the query, and the token of the context window is used as the key and value. The attention is calculated, and the value after softmax is used as the weight. It is multiplied by the value to get the token of the updated local window.
[0041]
[0042] , , Represents the local query, context key and context value respectively. , Represents the token of the local window and the context window respectively (an embedding vector), is the updated local window token after attention calculation. is the attention of the context. D and n are the dimension of embedding and the number of heads of multi-head attention, respectively.
[0043] Transformer hyperparameters include: L is the number of blocks, n is the number of heads, p is the size of the embedding parameter, and d is the dimension of the embedding. p is set to 1 to preserve spatial information. D is set to 512, which is the number of output channels of the context encoder. L and n are set based on experimental results.
[0044] The output of Transformer reflects the relationship between local and contextual features, integrates contextual and local features, and replaces the features encoded by M1 branch ResNet50.
[0045] 4. Use conv1x1 convolution on the output feature map of Transformer to classify each pixel.
[0046] 5. Upsample the classified feature map and restore it to the local window size to obtain the classification of each pixel in the local window, which is the semantic segmentation of the local window.
[0047] 6. Concatenate the semantic segmentation of each local window to obtain the semantic segmentation of each large image.
[0048] 7. The context window is also segmented, and the loss function is the weighted sum of the context window segmentation loss function and the local window segmentation loss function. The weighting parameter α is dynamically calculated at each iteration. In this way, the value of the context window loss function decreases with iteration, and more and more attention is paid to the accuracy of local window segmentation.
[0049]
[0050] is the total loss function, , are the loss functions for segmenting the local window and the context window, respectively. They are the segmentation result output by the model for segmenting the local window and the actual segmentation result. They are the segmentation result output by the model of the segmentation context window and the actual segmentation result.
[0051] The current WicoNet has deficiencies in the width, depth, and fusion method of the context window and local window feature fusion:
[0052] The insufficient fusion width is because the context window is small and single, with only one context size that is three times the length and width of the local window.
[0053] Insufficient fusion depth means that the feature map only contains the feature map after the local window Transformer, and does not fully utilize the feature map after the local window Resnet50 encoding.
[0054] Insufficient fusion means that the Transformer fusion method only adds the embedding vector feature and the weighted value, without reflecting the weight.
[0055] To address the deficiencies of the prior art, this application proposes a multi-level deep feature fusion transformer, namely a multi-level vision feature deep fusion transformer network (MFFT), which achieves the purpose of omni-directional deep feature fusion of context and local window semantics. This is described in detail below.
[0056] MFFT includes three CNNs, which extract features from the local window and two wider context windows respectively. The local branch uses ResNet to extract local features. The Transformer designed by MFFT has three levels of context, and the local window is also a context, which is used to mine the correlation between local windows. The feature fusion of the context is not a simple addition, but a weighted fusion mechanism is adopted to learn the weight of the feature. In addition, a fusion mechanism of Transform features and local window convolution features is designed. The context information is embedded into the local window through the Transformer, and then the convolution features are fused.
[0057] Figure 2FIG. 1 is a schematic diagram of a multi-level visual feature deep fusion converter network MFFT according to an embodiment of the present application. The MFFT structure is as follows: Figure 2 As shown, Figure 2 The shallow encoder in the figure represents a simple 11-layer CNN encoder, and the deep encoder represents a Resnet50 encoder. MFFT is used for the segmentation of pathological slice images. The large tissue area, medium tissue area, cell organization, and cell correspond to the large context window, medium context window, local window, and the number of pixels in the original image corresponding to a single pixel of the local window feature map, that is, the downsampling ratio of the encoding. The feature map is the feature map encoded by CNN. Feature embedding is to map the high-dimensional and sparse input data of Transformer to a low-dimensional and dense vector space to reduce the dimension of the data and improve computational efficiency. The attention map is the feature map output by the Concat-Excitation-and-Split Network (CESNet). The rich feature map is the fusion of the feature map and the attention map. The segmentation result seg is the segmentation result based on the rich feature map. The patch segmentation result patch seg is the segmentation result of upsampling the rich feature map to a local window.
[0058] exist Figure 2 For the sake of simplicity, the segmentation process of the two contexts is not shown. In fact, the segmentation results of the context are also needed to calculate the loss function.
[0059] Next, the algorithm steps of MFFT are described in conjunction with a specific embodiment.
[0060] Figure 3 is a flowchart of a method for semantic segmentation of an image according to an embodiment of the present application, such as Figure 3 As shown, the method comprises the following steps:
[0061] Step S302 , cropping a local window in the image to be processed, and cropping two context windows of different resolutions in a downsampled image obtained by downsampling the image to be processed.
[0062] According to an optional embodiment of the present application, before cropping two context windows of different resolution sizes from the downsampled image obtained by downsampling the processed image, the processed image is downsampled according to two different downsampling ratios to obtain two downsampled images of different resolution sizes, wherein the downsampled image with a high downsampling ratio is used to crop the first context window, and the downsampled image with a low downsampling ratio is used to crop the second context window, and the resolution of the first context window restored to before downsampling is higher than the resolution of the second context window restored to before downsampling.
[0063] In computer vision, local windows and context windows are two different types of windows used for image analysis, and there is a close relationship between them. Local windows and context windows complement each other in visual tasks and together help the model understand and analyze the information in the image.
[0064] A local window is a smaller area cropped from the original image, which usually contains a specific part or object in the image. The local window is mainly used to capture local features in the image, such as texture, edges, etc. It can help the model focus on a part of the image and extract useful information from it.
[0065] The context window refers to a larger area that includes one or more local windows and the surrounding environment. It includes not only the information of the target object itself, but also the environmental information around the target. The context window can help the model understand the relationship between local features and the overall scene, which is very important for many advanced visual tasks.
[0066] In an embodiment of the present application, the large image (i.e., the image to be processed) is downsampled 4 times and 8 times, respectively. "4 times" and "8 times" generally refer to the reduction ratio relative to the original image size. In the original image and the downsampled image, the local window and two sizes of context windows are cropped respectively. The local window has a ratio of 512*512, and the size of the context window is set to 9 times (w = 3*512, h = 3*512) and 36 times (w = 6*512, h = 6*512) of the local window size. The local window and context window sizes are configurable parameters. As mentioned above, the parameter values have practical business meanings.
[0067] It should be noted that two other different downsampling ratios can be performed on the processed image. Downsampling is to crop the context window from a macroscopic perspective. The higher the downsampling ratio, the larger the context window restored to the original image size. In the above embodiment, a context window with a lower resolution is cropped in the downsampled image obtained by 4 times downsampling, and a context window with a higher resolution is cropped in the downsampled image obtained by 8 times downsampling.
[0068] There are three sizes of contexts: the local window size is 512*512, and the two sizes of contexts are 1536*1536 and 3076*3076 respectively. The local window itself is also the source of context. Exploring the connections between local windows broadens the context vision and improves the level of fusion.
[0069] It should be noted that the image to be processed in step S302 is an image with ultra-high resolution.
[0070] Step S304, feature encoding is performed on the local window and the two context windows respectively to obtain corresponding convolution feature maps.
[0071] According to an optional embodiment of the present application, step S304 is executed to perform feature encoding on the local window and the two context windows respectively, including the following steps: using a deep residual convolutional neural network model to perform feature encoding on the local window, wherein the deep residual convolutional network model includes 49 convolutional layers, and the convolutional layers are used for feature encoding; using a convolutional neural network model to perform feature encoding on the two context windows respectively, wherein the convolutional neural network model is composed of multiple convolutional layers and multiple pooling layers connected in sequence.
[0072] In an embodiment of the present application, a CNN encoder is used to encode features for the local window and the context window, that is, the three branches. The local window uses a deep residual convolutional network model including 49 convolutional layers as the feature extraction network, and the downsampling ratio is 1 / 16 to reduce the computational complexity of the Transformer. The two context windows each use an 11 sequentially connected layer (including 8 convolutional layers and 3 maximum pooling layers) to extract features, and the downsampling ratio is the same as ResNet50 (×1 / 16). The downsampling ratio is also a parameter that can be set. As mentioned above, the parameter value also has practical business significance.
[0073] Step S306, based on the convolution feature map, attention calculation is performed on the local window and itself, and on the local window and two context windows, and weighted fusion is performed to obtain the attention feature map of the local window.
[0074] According to an optional embodiment of the present application, step S306 is performed based on the convolution feature map to perform a local window and itself ( Figure 2 The self-attention mechanism in the local window and the two context windows perform attention calculations respectively ( Figure 2 Attention mechanism 1 and attention mechanism 2 in the above diagram are combined and weighted to obtain the attention feature map of the local window, including the following steps:
[0075] The feature embedding calculation is performed on the local window, the first context window and the second context window respectively to obtain the embedding features of the local window, the first context window and the second context window respectively; the query, key and value of the local window, the key and value of the first context window, and the key and value of the second context window are calculated based on the embedding features; the updated embedding feature of the local window and the first local window after attention calculation is calculated according to the query, key and value of the local window; the updated embedding feature of the second local window after attention calculation is calculated according to the query of the local window, the key and value of the first context window; the updated embedding feature of the third local window after attention calculation is calculated according to the query of the local window, the key and value of the second context window; the weights corresponding to the updated embedding feature of the first local window, the updated embedding feature of the second local window and the updated embedding feature of the third local window are calculated respectively through the Concat-Excitation-and-Split Network, and the updated embedding feature of the first local window, the updated embedding feature of the second local window and the updated embedding feature of the third local window are multiplied by the corresponding weights and then added to obtain the attention feature map of the local window.
[0076] First, feature embedding calculation (patch embedding) is performed on the local window, the first context window, and the second context window respectively to obtain the embedded features of the local window, the first context window, and the second context window (that is, the token mentioned above is an embedding vector).
[0077] Next, the query, key and value of the local window, the key and value of the first context window, and the key and value of the second context window are calculated using the embedded features of the local window, the first context window, and the second context window, respectively. The calculation formula is as follows:
[0078]
[0079] q, k, v represent query, key and value respectively. The subscript letters l, c1, c2 represent local window, medium-sized context window (second context window) and large-sized context window (first context window) respectively.
[0080] In the fields of natural language processing and computer vision, especially in the attention mechanism, "query", "key" and "value" are three core concepts. These concepts were first introduced in the Transformer model to process sequence data.
[0081] A query is a vector or set of vectors that represents the current focus or interest. It is used to find relevant information from other vectors. The query is compared with the key to determine how to assign attention weights. In Transformer, the query is calculated by multiplying the embedding vector of the input sequence by a learnable weight matrix.
[0082] A key is also a vector or a set of vectors that indicate the location of information. They are matched with the query to determine the importance of the value. The key is compared with the query to calculate the attention weight, which determines which value should be considered more. The key is also calculated by multiplying the embedding vector of the input sequence by a learnable weight matrix.
[0083] The values contain the actual information to be attended to. They are weighted according to the attention weights calculated from the query and key. The values are weighted and combined according to the calculated attention weights to produce the output. The values are calculated by multiplying the embedding vector of the input sequence by a learnable weight matrix.
[0084] A and W represent attention and attention weight respectively, and the subscripts represent the objects of attention.
[0085] , , They represent the local window tokens updated based on the attention of other local windows, medium-sized context windows, and large-sized context windows, respectively. , , respectively correspond to the first local window update embedding feature, the second local window update embedding feature and the third local window update embedding feature, The updated value after weighting different attention objects (i.e., the attention feature map of the local window).
[0086] The output attention map of the weighted multi-attention fusion Transformer combines the features of the local window and the multi-level context window.
[0087] In MFFT, there are three kinds of contexts for local windows. The local window itself is also a context, which is used to model the correlation between different local windows. For the three designed context attentions, Concat-Excitation-and-Split Network is used for fusion. The Weighted Multi-attention Transformer and Concat-Excitation-and-Split Network technology are universal and decoupled from the present invention. The relevant content of Concat-Excitation-and-Split Network is explained below.
[0088] The Concat-Excitation-and-Split Network module is used to fuse multi-channel attention maps from different sources. Its structure and process are as follows Figure 6a As shown, there are mainly 4 operations, the steps are as follows:
[0089] 1. Concat operation. For attention maps from multiple sources, concat operation is performed on the channel dimension. The three sources shown in the figure are for reference only, and the number of sources should be increased or decreased according to the actual number. c is the number of channels, h and w are the height and width respectively.
[0090] 2. Excitation operation. Figure 6b As shown:
[0091] a) The input is the concatenated multi-channel attention map, not the multi-channel real number. The output is not the multi-channel weight, but the multi-channel weight map.
[0092] b) Replace the full connection with con2d(1,1). That is, use a convolution kernel with the number of c / r, the number of channels c, and the size (1,1) for convolution operation.
[0093] c) The relu operation after the full connection is cancelled.
[0094] 3. Weighting. Figure 6c As shown, the Weighting operation also includes 2 steps:
[0095] 1) Split: Split in the channel dimension according to the number of dimensions before the concatenation operation.
[0096] 2) Softmax. Before softmax, dimension 1 needs to be inserted into dimension 0, and then normalized on the source dimension to obtain the weight map of each source attention map. The weight map has the same size and the same number of channels as the attention map.
[0097] 4. Fusion operation: The weight map of each channel of each source and the attention map of each channel are multiplied and then added together to obtain the converged attention feature.
[0098] Figure 6a The circle with a dot in it represents dot product, and the circle with a + sign represents addition.
[0099] In an embodiment of the present application, the attention maps of three contexts are fused using a concat-excitation-and-weighting net, which learns the weights of different graphs from the data. The weights of the weighted addition can be learned, which improves the fusion effect compared to simple addition.
[0100] Step S308, obtaining a fused feature map of the local window based on the convolution feature map and the attention feature map of the local window.
[0101] Step S310, performing semantic segmentation processing on the local window based on the fused feature map of the local window.
[0102] The above method provided in the embodiment of the present application performs semantic segmentation on the pathological slice image through Transformer, thereby achieving the purpose of all-round deep feature fusion of context and local window semantics, thereby realizing the technical effect of improving the semantic segmentation accuracy of ultra-high resolution pathological slice images.
[0103] According to an optional embodiment of the present application, step S308 is executed to obtain a fused feature map of the local window based on the convolution feature map and the attention feature map of the local window, which is achieved by the following method: the convolution feature map of the local window and the attention feature map of the local window are spliced in the channel dimension to obtain a fused feature map of the local window.
[0104] Figure 4 is a schematic diagram of splicing a convolution feature map and an attention map of a local window according to an embodiment of the present application, such as Figure 4As shown in the figure, the convolutional feature map feature map of the local window encoded by the local window Resnet50 and the attention map attention map output by the Transformer are fused and spliced in the channel dimension to obtain a rich feature map richfeature map.
[0105] In some optional embodiments of the present application, step S310 is executed to perform semantic segmentation processing on the local window based on the fused feature map of the local window, including the following steps: classifying each pixel point in the fused feature map corresponding to the local window to obtain the semantic segmentation result of the fused feature map of the local window; upsampling the semantic segmentation result of the fused feature map of the local window to obtain the semantic segmentation result of the local window; splicing the semantic segmentation results of each local window to obtain the semantic segmentation result of the image to be processed.
[0106] Based on the rich feature map, conv1x1 convolution is used to classify each pixel. The feature map is upsampled and restored to the local window size to obtain the classification of each pixel in the local window, that is, the semantic segmentation of the local window. The semantic segmentation of each local window is spliced to obtain the semantic segmentation of each large image.
[0107] It should be noted that, as mentioned above, the local window uses a deep residual convolutional network model as the feature extraction network, and the downsampling ratio is 1 / 16 to reduce the computational complexity of the Transformer. Therefore, after classifying each pixel, the feature map needs to be upsampled (with a sampling ratio of 16) to restore it to the local window size.
[0108] In an embodiment of the present application, a fusion mechanism for splicing convolutional features and Transformer features in the channel dimension is designed, which fully utilizes the advantages of Transformer encoding and Resnet50 encoding.
[0109] As some optional embodiments of the present application, before executing step S310, each pixel point in the convolution feature map corresponding to the first context window is classified to obtain the semantic segmentation result of the convolution feature map corresponding to the first context window, and the semantic segmentation result of the convolution feature map corresponding to the first context window is upsampled to obtain the semantic segmentation result of the first context window.
[0110] And classify each pixel point in the convolution feature map corresponding to the second context window to obtain the semantic segmentation result of the convolution feature map corresponding to the second context window; upsample the semantic segmentation result of the convolution feature map corresponding to the second context window to obtain the semantic segmentation result of the second context window.
[0111] The context windows of both resolutions also need to be segmented to calculate the loss function. The segmentation method of the context window is similar to that of the local window. Here, only the segmentation process of the first context window is briefly described (the segmentation process of the second context window is the same).
[0112] Based on the convolution feature map corresponding to the first context window, use conv1x1 convolution to classify each pixel. Upsample the semantic segmentation result of the convolution feature map and restore it to the size of the first context window to obtain the classification of each pixel in the first context window, that is, the semantic segmentation of the first context window.
[0113] As an optional embodiment of the present application, executing step S310 to perform semantic segmentation processing on the local window based on the fused feature map of the local window, also includes the following technical solutions: determining a first loss function of a semantic segmentation model for performing semantic segmentation on the first context window; determining a second loss function of a semantic segmentation model for performing semantic segmentation on the second context window; determining a third loss function of a semantic segmentation model for performing semantic segmentation on the local window; weighting the sum of the first loss function and the second loss function and adding the weighted sum to the third loss function to obtain a target loss function; and using the target loss function to control the segmentation accuracy of semantic segmentation processing on the local window.
[0114] The loss function is a weighted sum of two context window segmentation losses and local window segmentation losses.
[0115]
[0116] is the total loss function, , , are the loss functions for segmentation of local window, medium-sized context window and large-sized context window respectively, They are the segmentation result output by the model for segmenting the local window and the actual segmentation result. , The segmentation results output by the model for segmenting context windows of two sizes and the actual segmentation results.
[0117] After the image to be processed is segmented, the segmentation results are verified using the BACH dataset. The BACH dataset is the data of the challenge held by the International Conference on Image Analysis and Recognition (ICIAR), including 10 large-scale image data from A01 to A10. Each large-scale image includes the segmentation results of normal areas, benign tumor areas, in situ tumor areas, and invasive tumor areas. A01 to A09 are used for training, and A10 is used for verification. The training is repeated for 50 epochs. Different versions of WiCoNet and MFFT are used for training and verification. The verification results are shown in Table 1 below. WiCoNet has an 8% improvement in accuracy compared to the traditional sliding window segmentation algorithm. The accuracy of each version of MFFT is significantly improved compared to WiCoNet. This article describes MFFT-tranformer&resnet50, which has an 8% improvement. MFFT-tranformer&resnet50 (n = 4, l = 4) and MFFT-tranformer&resnet50 have different parameter settings.
[0118] Table 1 Verification results of MFFT network
[0119]
[0120] The embodiment of the present application also provides an electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication. The memory stores a computer program executable by the at least one processor, and the computer program is used to enable the electronic device to perform the method of the embodiment of the present application when executed by the at least one processor.
[0121] An embodiment of the present application also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiment of the present application.
[0122] The embodiment of the present application also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiment of the present application.
[0123] refer to Figure 5, the structural block diagram of the electronic device that can be used as the server or client of the embodiment of the present application will now be described, which is an example of the hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples, and are not intended to limit the implementation of the present application described and / or required herein.
[0124] like Figure 5 As shown, the electronic device includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 to a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0125] Multiple components in the electronic device are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 can be any type of device that can input information to the electronic device, and the input unit 506 can receive input digital or character information, and generate key signal input related to user settings and / or function control of the electronic device. The output unit 507 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 508 can include but is not limited to a disk, an optical disk. The communication unit 509 allows the electronic device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0126] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU, a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the method embodiments of the present application may be implemented as a computer program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on an electronic device via a ROM 502 and / or a communication unit 509. In some embodiments, the computing unit 501 may be configured to perform the above-described method in any other appropriate manner (e.g., by means of firmware).
[0127] The computer program for implementing the method of the embodiment of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0128] In the context of the embodiments of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0129] It should be noted that the term "including" and its variations used in the embodiments of the present application are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present application are illustrative and not restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0130] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0131] The various steps described in the method implementation methods provided in the embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method implementation methods may include additional steps and / or omit the steps shown. The scope of protection of the present application is not limited in this respect.
[0132] The term "embodiment" in this specification refers to specific features, structures or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. The various embodiments in this specification are described in a related manner, and the same or similar parts between the various embodiments refer to each other. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiment.
[0133] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.
Claims
1. A semantic segmentation method for an image, characterized in that: include: Cropping a local window in the image to be processed, and cropping two context windows of different resolution sizes in a downsampled image obtained by downsampling the image to be processed; Before cropping two context windows of different resolution sizes from the downsampled image obtained by downsampling the image to be processed, the method further includes: downsampling the image to be processed according to two different downsampling ratios to obtain the downsampled images of two different resolution sizes, wherein the downsampled image with a high downsampling ratio is used to crop a first context window, and the downsampled image with a low downsampling ratio is used to crop a second context window, and the resolution of the first context window restored to before downsampling is higher than the resolution of the second context window restored to before downsampling; Performing feature encoding on the local window and the two context windows respectively to obtain corresponding convolution feature maps; Based on the convolution feature map, attention calculation is performed on the local window and itself, and the local window and the two context windows respectively, and weighted fusion is performed to obtain the attention feature map of the local window, including: Perform feature embedding calculations on the local window, the first context window, and the second context window respectively to obtain embedding features of the local window, the first context window, and the second context window; Based on the embedded features, respectively, the query, key and value of the local window, the key and value of the first context window, and the key and value of the second context window are calculated; Calculate the local window and the first local window after attention calculation based on the query, key and value of the local window to update the embedding feature; According to the query of the local window, the key and value of the first context window are calculated, and the local window and the second local window after the attention calculation of the first context window are updated with embedded features; According to the query of the local window, the key and value of the second context window are calculated to update the embedding feature of the third local window after the local window and the second context window are calculated by attention; The weights corresponding to the first local window updated embedding feature, the second local window updated embedding feature, and the third local window updated embedding feature are calculated respectively through Concat-Excitation-and-Split Network, and the first local window updated embedding feature, the second local window updated embedding feature, and the third local window updated embedding feature are multiplied by their respective weights and then added to obtain an attention feature map of the local window; Obtaining a fused feature map of the local window based on the convolution feature map and the attention feature map of the local window; Perform semantic segmentation processing on the local window based on the fused feature map of the local window.
2. The method according to claim 1, characterized in that: Obtaining a fusion feature map of the local window based on the convolution feature map and the attention feature map of the local window, including: The convolution feature map of the local window and the attention feature map of the local window are spliced in the channel dimension to obtain a fusion feature map of the local window.
3. The method according to claim 1, characterized in that Performing semantic segmentation processing on the local window based on the fused feature map of the local window includes: Classifying each pixel point of the fused feature map of the local window to obtain a semantic segmentation result of the fused feature map of the local window; Performing upsampling processing on the semantic segmentation result of the fused feature map of the local window to obtain the semantic segmentation result of the local window; The semantic segmentation results of each local window are spliced to obtain the semantic segmentation result of the image to be processed.
4. The method according to claim 1, characterized in that: Before performing semantic segmentation processing on the local window based on the fused feature map of the local window, the method further includes: Classifying each pixel point in the convolution feature map corresponding to the first context window to obtain a semantic segmentation result of the convolution feature map corresponding to the first context window; The semantic segmentation result of the convolution feature map corresponding to the first context window is upsampled to obtain the semantic segmentation result of the first context window.
5. The method according to claim 1, characterized in that Before performing semantic segmentation processing on the local window based on the fused feature map of the local window, the method further includes: Classify each pixel in the convolution feature map corresponding to the second context window to obtain a semantic segmentation result of the convolution feature map corresponding to the second context window; The semantic segmentation result of the convolution feature map corresponding to the second context window is upsampled to obtain the semantic segmentation result of the second context window.
6. The method according to claim 1, characterized in that Performing semantic segmentation processing on the local window based on the fused feature map of the local window, further comprising: Determining a first loss function of a semantic segmentation model for performing semantic segmentation on the first context window; Determine a second loss function of a semantic segmentation model for performing semantic segmentation on the second context window; Determine a third loss function of a semantic segmentation model for performing semantic segmentation on the local window; Performing weighted processing on the sum of the first loss function and the second loss function and adding the sum to the third loss function to obtain a target loss function; The target loss function is used to control the segmentation accuracy of the semantic segmentation processing performed on the local window.
7. The method according to claim 1, characterized in that The local window and the two context windows are respectively subjected to feature encoding, including: Performing feature encoding on the local window using a deep residual convolutional neural network model, wherein the deep residual convolutional neural network model includes 49 convolutional layers, and the convolutional layers are used for feature encoding; A convolutional neural network model is used to perform feature encoding on the two context windows respectively, wherein the convolutional neural network model is composed of a plurality of convolutional layers and a plurality of pooling layers connected in sequence.
8. An electronic device comprising: A processor and a memory storing a program, wherein the program comprises instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 7.
9. A non-transitory machine-readable medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Image semantic segmentation method based on Transform visual upsampling module
CN113888744A
Image segmentation method of global context attention network based on multi-scale fusion
CN115375711A