Image segmentation method and device, storage medium and electronic equipment
By combining super-resolution networks and semantic segmentation networks in medical image segmentation, the super-resolution network is used to generate high-resolution feature maps and input them into the semantic segmentation network, the problem of inaccurate image segmentation under low-resolution images is solved, and higher image segmentation accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510173680.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
The accuracy of medical image segmentation cannot be guaranteed in the prior art, especially in low-resolution image data, deep learning algorithms are difficult to accurately characterize micro lesions.
By acquiring the images to be processed, super-resolution networks and semantic segmentation networks, the images are enhanced by using super-resolution networks to generate feature maps of different resolutions, and input these feature maps into multiple intermediate layers of the semantic segmentation network to improve the accuracy of image segmentation.
By improving image resolution and deepening data interaction between super-resolution networks and semantic segmentation networks, the accuracy and efficiency of image segmentation are significantly improved.
Smart Images

Figure CN120107585A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an image segmentation method, device, storage medium and electronic equipment. Background Art
[0002] Nowadays, more and more researchers are using deep learning to segment medical images in order to extract effective information from medical images to assist doctors in diagnosis.
[0003] However, the medical image data collected by medical devices often have low resolution, making it difficult for deep learning algorithms to accurately characterize tiny lesions, which may affect the performance of deep learning algorithms and make it impossible to guarantee the accuracy of image segmentation. Summary of the invention
[0004] In view of this, the embodiments of the present invention are directed to providing an image segmentation method, apparatus, storage medium and electronic device to solve the problem in the prior art that the accuracy of image segmentation cannot be guaranteed.
[0005] One aspect of the present invention provides an image segmentation method, the method comprising:
[0006] Acquire an image to be processed, a super-resolution network, and a semantic segmentation network, wherein the semantic segmentation network includes multiple intermediate layers;
[0007] Performing image enhancement processing on the image to be processed through the super-resolution network to obtain multiple feature maps with different resolutions;
[0008] During the process of the semantic segmentation network performing semantic segmentation on the image to be processed, the multiple feature maps are respectively input into designated intermediate layers among the multiple intermediate layers, so that the semantic segmentation network outputs an image segmentation result of the image to be processed based on the multiple feature maps.
[0009] In one embodiment, the multiple intermediate layers include multiple encoding layers and multiple decoding layers, the multiple encoding layers correspond to the multiple decoding layers one by one, and the inputting the multiple feature maps into the designated intermediate layers of the multiple intermediate layers respectively so that the semantic segmentation network outputs the image segmentation result of the image to be processed based on the multiple feature maps includes:
[0010] Downsampling the image to be processed and the multiple feature maps respectively through the multiple encoding layers to obtain downsampling results corresponding to each of the encoding layers;
[0011] For each of the decoding layers, generating input features of the decoding layer according to a downsampling result output by the encoding layer corresponding to the decoding layer;
[0012] Upsampling the input feature through the decoding layer to output an upsampling result;
[0013] The image segmentation result is determined according to the up-sampling result.
[0014] In one embodiment, for each of the decoding layers, generating the input features of the decoding layer according to the downsampling result output by the encoding layer corresponding to the decoding layer includes:
[0015] For a first decoding layer among the multiple decoding layers, fusing a feature map with the smallest resolution among the multiple feature maps with a downsampling result output by a coding layer corresponding to the first decoding layer to obtain input features of the first decoding layer;
[0016] For a decoding layer other than the first decoding layer among the multiple decoding layers, fusing a down-sampling result output by an encoding layer corresponding to the decoding layer with an up-sampling result output by a previous decoding layer of the decoding layer to obtain an input feature of the decoding layer;
[0017] Determining the image segmentation result according to the upsampling result includes:
[0018] An upsampling result output by a last decoding layer among the multiple decoding layers is determined as the image segmentation result.
[0019] In one embodiment, the super-resolution network includes a feature extraction module and an enhancement module, and the image enhancement processing is performed on the image to be processed by the super-resolution network to obtain multiple feature maps with different resolutions, including:
[0020] Extracting features of the image to be processed by the feature extraction module to obtain a plurality of initial feature maps with different resolutions;
[0021] The initial feature map is subjected to image enhancement processing by the enhancement module to obtain the feature map.
[0022] In one embodiment, the feature extraction module includes multiple feature extraction layers, the feature extraction layer includes at least one dilated convolution layer, the number of dilated convolution layers in the feature extraction layer is positively correlated with the network depth of the feature extraction layer in the super-resolution network, and the feature extraction of the image to be processed by the feature extraction layer to obtain multiple initial feature maps of different resolutions includes:
[0023] For a first feature extraction layer among the multiple feature extraction layers, convolution processing is performed on the image to be processed respectively through the dilated convolution layer in the first feature extraction layer to obtain a first convolution result, and the first convolution result is fused to obtain an initial feature map output by the first feature extraction layer;
[0024] For feature extraction layers other than the first feature extraction layer among the multiple feature extraction layers, the initial feature maps output by the previous feature extraction layer of the feature extraction layer are convolved through the dilated convolution layers in the feature extraction layers to obtain a second convolution result, and the second convolution results are fused to obtain the initial feature map output by the feature extraction layer.
[0025] In one embodiment, the enhancement module includes a separation module submodule and an enhancement submodule, and the super-resolution processing of the initial feature map by the enhancement module to obtain the feature map includes:
[0026] Separating the region containing semantic information in the initial feature map by the separation module submodule to obtain a separation result;
[0027] The separation result is subjected to super-resolution processing by the enhancer module to obtain the feature map.
[0028] In one embodiment, obtaining the image to be processed includes:
[0029] Obtaining initial images of multiple modalities;
[0030] The initial images of the multiple modes are spliced to obtain the image to be processed.
[0031] Another aspect of the present invention provides an image segmentation device, the device comprising:
[0032] An acquisition unit, used to acquire an image to be processed, a super-resolution network, and a semantic segmentation network, wherein the semantic segmentation network includes multiple intermediate layers;
[0033] A feature extraction unit, used to perform image enhancement processing on the image to be processed through the super-resolution network to obtain a plurality of feature maps with different resolutions;
[0034] A segmentation unit is used to input the multiple feature maps into designated intermediate layers among the multiple intermediate layers respectively during the process of the semantic segmentation network performing semantic segmentation on the image to be processed, so that the semantic segmentation network outputs the image segmentation result of the image to be processed based on the multiple feature maps.
[0035] Yet another aspect of the present invention provides a computer-readable storage medium having computer-executable instructions stored thereon, wherein the executable instructions, when executed by a processor, implement the image segmentation method as described in any of the above embodiments.
[0036] Another aspect of the present invention provides an electronic device, the electronic device comprising:
[0037] processor;
[0038] a memory for storing instructions executable by the processor;
[0039] The processor is used to execute the image segmentation method described in any one of the above embodiments.
[0040] Compared with the related art, the image segmentation method provided by the present invention has the following beneficial effects:
[0041] The image segmentation method provided by the present invention includes: obtaining an image to be processed, a super-resolution network and a semantic segmentation network, wherein the semantic segmentation network includes multiple intermediate layers; and performing image enhancement processing on the image to be processed through the super-resolution network to obtain multiple feature maps of different resolutions; in the process of semantic segmentation of the image to be processed by the semantic segmentation network, the multiple feature maps are respectively input to the designated intermediate layers in the multiple intermediate layers, so that the semantic segmentation network outputs the image segmentation result of the image to be processed based on the multiple feature maps. Since the resolution of the image to be processed can be improved after image enhancement through the super-resolution network, the semantic segmentation network can perform image segmentation based on the image to be processed after the resolution is improved, thereby improving the accuracy of image segmentation. In addition, by inputting multiple feature maps of different resolutions obtained by the super-resolution network into the designated intermediate layers in the multiple intermediate layers of the semantic segmentation network, respectively, to assist the semantic segmentation network in performing image segmentation, the data interaction between the super-resolution network and the semantic segmentation network can be deepened, and the mutual adaptability between the semantic segmentation network and the super-resolution network is improved, thereby ensuring the accuracy of image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 FIG. 4 is a flow chart of an image segmentation method provided by an embodiment of the present invention.
[0043] Figure 2 FIG. 4 is a flow chart of an image segmentation method provided by another embodiment of the present invention.
[0044] Figure 3 FIG. 4 is a schematic diagram of the structure of a feature extraction network provided by an embodiment of the present invention.
[0045] Figure 4 FIG. 4 is a schematic diagram of the structure of a dynamic multi-branch module provided by an embodiment of the present invention.
[0046] Figure 5 FIG. 1 is a schematic diagram of the structure of an enhancement module provided in one embodiment of the present invention.
[0047] Figure 6 FIG. 4 is a schematic diagram of the structure of a first hash transformation module provided by an embodiment of the present invention.
[0048] Figure 7 Shown is a schematic diagram of the structure of the semantic segmentation network provided by an embodiment of the present invention.
[0049] Figure 8 FIG. 1 is a schematic diagram of the structure of a connection module provided in one embodiment of the present invention.
[0050] Fig. 9 FIG. 1 is a schematic diagram of parameter comparison provided by an embodiment of the present invention.
[0051] Fig.10 FIG. 4 is a schematic diagram of segmented image comparison provided by an embodiment of the present invention.
[0052] Fig.11 Shown is a loss curve diagram provided by an embodiment of the present invention.
[0053] Fig.12 Shown is a parameter comparison schematic diagram provided by another embodiment of the present invention.
[0054] Fig.13 FIG. 4 is a schematic diagram of parameter comparison provided by yet another embodiment of the present invention.
[0055] Fig.14 FIG. 4 is a principle block diagram of an image segmentation device provided by an embodiment of the present invention.
[0056] Fig.15 FIG. 1 is a block diagram of an electronic device provided in accordance with an embodiment of the present invention. DETAILED DESCRIPTION
[0057] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0058] Medical imaging is an important tool for identifying abnormal cell growth. For example, more and more researchers are using deep learning to segment medical images to extract tumor location information to assist doctors in diagnosis. However, low-resolution medical images captured by medical devices may affect the performance of deep learning algorithms.
[0059] Therefore, in addition to the need for medical image segmentation, it is often necessary to use super-resolution (SR) technology to perform image enhancement processing on low-resolution medical images to improve the clarity of medical images, thereby improving the accuracy of subsequent tasks (such as segmentation tasks).
[0060] In some related technologies, when using deep learning to segment medical images, the super-resolution network used to perform super-resolution tasks and the semantic segmentation network used to perform image segmentation tasks are usually trained separately, that is, the super-resolution network and the semantic segmentation network form a serial structure. Since the serial structure is trained one by one, the super-resolution network and the semantic segmentation network are independent of each other and lack information interaction, which makes the training process less robust, resulting in the data output by the super-resolution network may not be well adapted to the downstream image segmentation task, which may affect the accuracy of image segmentation.
[0061] In other related technologies, the super-resolution network and the semantic segmentation network are trained synchronously, that is, a common loss function is used to correct the two networks, so that the super-resolution network and the semantic segmentation network form a parallel structure, thereby establishing a connection between the two networks. Although such a parallel structure improves the adaptability of the super-resolution network and the semantic segmentation network to a certain extent, on the one hand, there is no data exchange between the super-resolution network and the semantic segmentation network, and the two networks are only adjusted through the loss function. It is still impossible to guarantee good adaptability between the super-resolution network and the semantic segmentation network, and thus the accuracy of image segmentation cannot be guaranteed. On the other hand, synchronous training may bring a large amount of computation and reduce the efficiency of image segmentation.
[0062] In view of the above problems, an embodiment of the present invention provides an image segmentation method, which can be executed by a computer device (for example, a server or a user terminal). Figure 1 As shown, the image segmentation method may include:
[0063] 110. Obtain an image to be processed, a super-resolution network, and a semantic segmentation network, wherein the semantic segmentation network includes multiple intermediate layers.
[0064] The image to be processed refers to the image that needs to be processed in the image segmentation task, which is usually the original data or the pre-processed image. For example, taking the image to be processed as a medical image, the image to be processed can be image data acquired by a medical imaging device, and the medical imaging device includes but is not limited to: computed tomography (CT), magnetic resonance imaging (MRI), ultrasound, etc. The goal of the image segmentation task can be to separate a specific area (such as a tumor, organ, diseased tissue, etc.) from the background or other tissues.
[0065] It is understandable that the image to be processed may be other images that require target segmentation in addition to medical images, which is not limited here.
[0066] Among them, the super-resolution network can be a pre-trained deep learning model used to perform super-resolution tasks, that is, to restore or generate high-resolution (HR) images from low-resolution (LR) images. Its core goal is to enhance the image details through algorithms and improve the clarity and quality of the image.
[0067] Among them, the semantic segmentation network can be a pre-trained deep learning model for performing image segmentation tasks, that is, segmenting the input image into multiple semantic regions, each region corresponding to a specific category in the image, such as cells, background, etc.
[0068] The semantic segmentation network may include multiple intermediate layers, each of which has a corresponding level. In the semantic segmentation network, the closer the intermediate layer is to the input, the lower the level is, the smaller the network depth is, and the closer the intermediate layer is to the output, the higher the level is, and the greater the network depth is. Optionally, the semantic segmentation network may be a U-net network (U-Net), and the multiple intermediate layers of the semantic segmentation network may include multiple encoding layers and multiple decoding layers symmetrical to the multiple encoding layers.
[0069] In some embodiments, the image to be processed, the super-resolution network, and the semantic segmentation network may be pre-stored locally on the computer device, or stored in a cloud database that is communicatively connected to the computer device, so as to be directly called when needed by the computer device.
[0070] 120. The image to be processed is enhanced through a super-resolution network to obtain multiple feature maps with different resolutions.
[0071] Among them, image enhancement processing may include feature map extraction processing, super-resolution processing, and the like.
[0072] In some embodiments, the super-resolution network may include multiple encoders, each encoder is used to extract feature maps and perform image enhancement processing on the input image to obtain a feature map of the corresponding resolution of the encoder.
[0073] Exemplarily, for example, the super-resolution network includes a first encoder, a second encoder, a third encoder and a fourth encoder arranged in sequence, wherein the extraction resolutions corresponding to the first encoder, the second encoder, the third encoder and the fourth encoder are 1 / 2, 1 / 4, 1 / 8 and 1 / 16 of the initial resolutions corresponding to the image to be processed, respectively.
[0074] The first encoder can perform feature extraction on the image to be processed to obtain a first extraction result with an initial resolution of 1 / 2, and output the first extraction result to a subsequent encoder. At the same time, the first extraction result can be subjected to image enhancement processing to obtain a first feature map.
[0075] The second encoder can perform feature extraction on the first extraction result output by the first encoder to obtain a second extraction result with an initial resolution of 1 / 4, and output the second extraction result to a subsequent encoder, while performing image enhancement processing on the second extraction result to obtain a second feature map.
[0076] The third encoder can perform feature extraction on the second extraction result output by the second encoder to obtain a third extraction result with an initial resolution of 1 / 8, and output the third extraction result to a subsequent encoder, while performing image enhancement processing on the third extraction result to obtain a third feature map.
[0077] The fourth encoder can perform feature extraction on the third extraction result output by the third encoder to obtain a fourth extraction result with an initial resolution of 1 / 16, and output the fourth extraction result to a subsequent encoder, and at the same time perform image enhancement processing on the fourth extraction result to obtain a fourth feature map.
[0078] In this example, the first feature map, the second feature map, the third feature map, and the fourth feature map are multiple feature maps with different resolutions obtained by the super-resolution network.
[0079] 130. In the process of semantic segmentation of the image to be processed by the semantic segmentation network, multiple feature maps are respectively input into designated intermediate layers among multiple intermediate layers, so that the semantic segmentation network outputs an image segmentation result of the image to be processed based on the multiple feature maps.
[0080] The image segmentation result may be a segmentation mask, which represents the prediction values of the image segmentation, and these prediction values represent the prediction results of the model for the category to which each pixel or voxel in the input image belongs.
[0081] In some embodiments, the image to be processed can be input into a semantic segmentation network, and the image to be processed is processed sequentially through each intermediate layer in the semantic segmentation network. In the process of semantic segmentation, multiple feature maps are respectively input into a specified intermediate layer among multiple intermediate layers to assist the semantic segmentation network in performing semantic segmentation, thereby obtaining an image segmentation result. As an example, the above intermediate layer can be used to process features of a resolution. Therefore, for an intermediate layer, a feature map with a resolution consistent with that corresponding to the intermediate layer in multiple feature maps can be input into the intermediate layer, so that the intermediate layer can perform segmentation processing in combination with the feature map, thereby establishing data interaction between the semantic segmentation network and the super-resolution network through the feature map.
[0082] It can be seen that in this embodiment, by obtaining the image to be processed, the super-resolution network and the semantic segmentation network, the semantic segmentation network includes multiple intermediate layers; and the image to be processed is enhanced by the super-resolution network to obtain multiple feature maps of different resolutions; in the process of semantic segmentation of the image to be processed by the semantic segmentation network, the multiple feature maps are respectively input into the designated intermediate layers in the multiple intermediate layers, so that the semantic segmentation network outputs the image segmentation result of the image to be processed based on the multiple feature maps. Since the resolution of the image to be processed can be improved after image enhancement by the super-resolution network, the semantic segmentation network can perform image segmentation based on the image to be processed after the resolution is improved, thereby improving the accuracy of image segmentation. In addition, by inputting multiple feature maps of different resolutions obtained by the super-resolution network into the designated intermediate layers in the multiple intermediate layers of the semantic segmentation network, to assist the semantic segmentation network in image segmentation, the data interaction between the super-resolution network and the semantic segmentation network can be deepened, and the mutual adaptability between the semantic segmentation network and the super-resolution network is improved, thereby ensuring the accuracy of image segmentation.
[0083] Figure 2 FIG. 2 is a flow chart of an image segmentation method provided by another embodiment of the present invention. The method may be executed by a computer device (eg, a server or a user terminal). Figure 2 As shown, the method may include the following contents:
[0084] 210. Obtain an image to be processed, a super-resolution network, and a semantic segmentation network, wherein the semantic segmentation network includes multiple intermediate layers; the multiple intermediate layers include multiple encoding layers and multiple decoding layers, and the multiple encoding layers and the multiple decoding layers correspond one to one.
[0085] In some implementations, in step 210, the specific implementation of obtaining the image to be processed may include:
[0086] 211. Obtain initial images of multiple modalities.
[0087] Among them, modality refers to the characteristics and forms of expression of images as a specific type of data.
[0088] As an example, initial images of multiple modalities can form an initial image set in, It represents the initial image of the mth modality among M modalities, where H, W and C m Respectively represent the height, width and number of channels of the initial image of the mth modality. Among them, all modalities of the initial image set can share a label set in, Represents a set of N semantic labels, which can characterize a specific category in the initial image and are used as image segmentation results in the future.
[0089] 212. Initial images of multiple modalities are stitched together to obtain an image to be processed.
[0090] In some embodiments, the initial images of multiple modalities may be spliced by a preset splicing method to obtain a to-be-processed image, so that the to-be-processed image contains information of multiple modalities, which facilitates subsequent image segmentation processing. Optionally, the preset splicing method may include channel splicing, spatial splicing, and the like.
[0091] In some embodiments, after the initial images of multiple modalities are spliced, the spliced images may be subjected to patch embedding processing to obtain the image to be processed. Patch embedding may divide the image into small blocks and embed them into a high-dimensional space to facilitate subsequent image segmentation.
[0092] 220. The image to be processed is enhanced through a super-resolution network to obtain multiple feature maps with different resolutions.
[0093] In some embodiments, Figure 3 As shown, the image super-resolution network may include a feature extraction module 310 and an enhancement module 320 , wherein the enhancement module 320 is connected to the feature extraction module 310 .
[0094] In this embodiment, the specific implementation of step 220 may include:
[0095] 221. The feature extraction module is used to extract features from the image to be processed, and multiple initial feature maps with different resolutions are obtained.
[0096] For example, please refer again to Figure 3 , the feature extraction module 310 may include multiple feature extraction layers 311, each of which is used to extract an initial feature map of a certain resolution. As the network depth of the multiple feature extraction layers 311 increases, the resolution of the extracted initial feature map gradually decreases. For example, Figure 3 In the super-resolution network shown, the initial features extracted by the first feature extraction layer are Figure 1 The resolution is 1 / 2 of the resolution of the image to be processed. The initial features extracted by the second feature extraction layer Figure 2 The resolution is 1 / 4 of the resolution of the image to be processed. The initial features extracted by the third feature extraction layer Figure 3 The resolution is 1 / 8 of the resolution of the image to be processed. The initial features extracted by the fourth feature extraction layer are Figure 4 The resolution is 1 / 16 of the resolution of the image to be processed.
[0097] In some embodiments, the feature extraction layer 311 may include a sliding window transformation module (SwinTransformer), a dynamic multi-branch (DMB) module and a regularization layer connected in sequence. Among them, the dynamic multi-branch module is used to perform multi-scale feature extraction on the input image to obtain feature maps of different resolutions. The DMB module is used to dynamically adjust the number of branches according to the feature depth, and combine the expansion convolution with different void rates to enhance the multi-scale feature extraction. The regularization layer is used to prevent the super-resolution network from overfitting. Among them, the sliding window transformation module is the input layer of the feature extraction layer, and the regularization layer is the output layer of the feature extraction layer.
[0098] Optionally, the dynamic multi-branch module may include at least one dilated convolutional layer, and the number of dilated convolutional layers in the feature extraction layer is positively correlated with the network depth of the feature extraction layer in the super-resolution network. The specific implementation of step 221 may include:
[0099] For the first feature extraction layer among the multiple feature extraction layers, the images to be processed are respectively convolved through the dilated convolution layer in the first feature extraction layer to obtain the first convolution result, and the first convolution result is fused to obtain the initial feature map output by the first feature extraction layer.
[0100] For feature extraction layers other than the first feature extraction layer in the multiple feature extraction layers, the initial feature maps output by the previous feature extraction layer of the feature extraction layer are convolved by the dilated convolutional layers in the feature extraction layer to obtain a second convolution result, and the second convolution results are fused to obtain the initial feature map output by the feature extraction layer.
[0101] As an example, Figure 4 As shown, the dynamic multi-branch module may include a first convolution layer 410, a first regularization layer 420, an expansion convolution layer 430, a second regularization layer 440 and a second convolution layer 450 connected in sequence.
[0102] In practical applications, the dynamic multi-branch module can receive the output result of the sliding window conversion module connected thereto, and perform preliminary feature extraction on the output result through the first convolution layer 410 to generate relevant residual features in the output result in preparation for subsequent morphological filtering. Optionally, the convolution kernel of the first convolution layer 410 can be a 3×3 convolution kernel.
[0103] Then, the generated relevant residual features can be output to the dilated convolution layer 430 through the first regularization layer 420, wherein the number of dilated convolution layers 430 is multiple, and the residual features will be respectively input to each dilated convolution layer 430 for convolution processing to obtain the convolution result of each dilated convolution layer 430. The dilated convolution layer 430 can be regarded as a branch of the dynamic multi-branch module, and the number of branches (i.e., the number of dilated convolution layers 430) depends on the current network depth of the feature extraction layer corresponding to the dynamic multi-branch module. Figure 3 For example, the number of feature extraction layers is 4, and i can be used to represent the network depth of the feature extraction layer in the super-resolution network, i∈{1, 2, 3, 4}, where the feature extraction layer closest to the input of the super-resolution network has a network depth i of 1, and the feature extraction layer closest to the output of the super-resolution network has a network depth i of 4. Optionally, the convolution kernel of the expanded convolution layer 430 can be 3×3. Optionally, a rectified linear unit (ReLU) for activating and simplifying regional features can also be provided between the first regularization layer 420 and the expanded convolution layer 430.
[0104] As an implementation, if the current network depth of the feature extraction layer is i, the feature extraction layer has i dilated convolution layers 430. Since semantic enhancement enables convolution to form a wider spatial connection, the dilated convolution in each branch may have a different dilation rate. For example, the dilation rate of the dilated convolution may vary with i. If the current network depth is i, the dilation rates of the i dilated convolutions are 1 to (2i-1) respectively.
[0105] After obtaining the convolution results output by each dilated convolution layer 430, the convolution results output by each dilated convolution layer 430 can be fused to obtain a fused convolution result, and then the fused convolution result can be input to the second convolution layer 450 through the second regularization layer 440, so as to adjust the size of the fused convolution result through the second convolution layer 450 to facilitate subsequent feature processing. Among them, the features output by the second convolution layer 450 after adjusting the size of the fused convolution result are processed by the regularization layer in the feature extraction layer with a network depth of i, and the initial feature map corresponding to the feature extraction layer with a network depth of i can be obtained. Optionally, the convolution kernel of the second convolution layer can be a 1×1 convolution kernel.
[0106] Considering that the current processing of multi-scale features is often for features of different scales, as the feature depth increases, the depth of the module is deepened (i.e., the number of modules is increased), which is not comprehensive for multi-scale features because it does not consider changing the width of the module when processing multi-scale features. To this end, in this embodiment, by setting a DMB module in the feature extraction layer, the number of its branches increases with the increase of the network depth, and the capacity of its receptive field will also expand. On the one hand, the information of features of different depths can be appropriately extracted, and on the other hand, the computational complexity of the model can be reduced to a certain extent. In addition, each branch contains a hole convolution with different hole rates to capture feature information of different scales. This multi-scale feature processing method helps the segmentation network to better identify and locate the area that needs to be segmented.
[0107] 222. Perform image enhancement processing on the initial feature map through the enhancement module to obtain a feature map.
[0108] Optionally, the enhancement module can be an image enhancement module based on sparse transformation super-resolution (SR based on sparseTransformer, SR-SP) technology (hereinafter also referred to as SR-SP module). SR-SP is a method that combines super-resolution (SR) technology and sparse Transformer, which aims to improve the resolution of images or feature maps, while using sparse attention mechanism to reduce computational complexity and improve efficiency. It is suitable for scenarios that need to process high-resolution data, such as medical image segmentation, video super-resolution, etc.
[0109] In some embodiments, please refer again to Figure 3 The number of enhancement modules 320 can be multiple, and the multiple enhancement modules 320 correspond to the multiple feature extraction layers 310 one by one. Each enhancement module 320 is used to receive the initial feature map output by its corresponding feature extraction layer 310, and perform image enhancement processing on the initial feature map to obtain a feature map. For example, the initial feature map Figure 1 After image enhancement processing through the enhancement module, the features can be obtained Figure 1 , initial features Figure 2 After image enhancement processing through the enhancement module, the features can be obtained Figure 2 , initial features Figure 3 After image enhancement processing through the enhancement module, the features can be obtained Figure 3 , initial features Figure 4 After image enhancement processing through the enhancement module, the features can be obtained Figure 4 .
[0110] In some embodiments, Figure 5As shown, the enhancement module may include a separation module submodule 510 and an enhancement submodule 520. The specific implementation of step 222 may include:
[0111] 2221. The region containing semantic information in the initial feature map is separated by a separation module submodule to obtain a separation result.
[0112] In some embodiments, the separation submodule 510 may include a downsampling layer 511, a first separation coding layer 512, a second separation coding layer 513, a separation pooling layer 514, a first separation decoding layer 515, a second separation decoding layer 516, a separation convolution layer 517 and a channel separation module 518 arranged in sequence.
[0113] In practical applications, the separation module submodule 510 can first perform downsampling processing on the input initial feature map through the downsampling layer 511, and sequentially pass the downsampling result through the first separation coding layer 512 and the second separation coding layer 513 for encoding processing to obtain the encoding result; then, the encoding result is input to the separation pooling layer 513 for pooling processing of different scales, and the pooling result is sequentially passed through the first separation decoding layer 515 and the second separation decoding layer 516 for decoding processing to obtain the decoding result, wherein the first separation coding layer 512 and the second separation coding layer 513 can be respectively separated from the first separation decoding layer 515 and the second separation decoding layer 516. Layer 516 is jump-connected to realize information interaction with the first separation decoding layer 515 and the second separation decoding layer 516, so that the output results of the first separation coding layer 512 and the second separation coding layer 513 are decoded by assisting the first separation decoding layer 515 and the second separation decoding layer 516. Optionally, the first separation decoding layer 515 and the second separation decoding layer 516 may also be provided with a transposed convolution layer for changing the size of the feature map; then, the decoding result may be input into the separation convolution layer 517 for convolution processing, and the convolution result may be input into the channel separation module 517 for channel separation processing to obtain a separation result.
[0114] As an example, for example, the size of the initial feature map is H i ×W i ×C i , where i is the network depth of the feature extraction module that outputs the initial feature map, which can also be regarded as the stage of the initial feature map. Figure 3 For example, for the initial feature Figure 2 , the corresponding i is 2, for the initial feature Figure 3 , the corresponding i is 3, where H i , W i , C iRespectively represent the length, height, and number of channels of the initial feature map of the i-th stage. After the initial feature map is processed by the downsampling layer 511, the first separation coding layer 512, the second separation coding layer 513, the separation pooling layer 514, the first separation decoding layer 515, the second separation decoding layer 516, and the separation convolution layer 517 in the above-mentioned separation module submodule 510, the size of the feature map can be adjusted to H i ×W i ×(C i +1), then the size is H i ×W i ×(C i +1) is processed by the channel separation module to obtain a feature map of size H i ×W i ×C i The first separated feature map and the size is H i ×W i ×1 second separation feature map, and finally, the size of H i ×W i The second separation feature map of ×1 is used as the separation result of the separation module submodule 510.
[0115] Optionally, the first separation coding layer 512 may include a first coding convolution layer, a second coding convolution layer, and a third coding convolution layer arranged in sequence, wherein the convolution kernels of the first coding convolution layer and the third coding convolution layer may be 1×1 convolution kernels, and the convolution kernel of the second coding convolution layer may be 3×3 convolution kernels. The separation pooling layer 514 may be an Atrous Spatial Pyramid Pooling (ASPP) layer. The first separation decoding layer 515 may include a transposed convolution layer, a first decoding convolution layer, and a second decoding convolution layer arranged in sequence, wherein the transposed convolution layer has a convolution kernel of 1×1, the convolution kernel of the second decoding convolution layer is a 3×3 convolution kernel, and the convolution kernel of the third decoding convolution layer is a 1×1 convolution kernel.
[0116] It can be seen that the role of the separation submodule is to extract the area containing semantic information in the initial feature map to obtain the separation result. That is to say, in the subsequent tasks, it is only necessary to calculate the attention of the area contained in the separation result, thereby reducing the amount of network calculation and effectively improving the efficiency of image segmentation. Among them, the separation result can be used express, Wherein, S represents the image patch size of the separated image. Wherein, the separation result may be the binary sparseness of the i-th stage obtained by downsampling and binarization.
[0117] 2222. Perform super-resolution processing on the separation result through the enhancer module to obtain a feature map.
[0118] In some embodiments, the enhancement submodule 520 may include a first enhancement coding layer 521 , a second enhancement coding layer 522 , an attention module 523 , a first enhancement decoding layer 524 , a second enhancement decoding layer 525 , and an upsampling layer 526 .
[0119] Optionally, the attention module 523 may be a hash transformation module (HSFormer), specifically a Transformer model with a hashing mechanism, for performing attention calculation on the input features. The first enhanced coding layer 521 may include a first hash transformation module, a second hash transformation module, and a coding convolution layer arranged in sequence; the structure of the second enhanced coding layer 522 may be the same as that of the first enhanced coding layer 521, and the first enhanced coding layer 521 and the second enhanced coding layer 522 are used to encode the input features. Optionally, the convolution kernel of the coding convolution layer may be a 4×4 convolution kernel.
[0120] The first enhanced decoding layer 524 may include a transposed convolution layer, a third hash transformation module, and a fourth hash transformation module arranged in sequence; the structure of the second enhanced decoding layer 525 may be the same as that of the first enhanced decoding layer 524. The first enhanced decoding layer 524 and the second enhanced decoding layer 525 may be used to change the size of the input feature and the decoding process.
[0121] Among them, the upsampling layer 526 is used to increase the resolution of the input feature map.
[0122] Among them, the structures of the first hash transformation module, the second hash transformation module, the third hash transformation module, and the fourth hash transformation module may be the same.
[0123] In some embodiments, Figure 6 As shown, the first hash transformation module may include: a first preprocessing module 610 , a super-resolution module 620 , a second preprocessing module 630 , and a feed-forward layer 640 .
[0124] Among them, the first preprocessing module 610 and the second preprocessing module 630 can be layer normalization modules (LayerNormalization), which are used to normalize the input of each layer of the neural network, stabilize the training process and accelerate convergence. The super-resolution module 620 can be a sparse attention mechanism (Hashed Self-Attention, HSA) module based on hash clustering. The HAS module can be used to efficiently process large-scale data (such as high-resolution images or long sequences). Specifically, the HAS module can map data points to different buckets through a hash function, and only calculate attention between data points in the same bucket, thereby significantly reducing the amount of calculation, thereby reducing the complexity of global attention calculation. Among them, the feedforward layer 640.
[0125] For example, in practical applications, HSFormer can convert the initial feature map As the first input, As the second input, by comparing Filter out the Image blocks of information, including The image block of information is input into HSA after regularization to calculate attention, and then the features processed by HSA are input into the feedforward layer 640 to add activation function and perform linear transformation.
[0126] In this embodiment, the efficiency of the network in processing long sequence data can be effectively improved by introducing HSFormer in the enhancement submodule. In addition, HSFormer reduces the amount of calculation and memory consumption by introducing the HSA mechanism, so that the network can efficiently process large-scale data.
[0127] In the process of semantic segmentation of the processed image by the semantic segmentation network, the processed image and multiple feature maps are downsampled respectively through multiple encoding layers to obtain the downsampling result corresponding to each encoding layer.
[0128] For example, Figure 7 As shown, the multiple intermediate layers of the semantic segmentation network 710 may include multiple encoding layers and multiple decoding layers. The multiple encoding layers include a first encoding layer 711, a second encoding layer 712, a third encoding layer 713, and a fourth encoding layer 714 arranged in sequence, wherein the structures of the first encoding layer 711, the second encoding layer 712, the third encoding layer 713, and the fourth encoding layer 714 may be the same. The multiple decoding layers include a first decoding layer 715, a second decoding layer 716, a third decoding layer 717, and a fourth encoding layer 718 arranged in sequence, wherein the structures of the first decoding layer 715, the second decoding layer 716, the third decoding layer 717, and the fourth encoding layer 718 may be the same.
[0129] Optionally, the encoding layer may include a first embedding layer (Embedding), a first encoding regularization layer, a segmented encoding convolution layer, and a second encoding regularization layer arranged in sequence, wherein the convolution kernel of the segmented encoding convolution layer is a 1×1 convolution kernel (i.e., Conv1×1). The decoding layer may include a first segmented decoding convolution layer, a second embedding layer, a first decoding regularization layer, a second segmented decoding convolution layer, and a second decoding regularization layer arranged in sequence, wherein the convolution kernel of the first segmented decoding convolution layer may be Tran Conv2×1, and the convolution kernel of the second segmented decoding convolution layer may be Conv1×1.
[0130] As an example, see again Figure 7 and Figure 3 , the image to be processed can be input to the first coding layer 711, and the image to be processed can be downsampled by the first coding layer 711 to obtain the downsampling result of the first coding layer 711. Figure 3 Features obtained in the examples Figure 1 Input to the second encoding layer 712, through the second encoding layer 712 to feature Figure 1 Down-sampling is performed to obtain the down-sampling result of the second coding layer 712. Figure 3 Features obtained in the examples Figure 2 Input to the third encoding layer 713, through the third encoding layer 713 to feature Figure 2 Down-sampling is performed to obtain the down-sampling result of the third coding layer 713. Figure 3 Features obtained in the examples Figure 3 Input to the fourth encoding layer 714, through the fourth encoding layer 714 to feature Figure 3 Downsampling is performed to obtain a downsampling result of the fourth coding layer 714.
[0131] 240. For each decoding layer, generate input features of the decoding layer according to the downsampling result output by the encoding layer corresponding to the decoding layer.
[0132] In some implementations, specific implementations of step 240 may include:
[0133] 241. For a first decoding layer among multiple decoding layers, a feature map with the smallest resolution among multiple feature maps is fused with a downsampling result output by the encoding layer corresponding to the first decoding layer to obtain input features of the first decoding layer.
[0134] 242. For a decoding layer other than the first decoding layer among the multiple decoding layers, a down-sampling result output by the encoding layer corresponding to the decoding layer is fused with an up-sampling result output by a previous decoding layer of the decoding layer to obtain an input feature of the decoding layer.
[0135] 250. Upsample the input features through the decoding layer to output the upsampled results.
[0136] Continuing with the above example, for the first decoding layer among the multiple decoding layers, that is, the first decoding layer 715, the downsampling result of the fourth encoding layer 714 can be compared with Figure 3 Features obtained in the examples Figure 4 After being fused through the skip connection, they are used as input features of the first decoding layer 715, so that the first decoding layer 715 performs upsampling according to the input features and outputs the upsampling results.
[0137] For the second decoding layer 716, the upsampling result output by the first decoding layer 715 and the downsampling result output by the third encoding layer 713 can be fused through a jump connection as the input feature of the second decoding layer 716, so that the second decoding layer 716 performs upsampling according to the input feature and outputs the upsampling result.
[0138] For the third decoding layer 717, the upsampling result output by the second decoding layer 716 and the downsampling result output by the second encoding layer 712 can be fused through a jump connection as the input feature of the third decoding layer 717, so that the third decoding layer 717 performs upsampling according to the input feature and outputs the upsampling result.
[0139] For the fourth decoding layer 718, the upsampling result output by the third decoding layer 717 and the downsampling result output by the first encoding layer 711 can be fused through a jump connection as the input feature of the fourth decoding layer 718, so that the fourth decoding layer 718 performs upsampling according to the input feature and outputs the upsampling result.
[0140] 260. Determine an image segmentation result according to the upsampling result.
[0141] In some implementations, a specific implementation of step 260 may include: determining an upsampling result output by a last decoding layer among multiple decoding layers as an image segmentation result.
[0142] Continuing with the above example, the up-sampling result output by the fourth decoding layer 718 may be determined as the image segmentation result.
[0143] In some embodiments, the skip connection operation can be performed as follows: Figure 8The connection module shown is executed, and the connection module may include two identical branches, each branch may represent a feature that needs to be connected or fused, and the branch may include a first connection convolution layer 810, a first connection regularization layer 820, a second connection convolution layer 830, and a second connection regularization layer 840 arranged in sequence, wherein the convolution kernel of the first connection convolution layer 810 is Conv1×1, and the convolution kernel of the second connection convolution layer 830 is a 5×5 convolution kernel, specifically DW Conv5×5.
[0144] In practical applications, Figure 8 The third input in can be the input feature of one branch, and the fourth input can be the input feature of another branch. For example, Figure 7 When the first decoding layer 715 in the embodiment is skipped, the down-sampling result of the fourth encoding layer 714 can be used as the third input. Figure 3 Features obtained in the examples Figure 4 As a fourth input. For another example, when a jump connection is performed on the third decoding layer 717, the up-sampling result output by the second decoding layer 716 can be used as the third input, and the down-sampling result output by the second encoding layer 712 can be used as the fourth input.
[0145] In each branch, the input features of the branch can be first downsampled using the first connection convolution layer 810, and then the downsampled features are concatenated with the features after deep processing of the downsampled features by the second connection convolution layer 830, and then the concatenated features are randomly shuffled to obtain a shuffled result. The operation of the other branch is the same, and finally the shuffled results of the two branches are concatenated together to obtain a connection output, that is, the input features of each decoding layer in the above embodiment.
[0146] In this embodiment, the connection module of the skip connection can analyze the feature maps of all channels and assign weights to each channel according to its importance to the task, thereby eliminating redundant information while ensuring the comprehensiveness of the information.
[0147] Considering that the segmentation loss in the related art when training the image segmentation model can only consider the difference between the true value and the predicted value, but does not consider the impact of the features generated in the model on the model performance, it is impossible to adapt the model of the super-resolution network assisted semantic segmentation network. In this embodiment, a loss function is provided to adapt the image segmentation model composed of the super-resolution network and the semantic segmentation network in this embodiment for model training.
[0148] Among them, the loss function It can be composed of two parts: the segmentation loss corresponding to the semantic segmentation network Super-resolution (SR) loss corresponding to the super-resolution network For segmentation loss The definition is as follows:
[0149]
[0150] Among them, Y * The semantic label of the output of the above image segmentation model, Y is the semantic label pre-set for the initial image, where in Represents a set of N semantic tags.
[0151] For the SR loss function The area in the above embodiment can be defined as Reference features To ensure Contains the region of the image that contains semantic information. The definition is as follows:
[0152]
[0153] Among them, C i Indicates the original number of channels of each layer feature, F i Represents the multi-scale features of the i-th level or i-th network depth (also called feature depth), Represents the multi-scale features of the i-th level after SR.
[0154] Then, we can define the constraints Generated SR loss for:
[0155]
[0156] Among them, BCE() represents the binary cross entropy loss, which is used to evaluate all pixels and check each pixel one by one. 2 Represents the mean square error loss, which is used to maximize the peak signal-to-noise ratio in the SR algorithm.
[0157] Finally, the overall loss function of the model is Defined as:
[0158]
[0159] Among them, λ is the adjustment A hyperparameter of relative importance, optionally, λ can be 0.5.
[0160] It can be seen that in this embodiment, by providing a loss function to adapt the image segmentation model composed of the super-resolution network and the semantic segmentation network in this embodiment to perform model training, the accuracy of image segmentation can be effectively improved.
[0161] As an example, the image segmentation method of this embodiment is used to segment medical images as an example to illustrate the segmentation effect of the image segmentation method. The super-resolution network and semantic segmentation network in this embodiment can constitute a super-resolution-assisted dynamic multi-branch segmentation network (SR-DM SegNet).
[0162] Among them, the initial image data set used this time has a total of 1470 cases, of which 1251 cases are used as training sets and 219 cases are used as test sets. Each case contains four modalities: T1-weighted (T1), T1-weighted contrast agent (T1C), T2-weighted (T2), and fluid-attenuated inversion recovery (Flair). The resolution of each modality is 240×240×155, and they share segmentation labels. Due to the complex calculations and high requirements for equipment in processing three-dimensional (3D) data, in order to solve this problem, in this example, the 3D data of each modality of each patient is segmented into two-dimensional (2D) data according to the third dimension. The resolution of these 2D data is 240×240, and the segmentation labels are also operated in the same way to facilitate the one-to-one correspondence between slices and labels.
[0163] like Fig. 9 As shown, Fig. 9 The comparison results of the image segmentation method of this embodiment and the graphic segmentation methods in related technologies (such as VPTTA, nnUNet, etc.) in terms of Dice parameter index, sensitivity (Sensitivity, SEN) index, intersection over Union (IoU) index, pixel-wise prediction value (Pixel-wise Prediction Value, PPV) index and segmentation time (Time) index are shown, among which the Dice parameter (also called Dice similarity coefficient or F1 score) is one of the commonly used indicators for evaluating the similarity between the prediction result and the true label (ground truth) in the image segmentation task. According to Fig. 9 It can be seen that the graphic segmentation results of the image segmentation method of this embodiment are better than those of other segmentation models in terms of Dice parameter index, SEN index, and IoU index.
[0164] like Fig.10 As shown, Fig.10The segmented image of SR-DM SegNet in this embodiment and the segmented image obtained by the image segmentation method in the related art (such as VPTTA, ASSNet, etc.) are shown. Fig.10 It can be clearly seen that the SR-DM SegNet of this embodiment is superior to other segmentation models in capturing detail segmentation.
[0165] like Fig.11 As shown, Fig.11 The loss curves of the SR-DM SegNet and other models in the training process of this embodiment are shown, wherein the first loss is the loss curve of the SR-DM SegNet of this embodiment, and the second loss is the reference loss curve of other models. Fig.11 It can be seen that the loss of SR-DM SegNet gradually decreases and stabilizes during the entire training stage, indicating that the model prevents overfitting while learning effectively.
[0166] like Fig.12 As shown, Fig.12 The comparison results of the image segmentation method of this embodiment and the graphic segmentation method in the related art (such as SRCNN, VDSR, etc.) in terms of Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) are shown. Fig.12 It can be seen that the graphic segmentation result of the image segmentation method of this embodiment is superior to other image segmentation methods in terms of PSNR index and SSIM.
[0167] like Fig.13 As shown, Fig.13 The comparison results of the SR-DM SegNet of this embodiment and other models (such as ASSNet, UNETR, TransBTs, etc.) in terms of the number of parameters, floating point operation parameters (FLOP) and running time (Time) are shown. Fig.13 and Fig. 9 It can be seen that the introduction of the SR network in this embodiment leads to an increase in the number of model parameters, but at the same parameter level, the running time and FLOP of SR-DM SegNet are the smallest. It can be seen that the DMB module and SR-SP module in this embodiment help to simplify the model architecture, thereby reducing its computational complexity and shortening the inference time.
[0168] In summary, considering that the relationship between the SR network and the semantic segmentation network in the related art is a serial structure, that is, the high-resolution image is first obtained by the SR network, and then the segmentation network is used to segment it, and the serial structure lacks a certain robustness, which may lead to the defect that the SR network is incompatible with the segmentation network. And the two networks of the SR network and the semantic segmentation network are only transmitted through the final designed loss function, which may cause the SR network and the semantic segmentation network to converge at the same time due to the different convergence speeds, so that the final model cannot reach its optimal convergence point, resulting in a decrease in model performance. In this embodiment, by inputting multiple feature maps of different resolutions extracted from the SR network into the designated intermediate layers of the multiple intermediate layers in the semantic segmentation network, the semantic segmentation network is assisted in image segmentation, so that the SR network and the semantic segmentation network constitute a serial-parallel hybrid structure, which may effectively improve the accuracy of image segmentation. In addition, considering that directly splicing the complete SR network into the segmentation network will make the constructed network too complicated. In this embodiment, the SR network structure is simplified, and the introduction of HSformer is proposed, thereby simplifying the model architecture and reducing the amount of calculation, which is conducive to improving the efficiency of image segmentation.
[0169] Fig.14 FIG. 1 is a block diagram of an image segmentation device provided by an embodiment of the present invention. Fig.14 As shown, the image segmentation device 1400 includes: an acquisition unit 1410, a feature extraction unit 1420 and a segmentation unit 1430, wherein:
[0170] An acquisition unit 1410 is used to acquire an image to be processed, a super-resolution network, and a semantic segmentation network, where the semantic segmentation network includes multiple intermediate layers;
[0171] A feature extraction unit 1420 is used to perform image enhancement processing on the image to be processed through a super-resolution network to obtain multiple feature maps with different resolutions;
[0172] The segmentation unit 1430 is used to input multiple feature maps into designated intermediate layers among multiple intermediate layers respectively during the process of semantic segmentation of the image to be processed by the semantic segmentation network, so that the semantic segmentation network outputs the image segmentation result of the image to be processed based on the multiple feature maps.
[0173] In another embodiment of the present invention, the multiple intermediate layers include multiple encoding layers and multiple decoding layers, the multiple encoding layers and the multiple decoding layers correspond one to one, and the segmentation unit 1430 is specifically used to:
[0174] Down-sampling the image to be processed and the multiple feature maps respectively through multiple coding layers to obtain the down-sampling result corresponding to each coding layer;
[0175] For each decoding layer, the input features of the decoding layer are generated according to the downsampling results of the encoding layer output corresponding to the decoding layer;
[0176] Upsample the input features through the decoding layer to output the upsampled results;
[0177] The image segmentation result is determined according to the upsampling result.
[0178] In another embodiment of the present invention, the segmentation unit 1430 is further configured to:
[0179] For the first decoding layer among the multiple decoding layers, the feature map with the smallest resolution among the multiple feature maps is fused with the down-sampling result of the encoding layer output corresponding to the first decoding layer to obtain the input feature of the first decoding layer;
[0180] For decoding layers other than the first decoding layer among the multiple decoding layers, a down-sampling result output by the encoding layer corresponding to the decoding layer is fused with an up-sampling result output by a previous decoding layer of the decoding layer to obtain input features of the decoding layer.
[0181] Accordingly, the segmentation unit 1430 is further configured to:
[0182] The upsampling result output by the last decoding layer among the multiple decoding layers is determined as the image segmentation result.
[0183] In another embodiment of the present invention, the super-resolution network includes a feature extraction module and an enhancement module, and the feature extraction unit 1420 is specifically used to:
[0184] The feature extraction module is used to extract features from the image to be processed, and multiple initial feature maps with different resolutions are obtained;
[0185] The initial feature map is enhanced by the enhancement module to obtain the feature map.
[0186] In another embodiment of the present invention, the feature extraction module includes multiple feature extraction layers, the feature extraction layer includes at least one dilated convolution layer, the number of dilated convolution layers in the feature extraction layer is positively correlated with the network depth of the feature extraction layer in the super-resolution network, and the feature extraction unit 1420 is further specifically used for:
[0187] For a first feature extraction layer among the multiple feature extraction layers, convolution processing is performed on the image to be processed respectively through the dilated convolution layer in the first feature extraction layer to obtain a first convolution result, and the first convolution result is fused to obtain an initial feature map output by the first feature extraction layer;
[0188] For feature extraction layers other than the first feature extraction layer in the multiple feature extraction layers, the initial feature maps output by the previous feature extraction layer of the feature extraction layer are convolved by the dilated convolutional layers in the feature extraction layer to obtain a second convolution result, and the second convolution results are fused to obtain the initial feature map output by the feature extraction layer.
[0189] In another embodiment of the present invention, the enhancement module includes a separation module submodule and an enhancement submodule, and the feature extraction unit 1420 is further configured to:
[0190] The region containing semantic information in the initial feature map is separated by the separation module submodule to obtain a separation result;
[0191] The separation results are processed by super-resolution through the enhancer module to obtain the feature map.
[0192] In another embodiment of the present invention, the acquiring unit 1410 is specifically configured to:
[0193] Obtaining initial images of multiple modalities;
[0194] The initial images of multiple modalities are spliced to obtain the image to be processed.
[0195] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps of the image segmentation method in the above-mentioned embodiment, and will not be repeated here.
[0196] Fig.15 FIG. 1 is a block diagram of an electronic device 1500 provided according to an embodiment of the present invention.
[0197] Reference Fig.15 , the electronic device 1500 includes a processing component 1510, which further includes one or more processors, and a memory resource represented by a memory 1520 for storing instructions executable by the processing component 1510, such as an application. The application stored in the memory 1520 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1510 is configured to execute instructions to perform the above-mentioned image segmentation method.
[0198] The electronic device 1500 may also include a power supply component configured to perform power management of the electronic device 1500, a wired or wireless network interface configured to connect the electronic device 1500 to a network, and an input / output (I / O) interface. The electronic device 1500 may operate based on an operating system stored in the memory 1520, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.
[0199] A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of the electronic device 1500, enables the electronic device 1500 to perform the image segmentation method.
[0200] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0201] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0202] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0203] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0204] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0205] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program check codes.
[0206] In addition, it should be noted that the combination of the various technical features in this case is not limited to the combination described in the claims of this case or the combination described in the specific embodiments. All technical features described in this case can be freely combined or combined in any way unless there is a contradiction between them.
[0207] It should be noted that the above examples are only specific embodiments of the present invention, and the present invention is obviously not limited to the above examples, and there are many similar variations. All variations directly derived or associated from the contents disclosed by the technicians in this field should fall within the protection scope of the present invention.
[0208] It should be understood that the first, second, etc. qualifiers mentioned in the embodiments of the present invention are only used to more clearly describe the technical solutions of the embodiments of the present invention and cannot be used to limit the protection scope of the present invention.
[0209] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An image segmentation method, characterized in that: include: Acquire an image to be processed, a super-resolution network, and a semantic segmentation network, wherein the semantic segmentation network includes multiple intermediate layers; Performing image enhancement processing on the image to be processed through the super-resolution network to obtain multiple feature maps with different resolutions; During the process of the semantic segmentation network performing semantic segmentation on the image to be processed, the multiple feature maps are respectively input into designated intermediate layers among the multiple intermediate layers, so that the semantic segmentation network outputs an image segmentation result of the image to be processed based on the multiple feature maps.
2. The method according to claim 1, characterized in that: The plurality of intermediate layers include a plurality of encoding layers and a plurality of decoding layers, the plurality of encoding layers correspond to the plurality of decoding layers one by one, and the inputting the plurality of feature maps into designated intermediate layers among the plurality of intermediate layers respectively so that the semantic segmentation network outputs the image segmentation result of the image to be processed based on the plurality of feature maps includes: Downsampling the image to be processed and the multiple feature maps respectively through the multiple encoding layers to obtain downsampling results corresponding to each of the encoding layers; For each of the decoding layers, generating input features of the decoding layer according to a downsampling result output by the encoding layer corresponding to the decoding layer; Upsampling the input feature through the decoding layer to output an upsampling result; The image segmentation result is determined according to the up-sampling result.
3. The method according to claim 2, characterized in that The step of generating, for each of the decoding layers, input features of the decoding layer according to a downsampling result output by the encoding layer corresponding to the decoding layer, comprises: For a first decoding layer among the multiple decoding layers, fusing a feature map with the smallest resolution among the multiple feature maps with a downsampling result output by a coding layer corresponding to the first decoding layer to obtain input features of the first decoding layer; For a decoding layer other than the first decoding layer among the multiple decoding layers, fusing a down-sampling result output by the encoding layer corresponding to the decoding layer with an up-sampling result output by a previous decoding layer of the decoding layer to obtain an input feature of the decoding layer; Determining the image segmentation result according to the upsampling result includes: An upsampling result output by a last decoding layer among the multiple decoding layers is determined as the image segmentation result.
4. The method according to claim 1, characterized in that: The super-resolution network includes a feature extraction module and an enhancement module. The image to be processed is subjected to image enhancement processing by the super-resolution network to obtain a plurality of feature maps with different resolutions, including: Extracting features of the image to be processed by the feature extraction module to obtain a plurality of initial feature maps with different resolutions; The initial feature map is subjected to image enhancement processing by the enhancement module to obtain the feature map.
5. The method according to claim 4, characterized in that The feature extraction module includes a plurality of feature extraction layers, the feature extraction layer includes at least one dilated convolution layer, the number of dilated convolution layers in the feature extraction layer is positively correlated with the network depth of the feature extraction layer in the super-resolution network, and the feature extraction of the image to be processed by the feature extraction layer is performed to obtain a plurality of initial feature maps with different resolutions, including: For a first feature extraction layer among the multiple feature extraction layers, convolution processing is performed on the image to be processed respectively through the dilated convolution layer in the first feature extraction layer to obtain a first convolution result, and the first convolution result is fused to obtain an initial feature map output by the first feature extraction layer; For feature extraction layers other than the first feature extraction layer among the multiple feature extraction layers, the initial feature maps output by the previous feature extraction layer of the feature extraction layer are convolved through the dilated convolution layers in the feature extraction layers to obtain a second convolution result, and the second convolution results are fused to obtain the initial feature map output by the feature extraction layer.
6. The method according to claim 4, characterized in that The enhancement module includes a separation module submodule and an enhancement submodule, and the enhancement module performs super-resolution processing on the initial feature map to obtain the feature map, including: Separating the region containing semantic information in the initial feature map by the separation module submodule to obtain a separation result; The separation result is subjected to super-resolution processing by the enhancer module to obtain the feature map.
7. The method according to any one of claims 1 to 6, characterized in that The step of obtaining the image to be processed comprises: Obtaining initial images of multiple modalities; The initial images of the multiple modes are spliced to obtain the image to be processed.
8. An image segmentation device, characterized in that: include: An acquisition unit, used to acquire an image to be processed, a super-resolution network, and a semantic segmentation network, wherein the semantic segmentation network includes multiple intermediate layers; A feature extraction unit, used to perform image enhancement processing on the image to be processed through the super-resolution network to obtain a plurality of feature maps with different resolutions; A segmentation unit is used to input the multiple feature maps into designated intermediate layers among the multiple intermediate layers respectively during the process of the semantic segmentation network performing semantic segmentation on the image to be processed, so that the semantic segmentation network outputs the image segmentation result of the image to be processed based on the multiple feature maps.
9. A computer-readable storage medium having computer-executable instructions stored thereon, characterized in that: When the executable instructions are executed by a processor, the image segmentation method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to execute the image segmentation method according to any one of claims 1 to 7.